PluginBench
Skill
Pass
Audit score 90

llm-evaluation

wshobson/agents

Comprehensive evaluation framework for LLM applications using automated metrics, human feedback, and benchmarking.

What is llm-evaluation?

Implement systematic evaluation strategies for LLM applications across automated metrics (BLEU, ROUGE, BERTScore), human assessment, and LLM-as-Judge approaches. Use this when measuring model performance, comparing prompts or models, detecting regressions, or establishing quality baselines before production deployment.

  • Automated metrics for text generation (BLEU, ROUGE, METEOR, BERTScore, Perplexity)
  • Classification metrics (Accuracy, Precision/Recall/F1, Confusion Matrix, AUC-ROC)
  • Retrieval evaluation (MRR, NDCG, Precision@K, Recall@K)
  • Human evaluation framework across accuracy, coherence, relevance, fluency, safety, and helpfulness dimensions
  • LLM-as-Judge evaluation (pointwise, pairwise, reference-based, reference-free scoring)
  • EvaluationSuite class for running multiple metrics across test cases with aggregated results

How to install llm-evaluation

npx skills add https://github.com/wshobson/agents --skill llm-evaluation
Claude Code
Cursor
Windsurf
Cline

How to use llm-evaluation

  1. 1.Define your test cases with inputs, expected outputs, and optional context
  2. 2.Select appropriate metrics for your task (e.g., BLEU for translation, ROUGE for summarization, custom metrics for domain-specific evaluation)
  3. 3.Create an EvaluationSuite with your chosen metrics
  4. 4.Call evaluate() with your model and test cases to get aggregated scores and raw results
  5. 5.Analyze results to identify performance gaps, regressions, or areas for improvement

Use cases

Good for
  • Comparing performance between different LLM models or prompt variations before deployment
  • Detecting performance regressions in production systems through automated regression testing
  • Validating that prompt engineering improvements actually improve measurable quality metrics
  • Establishing baseline metrics and tracking progress over time for iterative model improvements
  • Debugging unexpected model behavior by analyzing error patterns across multiple evaluation dimensions
Who it's for
  • ML engineers building and deploying LLM applications
  • Data scientists measuring and comparing model performance
  • Product teams validating AI application quality before release
  • Researchers benchmarking LLM capabilities across different tasks

llm-evaluation FAQ

Which metric should I use for my task?

Use BLEU/ROUGE for text generation with reference answers, BERTScore for semantic similarity, classification metrics for categorical outputs, and retrieval metrics (MRR/NDCG) for RAG systems. For novel tasks, combine automated metrics with human evaluation or LLM-as-Judge.

How do I implement custom metrics?

Use Metric.custom(name, fn) to define your own metric function. The function receives prediction, reference, and context parameters and should return a numeric score.

When should I use LLM-as-Judge vs human evaluation?

LLM-as-Judge scales better and is faster for high-volume evaluation, but human evaluation is more reliable for nuanced quality judgments. Use both: human evaluation for validation and LLM-as-Judge for continuous monitoring.

How many test cases do I need?

Start with 50-100 representative test cases for initial evaluation. Use more (500+) for production monitoring and regression detection to ensure statistical significance.

Can I combine multiple metrics?

Yes, the EvaluationSuite is designed to run multiple metrics simultaneously and aggregate results. This gives you a more complete picture of model performance across different quality dimensions.

Full instructions (SKILL.md)

Source of truth, from wshobson/agents.


name: llm-evaluation description: Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.

LLM Evaluation

Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

When to Use This Skill

  • Measuring LLM application performance systematically
  • Comparing different models or prompts
  • Detecting performance regressions before deployment
  • Validating improvements from prompt changes
  • Building confidence in production systems
  • Establishing baselines and tracking progress over time
  • Debugging unexpected model behavior

Core Evaluation Types

1. Automated Metrics

Fast, repeatable, scalable evaluation using computed scores.

Text Generation:

  • BLEU: N-gram overlap (translation)
  • ROUGE: Recall-oriented (summarization)
  • METEOR: Semantic similarity
  • BERTScore: Embedding-based similarity
  • Perplexity: Language model confidence

Classification:

  • Accuracy: Percentage correct
  • Precision/Recall/F1: Class-specific performance
  • Confusion Matrix: Error patterns
  • AUC-ROC: Ranking quality

Retrieval (RAG):

  • MRR: Mean Reciprocal Rank
  • NDCG: Normalized Discounted Cumulative Gain
  • Precision@K: Relevant in top K
  • Recall@K: Coverage in top K

2. Human Evaluation

Manual assessment for quality aspects difficult to automate.

Dimensions:

  • Accuracy: Factual correctness
  • Coherence: Logical flow
  • Relevance: Answers the question
  • Fluency: Natural language quality
  • Safety: No harmful content
  • Helpfulness: Useful to the user

3. LLM-as-Judge

Use stronger LLMs to evaluate weaker model outputs.

Approaches:

  • Pointwise: Score individual responses
  • Pairwise: Compare two responses
  • Reference-based: Compare to gold standard
  • Reference-free: Judge without ground truth

Quick Start

from dataclasses import dataclass
from typing import Callable
import numpy as np

@dataclass
class Metric:
    name: str
    fn: Callable

    @staticmethod
    def accuracy():
        return Metric("accuracy", calculate_accuracy)

    @staticmethod
    def bleu():
        return Metric("bleu", calculate_bleu)

    @staticmethod
    def bertscore():
        return Metric("bertscore", calculate_bertscore)

    @staticmethod
    def custom(name: str, fn: Callable):
        return Metric(name, fn)

class EvaluationSuite:
    def __init__(self, metrics: list[Metric]):
        self.metrics = metrics

    async def evaluate(self, model, test_cases: list[dict]) -> dict:
        results = {m.name: [] for m in self.metrics}

        for test in test_cases:
            prediction = await model.predict(test["input"])

            for metric in self.metrics:
                score = metric.fn(
                    prediction=prediction,
                    reference=test.get("expected"),
                    context=test.get("context")
                )
                results[metric.name].append(score)

        return {
            "metrics": {k: np.mean(v) for k, v in results.items()},
            "raw_scores": results
        }

# Usage
suite = EvaluationSuite([
    Metric.accuracy(),
    Metric.bleu(),
    Metric.bertscore(),
    Metric.custom("groundedness", check_groundedness)
])

test_cases = [
    {
        "input": "What is the capital of France?",
        "expected": "Paris",
        "context": "France is a country in Europe. Paris is its capital."
    },
]

results = await suite.evaluate(model=your_model, test_cases=test_cases)

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.