llm-evaluation
wshobson/agents
Comprehensive evaluation framework for LLM applications using automated metrics, human feedback, and benchmarking.
What is llm-evaluation?
Implement systematic evaluation strategies for LLM applications across automated metrics (BLEU, ROUGE, BERTScore), human assessment, and LLM-as-Judge approaches. Use this when measuring model performance, comparing prompts or models, detecting regressions, or establishing quality baselines before production deployment.
- Automated metrics for text generation (BLEU, ROUGE, METEOR, BERTScore, Perplexity)
- Classification metrics (Accuracy, Precision/Recall/F1, Confusion Matrix, AUC-ROC)
- Retrieval evaluation (MRR, NDCG, Precision@K, Recall@K)
- Human evaluation framework across accuracy, coherence, relevance, fluency, safety, and helpfulness dimensions
- LLM-as-Judge evaluation (pointwise, pairwise, reference-based, reference-free scoring)
- EvaluationSuite class for running multiple metrics across test cases with aggregated results
How to install llm-evaluation
npx skills add https://github.com/wshobson/agents --skill llm-evaluationHow to use llm-evaluation
- 1.Define your test cases with inputs, expected outputs, and optional context
- 2.Select appropriate metrics for your task (e.g., BLEU for translation, ROUGE for summarization, custom metrics for domain-specific evaluation)
- 3.Create an EvaluationSuite with your chosen metrics
- 4.Call evaluate() with your model and test cases to get aggregated scores and raw results
- 5.Analyze results to identify performance gaps, regressions, or areas for improvement
Use cases
- Comparing performance between different LLM models or prompt variations before deployment
- Detecting performance regressions in production systems through automated regression testing
- Validating that prompt engineering improvements actually improve measurable quality metrics
- Establishing baseline metrics and tracking progress over time for iterative model improvements
- Debugging unexpected model behavior by analyzing error patterns across multiple evaluation dimensions
- ML engineers building and deploying LLM applications
- Data scientists measuring and comparing model performance
- Product teams validating AI application quality before release
- Researchers benchmarking LLM capabilities across different tasks
llm-evaluation FAQ
Use BLEU/ROUGE for text generation with reference answers, BERTScore for semantic similarity, classification metrics for categorical outputs, and retrieval metrics (MRR/NDCG) for RAG systems. For novel tasks, combine automated metrics with human evaluation or LLM-as-Judge.
Use Metric.custom(name, fn) to define your own metric function. The function receives prediction, reference, and context parameters and should return a numeric score.
LLM-as-Judge scales better and is faster for high-volume evaluation, but human evaluation is more reliable for nuanced quality judgments. Use both: human evaluation for validation and LLM-as-Judge for continuous monitoring.
Start with 50-100 representative test cases for initial evaluation. Use more (500+) for production monitoring and regression detection to ensure statistical significance.
Yes, the EvaluationSuite is designed to run multiple metrics simultaneously and aggregate results. This gives you a more complete picture of model performance across different quality dimensions.
Full instructions (SKILL.md)
Source of truth, from wshobson/agents.
name: llm-evaluation description: Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
LLM Evaluation
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
When to Use This Skill
- Measuring LLM application performance systematically
- Comparing different models or prompts
- Detecting performance regressions before deployment
- Validating improvements from prompt changes
- Building confidence in production systems
- Establishing baselines and tracking progress over time
- Debugging unexpected model behavior
Core Evaluation Types
1. Automated Metrics
Fast, repeatable, scalable evaluation using computed scores.
Text Generation:
- BLEU: N-gram overlap (translation)
- ROUGE: Recall-oriented (summarization)
- METEOR: Semantic similarity
- BERTScore: Embedding-based similarity
- Perplexity: Language model confidence
Classification:
- Accuracy: Percentage correct
- Precision/Recall/F1: Class-specific performance
- Confusion Matrix: Error patterns
- AUC-ROC: Ranking quality
Retrieval (RAG):
- MRR: Mean Reciprocal Rank
- NDCG: Normalized Discounted Cumulative Gain
- Precision@K: Relevant in top K
- Recall@K: Coverage in top K
2. Human Evaluation
Manual assessment for quality aspects difficult to automate.
Dimensions:
- Accuracy: Factual correctness
- Coherence: Logical flow
- Relevance: Answers the question
- Fluency: Natural language quality
- Safety: No harmful content
- Helpfulness: Useful to the user
3. LLM-as-Judge
Use stronger LLMs to evaluate weaker model outputs.
Approaches:
- Pointwise: Score individual responses
- Pairwise: Compare two responses
- Reference-based: Compare to gold standard
- Reference-free: Judge without ground truth
Quick Start
from dataclasses import dataclass
from typing import Callable
import numpy as np
@dataclass
class Metric:
name: str
fn: Callable
@staticmethod
def accuracy():
return Metric("accuracy", calculate_accuracy)
@staticmethod
def bleu():
return Metric("bleu", calculate_bleu)
@staticmethod
def bertscore():
return Metric("bertscore", calculate_bertscore)
@staticmethod
def custom(name: str, fn: Callable):
return Metric(name, fn)
class EvaluationSuite:
def __init__(self, metrics: list[Metric]):
self.metrics = metrics
async def evaluate(self, model, test_cases: list[dict]) -> dict:
results = {m.name: [] for m in self.metrics}
for test in test_cases:
prediction = await model.predict(test["input"])
for metric in self.metrics:
score = metric.fn(
prediction=prediction,
reference=test.get("expected"),
context=test.get("context")
)
results[metric.name].append(score)
return {
"metrics": {k: np.mean(v) for k, v in results.items()},
"raw_scores": results
}
# Usage
suite = EvaluationSuite([
Metric.accuracy(),
Metric.bleu(),
Metric.bertscore(),
Metric.custom("groundedness", check_groundedness)
])
test_cases = [
{
"input": "What is the capital of France?",
"expected": "Paris",
"context": "France is a country in Europe. Paris is its capital."
},
]
results = await suite.evaluate(model=your_model, test_cases=test_cases)
Detailed patterns and worked examples
Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.
Related skills
More from wshobson/agents and the wider catalog.

lora-qlora-recipes
Configure LoRA and QLoRA fine-tuning with current best-practice hyperparameters and module targeting.

market-sizing-analysis
Calculate TAM/SAM/SOM for market opportunities using top-down, bottom-up, and value theory methodologies.

memory-forensics
Acquire and analyze memory dumps using Volatility to detect malware, extract artifacts, and investigate incidents.

memory-safety-patterns
Cross-language RAII, ownership, and smart pointer patterns for memory-safe Rust, C++, and C code.

microservices-patterns
Design distributed systems with service boundaries, event-driven communication, and resilience patterns.

ml-pipeline-workflow
Build end-to-end MLOps pipelines from data preparation through model training, validation, and production deployment.