PluginBench
Skill
Review
Audit score 70

phoenix-evals

arize-ai/phoenix

Build and run evaluators for AI/LLM applications using Phoenix.

What is phoenix-evals?

Phoenix Evals is a framework for creating and executing evaluators that assess AI/LLM application quality. Use it to validate outputs with code-based rules, LLM judges, and human feedback, then integrate evaluations into CI/CD pipelines and production monitoring.

  • Build code-based evaluators for deterministic checks (retrieval quality, format validation)
  • Create LLM-based evaluators with custom prompts and templates for nuanced assessment
  • Run pre-built evaluators for common tasks (RAG, faithfulness, relevance)
  • Batch evaluate DataFrames and run experiments across datasets
  • Integrate evaluations into pytest, Vitest, and Jest as CI gates
  • Validate evaluator accuracy against human labels to ensure >80% TPR/TNR

How to install phoenix-evals

npx skills add https://github.com/arize-ai/phoenix --skill phoenix-evals
Prerequisites
  • Phoenix server running
  • Python: phoenix and openai packages
  • TypeScript: @arizeai/phoenix-client package
Claude Code
Cursor
Windsurf
Cline

How to use phoenix-evals

  1. 1.Set up Phoenix server and install language-specific dependencies (Python or TypeScript)
  2. 2.Instrument your application with tracing to capture spans and traces
  3. 3.Analyze errors in production traces using error analysis and axial coding
  4. 4.Define what to evaluate based on observed failures
  5. 5.Choose a judge model (code-based, LLM, or pre-built)
  6. 6.Build evaluators using code templates or LLM-based templates
  7. 7.Validate evaluator accuracy against human labels
  8. 8.Run evaluators on datasets or in batch mode

Use cases

Good for
  • Validate RAG retrieval quality and LLM faithfulness to source documents
  • Gate CI/CD pipelines by running evaluators as test assertions
  • Analyze failure patterns in production traces to identify systematic issues
  • Build custom evaluators from observed failures rather than generic metrics
  • Generate synthetic datasets and validate evaluator accuracy before deployment
Who it's for
  • ML/AI engineers building and deploying LLM applications
  • Data scientists validating model outputs and creating evaluation datasets
  • DevOps/platform engineers integrating evals into CI/CD pipelines
  • QA teams establishing quality gates for AI systems
  • Researchers studying evaluator accuracy and failure modes

phoenix-evals FAQ

Should I use code-based or LLM-based evaluators?

Start with code-based evaluators for deterministic checks (format, retrieval quality). Use LLM-based evaluators for nuanced judgments (faithfulness, relevance). Validate LLM judges against human labels to ensure >80% accuracy before production use.

How do I integrate evaluators into CI/CD?

Build evaluators, then use the pytest (Python) or Vitest/Jest (TypeScript) integrations to run them as test assertions. Use invariants (assert/expect) for hard requirements that fail the build, and acceptance criteria for quality signals that trend over time.

What's the difference between experiments and production evals?

Experiments are for offline evaluation on datasets to validate evaluator accuracy and compare model versions. Production evals monitor live traces continuously, detect regressions, and trigger guardrails when quality drops.

Can I use Phoenix Evals without a Phoenix server?

No, Phoenix Evals requires a running Phoenix server for tracing, storage, and evaluation execution.

How do I validate that my evaluator is accurate?

Compare evaluator scores against human labels using the validation references. Aim for >80% true positive rate (TPR) and true negative rate (TNR) before deploying to production.

Full instructions (SKILL.md)

Source of truth, from arize-ai/phoenix.


name: phoenix-evals description: Build and run evaluators for AI/LLM applications using Phoenix. license: Apache-2.0 compatibility: Requires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client. metadata: author: oss@arize.com version: "1.0.0" languages: "Python, TypeScript"

Phoenix Evals

Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans.

Quick Reference

TaskFiles
Setupsetup-python, setup-typescript
Decide what to evaluateevaluators-overview
Choose a judge modelfundamentals-model-selection
Use pre-built evaluatorsevaluators-pre-built
Build code evaluatorevaluators-code-python, evaluators-code-typescript
Build LLM evaluatorevaluators-llm-python, evaluators-llm-typescript, evaluators-custom-templates
Batch evaluate DataFrameevaluate-dataframe-python
Run experimentexperiments-running-python, experiments-running-typescript
Run evals in a test runner (CI gate)integrations-pytest, integrations-vitest-jest
Create datasetexperiments-datasets-python, experiments-datasets-typescript
Generate synthetic dataexperiments-synthetic-python, experiments-synthetic-typescript
Validate evaluator accuracyvalidation, validation-evaluators-python, validation-evaluators-typescript
Export spansobserve-tracing-setup
Write a span filter (SpanQuery().where)filter-expressions
Sample traces for reviewobserve-sampling-python, observe-sampling-typescript
Analyze errorserror-analysis, error-analysis-multi-turn, axial-coding
RAG evalsevaluators-rag
Avoid common mistakescommon-mistakes-python, fundamentals-anti-patterns
Productionproduction-overview, production-guardrails, production-continuous

Workflows

Starting Fresh: observe-tracing-setup → error-analysis → axial-coding → evaluators-overview

Building Evaluator: fundamentals → common-mistakes-python → evaluators-{code|llm}-{python|typescript} → validation-evaluators-{python|typescript}

RAG Systems: evaluators-rag → evaluators-code-* (retrieval) → evaluators-llm-* (faithfulness)

Gating CI: evaluators-{code|llm}-{python|typescript} → integrations-{pytest|vitest-jest} → production-continuous

Production: production-overview → production-guardrails → production-continuous

Reference Categories

PrefixDescription
fundamentals-*Types, scores, anti-patterns
observe-*Tracing, sampling
error-analysis-*Finding failures
axial-coding-*Categorizing failures
evaluators-*Code, LLM, RAG evaluators
experiments-*Datasets, running experiments
integrations-*Run evals from test runners (pytest, Vitest, Jest) as a CI gate
validation-*Validating evaluator accuracy against human labels
production-*CI/CD, monitoring

Key Principles

PrincipleAction
Error analysis firstCan't automate what you haven't observed
Custom > genericBuild from your failures
Code firstDeterministic before LLM
Validate judges>80% TPR/TNR
Binary > LikertPass/fail, not 1-5
Invariants gate, signals trendassert/expect hard invariants (CI red); log LLM-judge quality signals and gate the aggregate (acceptance criteria), not every case