phoenix-evals
arize-ai/phoenix
Build and run evaluators for AI/LLM applications using Phoenix.
What is phoenix-evals?
Phoenix Evals is a framework for creating and executing evaluators that assess AI/LLM application quality. Use it to validate outputs with code-based rules, LLM judges, and human feedback, then integrate evaluations into CI/CD pipelines and production monitoring.
- Build code-based evaluators for deterministic checks (retrieval quality, format validation)
- Create LLM-based evaluators with custom prompts and templates for nuanced assessment
- Run pre-built evaluators for common tasks (RAG, faithfulness, relevance)
- Batch evaluate DataFrames and run experiments across datasets
- Integrate evaluations into pytest, Vitest, and Jest as CI gates
- Validate evaluator accuracy against human labels to ensure >80% TPR/TNR
How to install phoenix-evals
npx skills add https://github.com/arize-ai/phoenix --skill phoenix-evals- Phoenix server running
- Python: phoenix and openai packages
- TypeScript: @arizeai/phoenix-client package
How to use phoenix-evals
- 1.Set up Phoenix server and install language-specific dependencies (Python or TypeScript)
- 2.Instrument your application with tracing to capture spans and traces
- 3.Analyze errors in production traces using error analysis and axial coding
- 4.Define what to evaluate based on observed failures
- 5.Choose a judge model (code-based, LLM, or pre-built)
- 6.Build evaluators using code templates or LLM-based templates
- 7.Validate evaluator accuracy against human labels
- 8.Run evaluators on datasets or in batch mode
Use cases
- Validate RAG retrieval quality and LLM faithfulness to source documents
- Gate CI/CD pipelines by running evaluators as test assertions
- Analyze failure patterns in production traces to identify systematic issues
- Build custom evaluators from observed failures rather than generic metrics
- Generate synthetic datasets and validate evaluator accuracy before deployment
- ML/AI engineers building and deploying LLM applications
- Data scientists validating model outputs and creating evaluation datasets
- DevOps/platform engineers integrating evals into CI/CD pipelines
- QA teams establishing quality gates for AI systems
- Researchers studying evaluator accuracy and failure modes
phoenix-evals FAQ
Start with code-based evaluators for deterministic checks (format, retrieval quality). Use LLM-based evaluators for nuanced judgments (faithfulness, relevance). Validate LLM judges against human labels to ensure >80% accuracy before production use.
Build evaluators, then use the pytest (Python) or Vitest/Jest (TypeScript) integrations to run them as test assertions. Use invariants (assert/expect) for hard requirements that fail the build, and acceptance criteria for quality signals that trend over time.
Experiments are for offline evaluation on datasets to validate evaluator accuracy and compare model versions. Production evals monitor live traces continuously, detect regressions, and trigger guardrails when quality drops.
No, Phoenix Evals requires a running Phoenix server for tracing, storage, and evaluation execution.
Compare evaluator scores against human labels using the validation references. Aim for >80% true positive rate (TPR) and true negative rate (TNR) before deploying to production.
Full instructions (SKILL.md)
Source of truth, from arize-ai/phoenix.
name: phoenix-evals description: Build and run evaluators for AI/LLM applications using Phoenix. license: Apache-2.0 compatibility: Requires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client. metadata: author: oss@arize.com version: "1.0.0" languages: "Python, TypeScript"
Phoenix Evals
Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans.
Quick Reference
Workflows
Starting Fresh: observe-tracing-setup → error-analysis → axial-coding → evaluators-overview
Building Evaluator: fundamentals → common-mistakes-python → evaluators-{code|llm}-{python|typescript} → validation-evaluators-{python|typescript}
RAG Systems: evaluators-rag → evaluators-code-* (retrieval) → evaluators-llm-* (faithfulness)
Gating CI: evaluators-{code|llm}-{python|typescript} → integrations-{pytest|vitest-jest} → production-continuous
Production: production-overview → production-guardrails → production-continuous
Reference Categories
| Prefix | Description |
|---|---|
fundamentals-* | Types, scores, anti-patterns |
observe-* | Tracing, sampling |
error-analysis-* | Finding failures |
axial-coding-* | Categorizing failures |
evaluators-* | Code, LLM, RAG evaluators |
experiments-* | Datasets, running experiments |
integrations-* | Run evals from test runners (pytest, Vitest, Jest) as a CI gate |
validation-* | Validating evaluator accuracy against human labels |
production-* | CI/CD, monitoring |
Key Principles
| Principle | Action |
|---|---|
| Error analysis first | Can't automate what you haven't observed |
| Custom > generic | Build from your failures |
| Code first | Deterministic before LLM |
| Validate judges | >80% TPR/TNR |
| Binary > Likert | Pass/fail, not 1-5 |
| Invariants gate, signals trend | assert/expect hard invariants (CI red); log LLM-judge quality signals and gate the aggregate (acceptance criteria), not every case |
Related skills
More from arize-ai/phoenix and the wider catalog.

phoenix-tracing
OpenInference tracing and instrumentation for Phoenix LLM observability.

phoenix-cli
Debug LLM applications from the terminal using Phoenix CLI—fetch traces, spans, sessions, and run GraphQL queries.

swiftui-design-principles
Design principles for building polished, native-feeling SwiftUI apps and widgets. Use this skill when creating or modifying SwiftUI views, iOS widgets (WidgetKit), or any native Apple UI. Ensures proper spacing, typography, colors, and widget implementations that look and feel like quality apps rather than AI-generated slop.
sd25-pe
Agent skill from arkdocs-en.tos-ap-southeast-1.volces.com.
sd25-pe
Agent skill from arkdocs.tos-cn-beijing.volces.com.

b2b-brand-marketing
Build B2B brand strategy with thought leadership, trust signals, and multi-stakeholder messaging for enterprise sales cycles.