DeepEval
deepeval
Add LLM evaluation, tracing, and dataset management to AI applications with DeepEval.
What is DeepEval?
DeepEval is a framework for evaluating and improving LLM applications. Use it to add evaluation metrics, trace execution, manage test datasets, generate Confident AI reports, and implement iterative improvement loops in your AI projects.
- Run LLM evaluations and quality metrics on AI outputs
- Trace execution and debug LLM application behavior
- Create and manage test datasets for evaluation
- Generate Confident AI reports for performance insights
- Implement iterative improvement loops based on evaluation results
How to install DeepEval
Claude Code install command.
/plugin install deepeval@claude-plugins-officialWorks with
Agents this plugin has been detected to support.
Related plugins

streaming-skills-plugin
Kafka, Flink, and Schema Registry skills for streaming application developers.

crowdsec
Operational and API skills for CrowdSec security engine, bouncers, WAF, and bot detection.

dash0
OpenTelemetry observability for Claude Code sessions with automatic tracing.

databricks
Databricks integration for CLI, data discovery, model serving, pipelines, and serverless migration.

datadog
Query Datadog logs, metrics, traces, and dashboards directly in Claude Code via MCP.

datahub-skills
DataHub development toolkit with connector planning, PR review, catalog search, and metadata management