DeepEval
deepeval
Add LLM evaluation, tracing, and dataset management to AI applications with DeepEval.
What is DeepEval?
DeepEval is a framework for evaluating and improving LLM applications. Use it to add evaluation metrics, trace execution, manage test datasets, generate Confident AI reports, and implement iterative improvement loops in your AI projects.
- Run LLM evaluations and quality metrics on AI outputs
- Trace execution and debug LLM application behavior
- Create and manage test datasets for evaluation
- Generate Confident AI reports for performance insights
- Implement iterative improvement loops based on evaluation results
How to install DeepEval
Claude Code install command.
/plugin install deepeval@claude-plugins-officialWorks with
Agents this plugin has been detected to support.
Related plugins
superpowers
Core skills library for Claude Code: TDD, debugging, collaboration patterns, and proven techniques
mattpocock-skills
Matt Pocock's engineering skills: TDD, code review, domain modeling, and spec-driven development for Claude Code.
chrome-devtools-mcp
Automate Chrome, debug deeply, and analyze performance via DevTools and Puppeteer
hyperframes
Write HTML, render it as video with animations, captions, and voiceovers using HyperFrames.
agent-sdk-dev
Development tools for building Claude agents with the Agent SDK.
asana
Connect Claude Code to Asana for task creation, project search, and progress tracking.