langsmith-online-eval-engineering
langchain-ai/langchain-skills
Iteratively build and test LangSmith online evaluators by inspecting traces and interviewing users.
What is langsmith-online-eval-engineering?
This skill guides you through creating LangSmith online evaluators one at a time. It walks you through inspecting traces, proposing evaluation criteria, building evaluators (LLM-as-judge or code), testing them against historical traces, and attaching them to your project with run rules.
- Fetch and inspect recent traces from a LangSmith project to understand input/output structure
- Propose evaluation criteria grounded in actual trace data
- Build LLM-as-judge evaluators with custom rubrics and variable mapping
- Build code evaluators using only builtins and standard library
- Test evaluators against historical traces before attaching to production
- Create and configure run rules to attach evaluators to projects with configurable sampling rates
How to install langsmith-online-eval-engineering
npx skills add https://github.com/langchain-ai/langchain-skills --skill langsmith-online-eval-engineering- Access to a LangSmith project with existing traces
- LangSmith API credentials configured
- Understanding of your application's input/output structure
How to use langsmith-online-eval-engineering
- 1.Provide your LangSmith project name when prompted
- 2.Review the inspected trace structure and confirm the input/output fields
- 3.Describe your quality concerns and whether you want a naming prefix for evaluators
- 4.Choose one of the proposed evaluation criteria
- 5.Review the evaluator configuration (schema, prompt, or code) and approve it
- 6.Specify a sampling rate for the run rule (e.g., 1.0 for all traces, 0.1 for 10%)
- 7.Optionally test the evaluator against historical traces to catch errors before production
- 8.Confirm the evaluator appears in your project with the correct attachment and scoring
Use cases
- Create a relevance evaluator to measure whether LLM responses address user questions
- Build a code-quality evaluator to assess whether generated code is syntactically correct
- Develop a safety evaluator to flag responses that violate content policies
- Test a custom scoring function against historical traces before deploying to new traffic
- Set up multiple evaluators for different quality dimensions on the same project
- LangSmith users building evaluation systems for LLM applications
- ML engineers iterating on evaluation criteria based on real trace data
- Teams that need to validate evaluators on historical data before production attachment
langsmith-online-eval-engineering FAQ
This skill creates online evaluators for use within LangSmith via run rules and attachments. Use 'eval-engineering' for Harbor-style offline evaluations instead.
Yes. For code evaluators, the skill executes the function against fetched historical traces. For LLM evaluators, it verifies the configuration. This is the only way to catch errors before new traffic arrives, since run rules only fire on new traces.
The skill tests against historical traces first, which catches runtime errors like wrong field names or missing data. If errors occur in production, you can fix and recreate the evaluator before reattaching.
No, it is optional. If you provide one (e.g., 'myapp-' or 'v2-'), it will be applied consistently to all evaluator names, prompt hub handles, and run rule display names in the session.
The skill asks you to specify this explicitly; it does not default silently. For initial testing, 1.0 (every trace) is recommended. You can use 0.5 (50%), 0.1 (10%), or a custom value depending on your traffic volume and cost constraints.
Full instructions (SKILL.md)
Source of truth, from langchain-ai/langchain-skills.
name: langsmith-online-eval-engineering description: Iteratively inspect traces, interview the user, and create LangSmith online evaluators one at a time. Use specifically for creating online evaluators for use within LangSmith -- use "eval-engineering" for Harbor-style online evaluations.
Online Eval Engineering
Build online evaluators iteratively:
inspect traces and interview user -> propose directions -> user chooses
-> build evaluator -> test, attach, verify -> review and repeat
Read references/langsmith-api.md before creating or modifying evaluators.
1. Inspect traces
Ask the user for their LangSmith project name. Fetch recent root-level traces and print their structure. Read references/trace-inspection.md. Find:
- run name and type;
- available input and output field names;
- the shape and content of the data (truncated samples);
- which fields carry the data an evaluator would need.
Summarize the trace structure in the conversation:
Project: name
Run type: chain | llm | tool | ...
Input fields: field names and what they contain
Output fields: field names and what they contain
Sample: one representative input/output pair (truncated)
Keep the user involved: explain the trace structure and what it implies, then ask only for information the traces cannot establish. For example: "What does this application do?", "What quality concern matters most?", or "What failure should never happen?"
Ask whether the user wants a naming prefix for evaluators in this session (e.g., myapp-, v2-, dogfood-). If they provide one, apply it to all evaluator names, prompt hub handles, and run rule display names. If they decline, use plain descriptive names.
Do not propose evaluators until the trace structure is understood and the user has described their concerns.
2. Discuss and choose an eval direction
Read references/evaluator-design.md. Propose two or three evaluation criteria grounded in the trace data. Apply the naming prefix from step 1 if the user provided one. For each, give:
Name: descriptive evaluator name (with prefix if set)
Type: LLM-as-judge or code
Measures: what quality dimension this evaluates
Scoring: bool, float (0-1), or int; what pass/fail means
Fields needed: which trace fields are used and how
Rationale: why this type and approach
Example:
Name: response-relevance
Type: LLM-as-judge
Measures: whether the response addresses the user's question
Scoring: bool; True = relevant, False = off-topic or non-responsive
Fields needed: input (user question), output (assistant response)
Rationale: relevance is semantic and requires reading comprehension; not decidable by code
Recommend one and ask the user which to build. Do not implement until the user chooses.
3. Build one evaluator
Read references/langsmith-api.md. Build the selected evaluator. Show the full configuration to the user and get approval before executing any API calls.
LLM-as-judge path. Define a ResponseSchema with reasoning first, then the score field. Write prompt messages with a clear rubric that assesses the result, not whether it matches a reference answer. Set variable_mapping using field names discovered in step 1. Present the schema, prompt, variable mapping, and evaluator name for approval. On approval, push the prompt and create the evaluator. Report the evaluator ID.
Code evaluator path. Write a perform_eval(run, example=None) function. It must be self-contained (only builtins and standard library), access run as a dict (run.get("outputs")), and return {"key": ..., "score": ..., "comment": ...}. Present the function code and evaluator name for approval. On approval, create the evaluator. Report the evaluator ID.
4. Test, attach, and verify
Before attaching, ask the user what sampling rate they want (1.0 = every trace, 0.5 = half, 0.1 = 10%, or custom). Do not default silently. If the user is unsure, recommend 1.0 for initial testing.
Ask the user whether they want to test the evaluator against a few existing traces before attaching. Run rules only fire on new traces, so historical testing is the only way to verify before new traffic arrives.
For code evaluators, execute perform_eval directly against fetched root-level traces, passing a dict with inputs, outputs, and attachments keys. This catches runtime errors (wrong field names, dict-vs-object access, missing data) before production. For LLM evaluators, verify the configuration: confirm variable_mapping keys match prompt placeholders, confirm mapped trace fields exist, and check that the mapped data is meaningful.
If testing reveals errors, fix and recreate before attaching. If the user declines testing, proceed to attach.
Create a run rule to connect the evaluator to the tracing project. Apply the user's naming prefix to the display_name. Confirm the evaluator appears in the evaluator list with the correct project attachment. Inspect:
- evaluator attachment and run rule status;
- recent trace feedback and scores (from historical testing or new traces);
- whether scores match expectations for the traces inspected;
- edge case handling (empty output, errored runs, unexpected structure).
Fix and reattach when the evaluator crashes, scores incorrectly, or fails on edge cases. Before approval, confirm the evaluator scored the intended quality dimension, not an infrastructure or data-shape failure.
5. Review with the user
Explain the evaluator name and ID, quality dimension and scoring approach, trace fields used, sampling rate, and any limitation. Ask the user to approve, revise, drop, or choose the next direction. If continuing, reuse the trace findings, then propose a distinct quality dimension.
Invariants
- One quality dimension per evaluator.
- No guessing field names; always inspect traces before implementing.
- Show configuration and get user approval before making API calls.
- Code evaluators must be self-contained: only builtins and standard library.
- Code evaluators receive
runas a plain dict; userun.get("inputs")andrun.get("outputs"), not attribute access. Theexampleparameter must default toNone. - Treat API failures, auth errors, and run rule failures as infrastructure errors, not evaluator bugs.
Related skills
More from langchain-ai/langchain-skills and the wider catalog.

managed-deep-agents
Build, test, and deploy Managed Deep Agents in LangSmith with the mda CLI.

swarm
Dispatch many independent items in parallel across subagents and aggregate results.

deep-agents-core
Framework for building multi-step AI agents with built-in planning, memory, and skill management

deep-agents-memory
Pluggable memory and file backends for Deep Agents: ephemeral, persistent, or hybrid routing.

langsmith-dataset
Create, manage, and upload evaluation datasets to LangSmith for testing and validation.

langsmith-evaluator
Build and run evaluation pipelines for LangSmith with LLM judges and custom code evaluators.