langsmith-trace
langchain-ai/langsmith-skills
Add tracing to LLM apps and query trace data with LangSmith CLI.
What is langsmith-trace?
This skill enables you to instrument LangChain/LangGraph applications with automatic tracing, or manually add tracing to custom LLM pipelines using decorators and wrappers. Use it to capture execution trees, debug agent behavior, and export traces for analysis and dataset creation.
- Automatically trace LangChain/LangGraph applications by setting environment variables
- Manually instrument Python and TypeScript apps with @traceable decorators and client wrappers
- Query and list traces (complete execution trees) and runs (individual nodes) via CLI
- Export traces to JSONL files for offline analysis and dataset generation
- Manage datasets, examples, and evaluators through the CLI
- Filter traces by project, name, metadata, and other criteria
How to install langsmith-trace
npx skills add https://github.com/langchain-ai/langsmith-skills --skill langsmith-trace- LangSmith API key (set as LANGSMITH_API_KEY environment variable)
- LangSmith CLI installed via provided install script
- OpenAI API key or other LLM provider credentials (for traced applications)
- Optional: LANGSMITH_PROJECT environment variable to specify target project
How to use langsmith-trace
- 1.Install the LangSmith CLI using the provided curl script
- 2.Set LANGSMITH_API_KEY and other required environment variables
- 3.For LangChain apps: set LANGSMITH_TRACING=true and run normally
- 4.For custom apps: import @traceable decorator or wrapper functions and wrap your LLM client calls
- 5.Use langsmith trace list to view traces, langsmith trace get <trace_id> for details, and langsmith trace export to download data
Use cases
- Debug multi-step agent workflows by examining complete trace hierarchies
- Export production traces to create evaluation datasets for model testing
- Instrument custom LLM pipelines (non-LangChain) with nested function tracing
- Compare trace performance across different prompts or model versions
- Analyze tool-use patterns and LLM call sequences in complex applications
- LLM application developers building with LangChain or custom frameworks
- ML engineers evaluating and iterating on agent behavior
- DevOps/SRE teams monitoring production LLM systems
- Researchers analyzing LLM decision-making and trajectory data
langsmith-trace FAQ
A trace is a complete execution tree (root run + all children), representing one full agent invocation. A run is a single node (one LLM call, tool call, etc.). Query traces first for complete context; use runs for specific node-level analysis.
No. For LangChain/LangGraph apps, tracing is automatic—just set LANGSMITH_TRACING=true and LANGSMITH_API_KEY environment variables.
Use the @traceable decorator (Python) or traceable() wrapper (TypeScript) on your functions, and wrap your LLM client with wrap_openai() or wrapOpenAI(). Nested functions should also be decorated for visibility.
Yes. Use langsmith trace export to save traces to JSONL files, one file per trace. You can then load and analyze them locally or upload to create datasets.
Check the LANGSMITH_PROJECT environment variable or .env file. If unset, the CLI defaults to 'default'. Specify a project explicitly with --project flag in CLI commands.
Full instructions (SKILL.md)
Source of truth, from langchain-ai/langsmith-skills.
name: langsmith-trace description: "INVOKE THIS SKILL when working with LangSmith tracing OR querying traces. Covers adding tracing to applications and querying/exporting trace data. Uses the langsmith CLI tool."
<oneliner> Two main topics: **adding tracing** to your application, and **querying traces** for debugging and analysis. Python and Javascript implementations are both supported. </oneliner> <setup> Environment VariablesLANGSMITH_API_KEY=lsv2_pt_your_api_key_here # REQUIRED
LANGSMITH_PROJECT=your-project-name # Optional: default project
LANGSMITH_WORKSPACE_ID=your-workspace-id # Optional: for org-scoped keys
Authentication is REQUIRED: either set the LANGSMITH_API_KEY environment variable, or pass the --api-key flag to CLI commands (preferred):
langsmith trace list --project my-project --api-key $LANGSMITH_API_KEY
IMPORTANT: Always check the environment variables or .env file for LANGSMITH_PROJECT before querying or interacting with LangSmith. This tells you which project contains the relevant traces and data. If the LangSmith project is not available, use your best judgement to identify the right one.
CLI Tool
curl -sSL https://raw.githubusercontent.com/langchain-ai/langsmith-cli/main/scripts/install.sh | sh
</setup>
<cli_feedback> After an actual CLI failure, undocumented workaround, confusing output, or missing capability, submit one concise product-feedback note per distinct issue in the task. Do not report routine successes or failures in the traced application itself.
CLI requirement: langsmith feedback requires LangSmith CLI v0.2.58 or later. Check the installed version with langsmith --version.
Check langsmith feedback --help for feedback <note> and --category; if unavailable, skip feedback without raw HTTP or unreleased builds. Use the existing authenticated profile, endpoint, and workspace. Feedback goes to LangSmith Cloud, including through the BYOC relay; skip standalone self-hosted. Respect user/organization restrictions and ask first if permission to send feedback is unclear.
Summarize expected versus observed CLI behavior and any workaround in your own words. Never send secrets, customer data, trace payloads, prompts, full stack traces, copied command output, raw arguments, environment-variable values, local paths, or resource identifiers. The CLI adds version/OS/architecture, but does not redact your note; skip it if it cannot be safely redacted.
Choose bug, feature-request, usability, documentation, or other. This is CLI product feedback, not run evaluation feedback. Example shape only—do not submit unless actually encountered:
langsmith feedback --category usability --format json "The trace list output made it hard to distinguish root runs from child runs."
Do not retry a failed or rate-limited feedback submission, switch credentials/endpoints to bypass a failure, or block the original task on feedback. </cli_feedback>
<trace_langchain_oss> For LangChain/LangGraph apps, tracing is automatic. Just set environment variables:
export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY=<your-api-key>
export OPENAI_API_KEY=<your-openai-api-key> # or your LLM provider's key
Optional variables:
LANGSMITH_PROJECT- specify project name (defaults to "default")LANGCHAIN_CALLBACKS_BACKGROUND=false- use for serverless to ensure traces complete before function exit (Python) </trace_langchain_oss>
<trace_other_frameworks> For non-LangChain apps, if the framework has native OpenTelemetry support, use LangSmith's OpenTelemetry integration.
If the app is NOT using a framework, or using one without automatic OTel support, use the traceable decorator/wrapper and wrap your LLM client.
<python> Use @traceable decorator and wrap_openai() for automatic tracing. ```python from langsmith import traceable from langsmith.wrappers import wrap_openai from openai import OpenAIclient = wrap_openai(OpenAI())
@traceable def my_llm_pipeline(question: str) -> str: resp = client.chat.completions.create( model="gpt-4o-mini", messages=[{"role": "user", "content": question}], ) return resp.choices[0].message.content
Nested tracing example
@traceable def rag_pipeline(question: str) -> str: docs = retrieve_docs(question) return generate_answer(question, docs)
@traceable(name="retrieve_docs") def retrieve_docs(query: str) -> list[str]: return docs
@traceable(name="generate_answer") def generate_answer(question: str, docs: list[str]) -> str: return client.chat.completions.create(...)
</python>
<typescript>
Use traceable() wrapper and wrapOpenAI() for automatic tracing.
```typescript
import { traceable } from "langsmith/traceable";
import { wrapOpenAI } from "langsmith/wrappers";
import OpenAI from "openai";
const client = wrapOpenAI(new OpenAI());
const myLlmPipeline = traceable(async (question: string): Promise<string> => {
const resp = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: question }],
});
return resp.choices[0].message.content || "";
}, { name: "my_llm_pipeline" });
// Nested tracing example
const retrieveDocs = traceable(async (query: string): Promise<string[]> => {
return docs;
}, { name: "retrieve_docs" });
const generateAnswer = traceable(async (question: string, docs: string[]): Promise<string> => {
const resp = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: `${question}\nContext: ${docs.join("\n")}` }],
});
return resp.choices[0].message.content || "";
}, { name: "generate_answer" });
const ragPipeline = traceable(async (question: string): Promise<string> => {
const docs = await retrieveDocs(question);
return await generateAnswer(question, docs);
}, { name: "rag_pipeline" });
</typescript>
Best Practices:
- Apply traceable to all nested functions you want visible in LangSmith
- Wrapped clients auto-trace all calls —
wrap_openai()/wrapOpenAI()records every LLM call - Name your traces for easier filtering
- Add metadata for searchability </trace_other_frameworks>
<traces_vs_runs>
Use the langsmith CLI to query trace data.
Understanding the difference is critical:
- Trace = A complete execution tree (root run + all child runs). A trace represents one full agent invocation with all its LLM calls, tool calls, and nested operations.
- Run = A single node in the tree (one LLM call, one tool call, etc.)
Generally, query traces first — they provide complete context and preserve hierarchy needed for trajectory analysis and dataset generation. </traces_vs_runs>
<command_structure> Two command groups with consistent behavior:
langsmith
├── trace (operations on trace trees - USE THIS FIRST)
│ ├── list - List traces (filters apply to root run)
│ ├── get - Get single trace with full hierarchy
│ └── export - Export traces to JSONL files (one file per trace)
│
├── run (operations on individual runs - for specific analysis)
│ ├── list - List runs (flat, filters apply to any run)
│ ├── get - Get single run
│ └── export - Export runs to single JSONL file (flat)
│
├── dataset (dataset operations)
│ ├── list - List datasets
│ ├── get - Get dataset details
│ ├── create - Create empty dataset
│ ├── delete - Delete dataset
│ ├── export - Export dataset to file
│ └── upload - Upload local JSON as dataset
│
├── example (example operations)
│ ├── list - List examples in a dataset
│ ├── create - Add example to a dataset
│ └── delete - Delete an example
│
├── evaluator (evaluator operations)
│ ├── list - List evaluators
│ ├── upload - Upload evaluator
│ └── delete - Delete evaluator
│
├── experiment (experiment operations)
│ ├── list - List experiments
│ └── get - Get experiment results
│
├── thread (thread operations)
│ ├── list - List conversation threads
│ └── get - Get thread details
│
└── project (project operations)
└── list - List tracing projects
Key differences:
traces * | runs * | |
|---|---|---|
| Filters apply to | Root run only | Any matching run |
--run-type | Not available | Available |
| Returns | Full hierarchy | Flat list |
| Export output | Directory (one file/trace) | Single file |
| </command_structure> |
<querying_traces>
Query traces using the langsmith CLI. Commands are language-agnostic.
# List recent traces (most common operation)
langsmith trace list --limit 10 --project my-project --api-key $LANGSMITH_API_KEY
# List traces with metadata (timing, tokens, costs)
langsmith trace list --limit 10 --include-metadata --api-key $LANGSMITH_API_KEY
# Filter traces by time
langsmith trace list --last-n-minutes 60 --api-key $LANGSMITH_API_KEY
langsmith trace list --since 2025-01-20T10:00:00Z --api-key $LANGSMITH_API_KEY
# Get specific trace with full hierarchy
langsmith trace get <trace-id> --api-key $LANGSMITH_API_KEY
# List traces and show hierarchy inline
langsmith trace list --limit 5 --show-hierarchy --api-key $LANGSMITH_API_KEY
# Export traces to JSONL (one file per trace, includes all runs)
langsmith trace export ./traces --limit 20 --full --api-key $LANGSMITH_API_KEY
# Filter traces by performance
langsmith trace list --min-latency 5.0 --limit 10 --api-key $LANGSMITH_API_KEY # Slow traces (>= 5s)
langsmith trace list --error --last-n-minutes 60 --api-key $LANGSMITH_API_KEY # Failed traces
# List specific run types (flat list)
langsmith run list --run-type llm --limit 20 --api-key $LANGSMITH_API_KEY
</querying_traces>
<filters> All commands support these filters (all AND together):Basic filters:
--trace-ids abc,def- Filter to specific traces--limit N- Max results--project NAME- Project name--last-n-minutes N- Time filter--since TIMESTAMP- Time filter (ISO format)--error / --no-error- Error status--name PATTERN- Name contains (case-insensitive)
Performance filters:
--min-latency SECONDS- Minimum latency (e.g.,5for >= 5s)--max-latency SECONDS- Maximum latency--min-tokens N- Minimum total tokens--tags tag1,tag2- Has any of these tags
Advanced filter:
--filter QUERY- Raw LangSmith filter query for complex cases (feedback, metadata, etc.)
# Filter traces by feedback score using raw LangSmith query
langsmith trace list --filter 'and(eq(feedback_key, "correctness"), gte(feedback_score, 0.8))' --api-key $LANGSMITH_API_KEY
</filters>
<export_format>
Export creates .jsonl files (one run per line) with these fields:
{"run_id": "...", "trace_id": "...", "name": "...", "run_type": "...", "parent_run_id": "...", "inputs": {...}, "outputs": {...}}
Use --include-io or --full to include inputs/outputs (required for dataset generation).
</export_format>
Related skills
More from langchain-ai/langsmith-skills and the wider catalog.

langsmith-dataset
Create, manage, and upload evaluation datasets to LangSmith for testing and validation.

langsmith-evaluator
Build and run evaluation pipelines for LangSmith with LLM judges and custom code evaluators.

backend-dev-guidelines
Guidelines for building and reviewing Langfuse backend code: tRPC routers, APIs, queues, services, and tests.

langfuse
Interact with Langfuse for AI observability, tracing, prompt management, datasets, and experimentation.
backend-code-review
Review backend code for quality, security, maintainability, and best practices.
component-refactoring
Refactor high-complexity React components in Dify frontend using extraction patterns and automated analysis.