langsmith-dataset
langchain-ai/langsmith-skills
Create, manage, and upload evaluation datasets to LangSmith for testing and validation.
What is langsmith-dataset?
This skill provides CLI and SDK tools to create, manage, and upload evaluation datasets to LangSmith. Use it when you need to build test datasets for agent evaluation, organize examples by type (final_response, single_step, trajectory, RAG), or export traces for analysis.
- List, create, delete, and export datasets via CLI or SDK
- Upload local JSON files as datasets to LangSmith
- Add and manage individual examples within datasets
- Export traces from LangSmith and convert them to dataset format
- Support multiple dataset types: final_response, single_step, trajectory, and RAG
How to install langsmith-dataset
npx skills add https://github.com/langchain-ai/langsmith-skills --skill langsmith-dataset- LANGSMITH_API_KEY environment variable set (required)
- langsmith CLI tool installed via curl script or npm/pip
- Python (langsmith package) or JavaScript (langsmith package) SDK if using programmatic creation
- LANGSMITH_PROJECT environment variable recommended to identify the correct project
How to use langsmith-dataset
- 1.Set LANGSMITH_API_KEY environment variable or use --api-key flag with commands
- 2.Run `langsmith dataset list` to view existing datasets
- 3.Create a new dataset with `langsmith dataset create --name <name>` or upload a JSON file with `langsmith dataset upload <file> --name <name>`
- 4.Add examples using `langsmith example create --dataset <name> --inputs <json> --outputs <json>` or via SDK
- 5.Export datasets with `langsmith dataset export <name> <output-file>` for backup or analysis
Use cases
- Export production traces and convert them into evaluation datasets for regression testing
- Create a RAG evaluation dataset with questions, retrieved chunks, and expected answers
- Build a trajectory dataset to validate that an agent calls tools in the correct order
- Upload a manually curated dataset of test cases for continuous evaluation
- Add new examples to an existing dataset as you discover edge cases
- AI/ML engineers building evaluation frameworks
- LangChain developers testing agent behavior
- Data scientists creating benchmark datasets
- Teams implementing continuous evaluation pipelines
langsmith-dataset FAQ
Use final_response for full conversation testing, single_step for individual node behavior, trajectory for tool call sequences, and RAG for retrieval quality evaluation.
Yes, the CLI prompts for confirmation before delete/overwrite unless you use --yes flag. Always wait for user input in interactive mode unless explicitly requested otherwise.
Export traces with `langsmith trace export`, then process the JSONL files using the provided Python or TypeScript code to extract inputs/outputs, then upload with `langsmith dataset upload`.
Yes, use the langsmith SDK (Python or JavaScript) to create datasets and add examples directly without CLI commands.
The skill will use best judgment to identify the correct project, but setting LANGSMITH_PROJECT is recommended to ensure you query the right project.
Full instructions (SKILL.md)
Source of truth, from langchain-ai/langsmith-skills.
name: langsmith-dataset description: "INVOKE THIS SKILL when creating evaluation datasets, uploading datasets to LangSmith, or managing existing datasets. Covers dataset types (final_response, single_step, trajectory, RAG), CLI management commands, SDK-based creation, and example management. Uses the langsmith CLI tool."
<oneliner> Create, manage, and upload evaluation datasets to LangSmith for testing and validation. </oneliner> <setup> Environment VariablesLANGSMITH_API_KEY=lsv2_pt_your_api_key_here # REQUIRED
LANGSMITH_PROJECT=your-project-name # Check this to know which project has traces
LANGSMITH_WORKSPACE_ID=your-workspace-id # Optional: for org-scoped keys
Authentication is REQUIRED: either set the LANGSMITH_API_KEY environment variable, or pass the --api-key flag to CLI commands (preferred):
langsmith dataset list --api-key $LANGSMITH_API_KEY
IMPORTANT: Always check the environment variables or .env file for LANGSMITH_PROJECT before querying or interacting with LangSmith. This tells you which project contains the relevant traces and data. If the LangSmith project is not available, use your best judgement to identify the right one.
Python Dependencies
pip install langsmith
JavaScript Dependencies
npm install langsmith
CLI Tool
curl -sSL https://raw.githubusercontent.com/langchain-ai/langsmith-cli/main/scripts/install.sh | sh
</setup>
<usage>
Use the `langsmith` CLI to manage datasets and examples.
Dataset Commands
langsmith dataset list- List datasets in LangSmithlangsmith dataset get <name-or-id>- View dataset detailslangsmith dataset create --name <name>- Create a new empty datasetlangsmith dataset delete <name-or-id>- Delete a datasetlangsmith dataset export <name-or-id> <output-file>- Export dataset to local JSON filelangsmith dataset upload <file> --name <name>- Upload a local JSON file as a dataset
Example Commands
langsmith example list --dataset <name>- List examples in a datasetlangsmith example create --dataset <name> --inputs <json>- Add an example to a datasetlangsmith example delete <example-id>- Delete an example
Experiment Commands
langsmith experiment list --dataset <name>- List experiments for a datasetlangsmith experiment get <name>- View experiment results
Common Flags
--limit N- Limit number of results--yes- Skip confirmation prompts (use with caution)
IMPORTANT - Safety Prompts:
- The CLI prompts for confirmation before destructive operations (delete, overwrite)
- If you are running with user input: ALWAYS wait for user input; NEVER use
--yesunless the user explicitly requests it - If you are running non-interactively: Use
--yesto skip confirmation prompts </usage>
<dataset_types_overview> Common evaluation dataset types:
- final_response - Full conversation with expected output. Tests complete agent behavior.
- single_step - Single node inputs/outputs. Tests specific node behavior (e.g., one LLM call or tool).
- trajectory - Tool call sequence. Tests execution path (ordered list of tool names).
- rag - Question/chunks/answer/citations. Tests retrieval quality. </dataset_types_overview>
<creating_datasets>
Creating Datasets
Datasets are JSON files with an array of examples. Each example has inputs and outputs.
From Exported Traces (Programmatic)
Export traces first, then process them into dataset format using code:
# 1. Export traces to JSONL files
langsmith trace export ./traces --project my-project --limit 20 --full --api-key $LANGSMITH_API_KEY
<python>
```python
import json
from pathlib import Path
from langsmith import Client
client = Client()
2. Process traces into dataset examples
examples = [] for jsonl_file in Path("./traces").glob("*.jsonl"): runs = [json.loads(line) for line in jsonl_file.read_text().strip().split("\n")] root = next((r for r in runs if r.get("parent_run_id") is None), None) if root and root.get("inputs") and root.get("outputs"): examples.append({ "trace_id": root.get("trace_id"), "inputs": root["inputs"], "outputs": root["outputs"] })
3. Save locally
with open("/tmp/dataset.json", "w") as f: json.dump(examples, f, indent=2)
</python>
<typescript>
```typescript
import { Client } from "langsmith";
import { readFileSync, writeFileSync, readdirSync } from "fs";
import { join } from "path";
const client = new Client();
// 2. Process traces into dataset examples
const examples: Array<{trace_id?: string, inputs: Record<string, any>, outputs: Record<string, any>}> = [];
const files = readdirSync("./traces").filter(f => f.endsWith(".jsonl"));
for (const file of files) {
const lines = readFileSync(join("./traces", file), "utf-8").trim().split("\n");
const runs = lines.map(line => JSON.parse(line));
const root = runs.find(r => r.parent_run_id == null);
if (root?.inputs && root?.outputs) {
examples.push({ trace_id: root.trace_id, inputs: root.inputs, outputs: root.outputs });
}
}
// 3. Save locally
writeFileSync("/tmp/dataset.json", JSON.stringify(examples, null, 2));
</typescript>
Upload to LangSmith
# Upload local JSON file as a dataset
langsmith dataset upload /tmp/dataset.json --name "My Evaluation Dataset" --api-key $LANGSMITH_API_KEY
Using the SDK Directly
<python> ```python from langsmith import Clientclient = Client()
Create dataset and add examples in one step
dataset = client.create_dataset("My Dataset", description="Evaluation dataset")
client.create_examples( inputs=[{"query": "What is AI?"}, {"query": "Explain RAG"}], outputs=[{"answer": "AI is..."}, {"answer": "RAG is..."}], dataset_name="My Dataset", )
</python>
<typescript>
```typescript
import { Client } from "langsmith";
const client = new Client();
// Create dataset and add examples
const dataset = await client.createDataset("My Dataset", {
description: "Evaluation dataset",
});
await client.createExamples({
inputs: [{ query: "What is AI?" }, { query: "Explain RAG" }],
outputs: [{ answer: "AI is..." }, { answer: "RAG is..." }],
datasetName: "My Dataset",
});
</typescript>
</creating_datasets>
<dataset_structures>
Dataset Structures by Type
Final Response
{"trace_id": "...", "inputs": {"query": "What are the top genres?"}, "outputs": {"response": "The top genres are..."}}
Single Step
{"trace_id": "...", "inputs": {"messages": [...]}, "outputs": {"content": "..."}, "metadata": {"node_name": "model"}}
Trajectory
{"trace_id": "...", "inputs": {"query": "..."}, "outputs": {"expected_trajectory": ["tool_a", "tool_b", "tool_c"]}}
RAG
{"trace_id": "...", "inputs": {"question": "How do I..."}, "outputs": {"answer": "...", "retrieved_chunks": ["..."], "cited_chunks": ["..."]}}
</dataset_structures>
<script_usage>
CLI Usage
# List all datasets
langsmith dataset list --api-key $LANGSMITH_API_KEY
# Get dataset details
langsmith dataset get "My Dataset" --api-key $LANGSMITH_API_KEY
# Create an empty dataset
langsmith dataset create --name "New Dataset" --description "For evaluation" --api-key $LANGSMITH_API_KEY
# Upload a local JSON file
langsmith dataset upload /tmp/dataset.json --name "My Dataset" --api-key $LANGSMITH_API_KEY
# Export a dataset to local file
langsmith dataset export "My Dataset" /tmp/exported.json --limit 100 --api-key $LANGSMITH_API_KEY
# Delete a dataset
langsmith dataset delete "My Dataset" --api-key $LANGSMITH_API_KEY
# List examples in a dataset
langsmith example list --dataset "My Dataset" --limit 10 --api-key $LANGSMITH_API_KEY
# Add an example
langsmith example create --dataset "My Dataset" \
--inputs '{"query": "test"}' \
--outputs '{"answer": "result"}' --api-key $LANGSMITH_API_KEY
# List experiments
langsmith experiment list --dataset "My Dataset" --api-key $LANGSMITH_API_KEY
langsmith experiment get "eval-v1" --api-key $LANGSMITH_API_KEY
</script_usage>
<example_workflow> Complete workflow from traces to uploaded LangSmith dataset:
# 1. Export traces from LangSmith
langsmith trace export ./traces --project my-project --limit 20 --full --api-key $LANGSMITH_API_KEY
# 2. Process traces into dataset format (using Python/JS code)
# See "Creating Datasets" section above
# 3. Upload to LangSmith
langsmith dataset upload /tmp/final_response.json --name "Skills: Final Response" --api-key $LANGSMITH_API_KEY
langsmith dataset upload /tmp/trajectory.json --name "Skills: Trajectory" --api-key $LANGSMITH_API_KEY
# 4. Verify upload
langsmith dataset list --api-key $LANGSMITH_API_KEY
langsmith dataset get "Skills: Final Response" --api-key $LANGSMITH_API_KEY
langsmith example list --dataset "Skills: Final Response" --limit 3 --api-key $LANGSMITH_API_KEY
# 5. Run experiments
langsmith experiment list --dataset "Skills: Final Response" --api-key $LANGSMITH_API_KEY
</example_workflow>
<troubleshooting> **Dataset upload fails:** - Verify LANGSMITH_API_KEY is set - Check JSON file is valid: each element needs `inputs` (and optionally `outputs`) - Dataset name must be unique, or delete existing first with `langsmith dataset delete`Empty dataset after upload:
- Verify JSON file contains an array of objects with
inputskey - Check file isn't empty:
langsmith example list --dataset "Name"
Export has no data:
- Ensure traces were exported with
--fullflag to include inputs/outputs - Verify traces have both
inputsandoutputspopulated
Example count mismatch:
- Use
langsmith dataset get "Name"to check remote count - Compare with local file to verify upload completeness </troubleshooting>
Related skills
More from langchain-ai/langsmith-skills and the wider catalog.

langsmith-evaluator
Build evaluation pipelines for LangSmith with LLM-as-Judge and custom code evaluators.

langsmith-trace
Add tracing to LangChain/LangGraph apps and query trace data via LangSmith CLI.

backend-dev-guidelines
Shared backend guide for Langfuse's Next.js, tRPC, BullMQ, and TypeScript monorepo. Use when creating or reviewing tRPC routers, public REST endpoints, BullMQ queue processors, backend services, middleware, Prisma or ClickHouse data access, OpenTelemetry instrumentation, Zod validation, env configuration, or backend tests across web, worker, or packages/shared.

langfuse
Query and modify Langfuse data via CLI, access documentation, and understand Langfuse features.
backend-code-review
Review backend code for quality, security, maintainability, and best practices based on established checklist rules. Use when the user requests a review, analysis, or improvement of backend files (e.g., `.py`) under the `api/` directory. Do NOT use for frontend files (e.g., `.tsx`, `.ts`, `.js`). Supports pending-change review, code snippets review, and file-focused review.
component-refactoring
Refactor high-complexity React components in Dify frontend using extraction patterns and automated analysis.