dataset-curation
wshobson/agents
Prepare, format, and validate datasets for supervised fine-tuning and preference training.
What is dataset-curation?
Converts raw examples into training-ready JSONL datasets with proper chat templates, loss masking, and packing. Use this after selecting a fine-tuning method to prepare data, validate formatting, mix synthetic data safely, and generate the dataset card required before launching training.
- Format selection by target method (SFT, DPO, ORPO, KTO, GRPO/RLVR) with row-count guidance
- Apply chat templates before packing and mask loss to assistant turns only to avoid silent failures
- Configure sequence packing to eliminate 40–70% of padding waste and recompute LR schedules accordingly
- Mix synthetic data while maintaining ≥25% real data to prevent quality collapse
- Validate packed sequences by decoding 5–10 examples before full training runs
- Generate a required dataset card with provenance, counts, dedup method, template, and packing config
How to install dataset-curation
npx skills add https://github.com/wshobson/agents --skill dataset-curation- Output from `finetuning-method-selection` skill (routing decision on which method to use)
- Raw training examples (demonstrations, preference judgments, or task prompts)
- Target model's chat template string or identifier
- Tokenizer for the target model
How to use dataset-curation
- 1.Select the JSONL format matching your fine-tuning method from the format table (SFT instruct/conversation, DPO preference pairs, etc.)
- 2.Apply the target model's chat template to each example before any concatenation or packing
- 3.Configure loss masking to -100 over system/user turns and template markers, keeping only assistant-turn tokens
- 4.Decode and manually inspect 5–10 packed sequences to verify example boundaries, template integrity, and loss masks are correct
- 5.Ensure your final dataset contains ≥25% real (non-generated-for-this-task) data to guard against quality collapse
- 6.Complete the dataset card with all six required fields: provenance, counts, synthetic/real ratio, dedup method, template used, and packing config
Use cases
- Converting raw demonstrations into multi-turn ChatML format for SFT with proper role masking
- Preparing preference pairs (prompt, chosen, rejected) for DPO or ORPO training
- Mixing Magpie-generated synthetic data with real examples while meeting the 25% real-data floor
- Configuring packing to fit more examples per batch and recalculating training schedules
- Validating that chat templates match between training and inference to avoid eval degradation
- ML engineers preparing datasets for model fine-tuning
- Researchers mixing synthetic and real data for preference optimization
- Teams implementing SFT, DPO, or ORPO training pipelines
- Anyone debugging silent training failures caused by template or packing mismatches
dataset-curation FAQ
Without packing, 40–70% of compute is wasted on padding between variable-length examples. Packing concatenates multiple examples into one sequence up to max length, cutting most waste. However, packing changes batch semantics — recompute your LR schedule milestones against packed-sequence count, not example count.
~1,000+ rows is the recommended floor, not a target. Below it, a handful of low-quality or duplicate examples can dominate the gradient. Above 1,000, prioritize quality over quantity — a smaller verified, deduplicated set beats a larger noisy one.
A model trained against one chat template but served or evaluated with a different one degrades without erroring. Verify the same template string is applied at training, inference, and eval time. This is a top silent failure mode that only surfaces in eval quality, hours after training completes.
Keep ≥25% real data as a collapse guard. Training on a growing share of model-generated data without a real-data floor drives measurable quality collapse over successive generations. General-domain replay rows count toward this floor.
Six mandatory fields: provenance (traceable to source), counts (total and per split), synthetic/real ratio, dedup method, exact template used, and packing config with confirmation of manual inspection. A dataset missing any field is not ready for training.
Full instructions (SKILL.md)
Source of truth, from wshobson/agents.
name: dataset-curation description: Prepare, format, and validate datasets for supervised fine-tuning and preference training. Use when converting raw data into training format, applying chat templates, configuring sequence packing, generating synthetic training data, or writing a dataset card before a run.
Dataset Curation
This skill assumes finetuning-method-selection
already routed here — the next step is preparing
data, not choosing a method. What follows: format
selection by target method, the template/packing
mechanics behind the most common silent training
failures, rules for mixing in synthetic data
without collapse, and the dataset card that closes
out Phase 2 before a run starts.
Input: raw examples (demonstrations, preference
judgments, or task prompts) plus a routing decision
from finetuning-method-selection.
Output format: a formatted, packed, validated
JSONL dataset plus a completed dataset card — the
Phase 2 artifact /finetune checks before launching
training.
Format Selection
| Method | Shape | Rows |
|---|---|---|
| SFT, single-turn | Instruct (instruction/response or prompt/completion) | ~1,000+ floor |
| SFT, multi-turn | Conversation / ChatML messages list | ~1,000+ floor |
| DPO / ORPO | Preference pair (prompt, chosen, rejected) | Method-dependent, see preference-optimization |
| KTO | Unpaired (prompt, completion, label) | Method-dependent, see preference-optimization |
| GRPO / RLVR | Prompt-only (prompt + verifier metadata) | Method-dependent, see grpo-rlvr-training |
-
~1,000+ rows is the recommended floor for SFT, not a target. Below it, a handful of low-quality or duplicate examples can dominate the gradient; above it, quality over quantity — a smaller verified, deduplicated set beats a larger noisy one.
-
The ChatML shape, for orientation; the other four formats plus a ShareGPT conversion note live in
references/formats-and-templates.md:{"messages": [ {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."} ]}
Chat Templates and Loss Masking
Apply the target model's chat template before any concatenation or packing, never after — packing raw text and templating the packed blob afterward corrupts turn boundaries, landing role markers in the wrong place relative to each example.
-
Train on assistant responses only. Mask the loss (
-100in the labels tensor) over system/user turns and the template's own role markers — only assistant-turn content tokens contribute to loss. -
Template/tokenizer mismatches are a top silent failure mode. A model trained against one chat template but served or evaluated with a different one degrades without erroring. Verify the same template string used in training is applied at inference and eval time.
-
Keep the dataset in
messagesshape and let the trainer template and mask it (assistant_only_loss=Truein current TRL) — pre-rendering to a flat text field destroys the turn boundaries masking needs. Full code sketch:references/formats-and-templates.md. Sanity-check before training — decode only unmasked positions; expect only assistant text:keep = batch["labels"][0] != -100 print(tokenizer.decode(batch["input_ids"][0][keep]))
Packing
Without packing, 40–70% of compute is spent on padding — variable-length examples batched at a fixed sequence length waste the gap between each example's length and the batch's max. Packing concatenates multiple examples into one sequence up to the max length, cutting most of that waste.
-
Packing changes batch semantics. A packed sequence can contain several original examples, so "steps per epoch" and any LR schedule keyed to example count shift once packing is on — recompute schedule milestones against packed-sequence count.
-
MANDATORY: decode and manually inspect 5–10 packed sequences before scaling to a full run. Confirm example boundaries land where expected, template markers are intact per sub-example, and the loss mask is still assistant-only within each packed sequence. Not optional — packing bugs are silent (the loss curve looks normal) and only surface in eval quality, hours later:
for seq in packed_dataset.select(range(10)): print(tokenizer.decode(seq["input_ids"]))
Synthetic Data Rules
- Keep ≥25% real data as a collapse guard.
Training on a growing share of model-generated
data without a real-data floor drives measurable
quality collapse over successive generations —
25% real is the minimum that holds the line.
General-domain replay rows
count toward this floor —
"real" means "not generated
for this task from this
student," not "human-authored."
An all-synthetic-by-construction
dataset can meet the ≥25% floor
through replay alone (see
references/synthetic-data.md's Replay-Mix Construction recipe); state which rows count as "real" in the dataset card rather than leaving the floor structurally unmeetable. - Magpie and rejection sampling are the workhorses. Magpie extracts prompts from the model's own template prior; rejection sampling generates several candidates per prompt and keeps only the ones a filter passes. Both beat naive single-shot generation.
- Targeted, student-aware generation beats static generation by 1.3–2x sample efficiency — aiming at the student's actual failure modes hits a quality bar with fewer filtered examples.
- Typical accept rates after filtering run 10–30%. Plan volume accordingly — a 10,000-row target at 15% accept needs ~65,000+ raw generations.
- Generation-method ranking, filter funnel, replay-
mix construction, and distillation pattern:
references/synthetic-data.md.
The Dataset Card
Every dataset that reaches training gets a card —
the required Phase 2 artifact /finetune checks
before launching. The card is not free-form
documentation; it MUST carry these fields:
- Provenance — where every row came from (real
source(s), synthetic method(s), or both),
traceable to
trace-to-training-dataoutput. - Counts — total rows, and rows per split (train/eval/held-out) if split.
- Synthetic/real ratio — the measured ratio, checked against the ≥25% real floor above.
- Dedup method — exact-match, semantic
(embedding threshold), or both; see the filter
funnel in
references/synthetic-data.md. - Template used — the exact chat template
string/identifier, kept consistent through
inference and eval — this is what ties an
eval-harness-firstrun back to the checkpoint. - Packing config — whether packing was used, max sequence length, and confirmation the 5–10-sequence manual inspection above was done.
A dataset missing any of these six fields isn't
ready for /finetune — the card is a gate, not a
summary written after the fact.
Phase 2 Exit Checklist
Before handing off to /finetune, confirm:
- Format matches the method (table above).
- Template applied before concatenation.
- Loss masked to assistant turns only.
- 5–10 packed sequences decoded and read.
- ≥25% real data in the final mix.
- Dataset card complete — all six fields.
References
references/formats-and-templates.md— JSONL examples per format, current-TRL masking code, and the ShareGPT conversion note.references/synthetic-data.md— generation-method ranking, filter funnel, replay-mix construction, and teacher→student distillation pattern.
Related skills: finetuning-method-selection routes
here; lora-qlora-recipes, vision-sft, and
preference-optimization consume the datasets this
skill produces; trace-to-training-data is the
provenance source for graded-trajectory datasets;
eval-harness-first grades the resulting checkpoint.
Related skills
More from wshobson/agents and the wider catalog.

dbt-transformation-patterns
Master dbt model organization, testing, documentation, and incremental strategies for analytics engineering.

debugging-strategies
Master systematic debugging techniques and root cause analysis to efficiently track down bugs across any codebase.

defi-protocol-templates
Production-ready Solidity templates for staking, AMMs, governance, and flash loans.

dependency-upgrade
Manage major dependency version upgrades with compatibility analysis, staged rollout, and comprehensive testing.

deployment-pipeline-design
Design multi-stage CI/CD pipelines with approval gates, security checks, and progressive delivery strategies.

design-system-patterns
Build scalable design systems with tokens, theming, and component architecture patterns.