finetuning
microsoft/azure-skills
Fine-tune models on Azure AI Foundry with SFT, DPO, or RFT training methods.
What is finetuning?
Fine-tune language and vision models on Azure AI Foundry using supervised (SFT), preference (DPO), or reinforcement (RFT) training. Use this skill when preparing training data, submitting and monitoring training jobs, calibrating graders, deploying fine-tuned models, and evaluating results.
- Submit and monitor SFT, DPO, and RFT training jobs on Azure AI Foundry
- Prepare, validate, and convert training datasets between formats
- Calibrate graders and pass thresholds for reinforcement fine-tuning
- Deploy fine-tuned models via ARM REST API
- Evaluate fine-tuned models using LLM judges
- Generate synthetic training data and score dataset quality
How to install finetuning
npx skills add https://github.com/microsoft/azure-skills --skill microsoft-foundry- Azure AI Foundry project and API credentials
- Training data in JSONL format (SFT, DPO, or RFT)
- Python 3.8+ with openai SDK >= 1.0
- For RFT: a grader function or tool endpoint
How to use finetuning
- 1.Prepare training data in the appropriate JSONL format (SFT, DPO, or RFT)
- 2.Run validation script to check data quality and format
- 3.Submit training job using submit_training.py with model, data files, and training type
- 4.Monitor job progress with monitor_training.py until completion
- 5.Analyze training curves and checkpoints with check_training.py
- 6.Deploy the selected fine-tuned model using deploy_model.py
- 7.Evaluate the deployed model on test data with evaluate_model.py
Use cases
- Fine-tune a base model on domain-specific supervised examples to improve task accuracy
- Use DPO to align model outputs with preference rankings from human feedback
- Train with RFT using custom graders to optimize for specific evaluation criteria
- Prepare and validate JSONL training datasets before submission
- Monitor training curves and select optimal checkpoints for deployment
- ML engineers building custom models for specific domains
- Data scientists optimizing model behavior with preference or reinforcement training
- Teams deploying fine-tuned models to production on Azure
- Researchers experimenting with different training types and hyperparameters
finetuning FAQ
Use SFT for supervised learning on labeled examples. Use DPO when you have preference pairs (better/worse responses). Use RFT when you have a grader function that can score outputs and want to optimize for specific criteria.
Format data as JSONL with required fields per training type. For SFT: messages array. For DPO: chosen and rejected completions. For RFT: prompt and expected behavior. Run validate_sft.py, validate_dpo.py, or validate_rft.py before submission.
Target 25-50% failure rate on the base model. Use calibrate_grader.py to find the optimal pass_threshold that achieves this range.
Always baseline the original model first. Compare metrics on a held-out test set using evaluate_model.py. Also measure token cost alongside accuracy for production decisions.
Check error messages in the job logs. Common issues: API version mismatch (upgrade openai SDK), under-provisioned tool endpoint for RFT (scale to S2+), or content safety blocks (review training data for PII). See platform-gotchas.md for detailed troubleshooting.
Full instructions (SKILL.md)
Source of truth, from microsoft/azure-skills.
name: finetuning description: "Fine-tune models on Azure AI Foundry using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset preparation, training job submission, deployment, and evaluation. USE FOR: fine-tune, SFT, DPO, RFT, training data, grader, distillation, fine-tuned model, training job, large file upload, calibrate grader, deploy fine-tuned model, evaluate fine-tuned model. DO NOT USE FOR: general model deployment without fine-tuning (use deploy-model), agent creation (use agents), prompt optimization without training (use prompt-optimizer)." license: MIT metadata: author: Microsoft version: "0.0.0-placeholder"
Fine-Tuning on Azure AI Foundry
Fine-tune models using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset prep, training, deployment, and evaluation.
When to Use
Use this sub-skill when the user asks about:
- Fine-tuning a model (SFT, DPO, or RFT)
- Preparing, validating, or formatting training data
- Submitting, monitoring, or diagnosing training jobs
- Calibrating graders or pass thresholds for RFT
- Deploying or evaluating a fine-tuned model
- Choosing between training types (SFT vs DPO vs RFT)
- Distillation, synthetic data generation, or dataset quality scoring
- Large file uploads for training data
- Cleaning up fine-tuning resources (files, deployments)
Do NOT use for: General model deployment without fine-tuning (use deploy-model), agent creation (use agents), prompt optimization without training (use prompt-optimizer).
Workflows
| Stage | Guide |
|---|---|
| Quick start | workflows/quickstart.md |
| Full pipeline | workflows/full-pipeline.md |
| Create data | workflows/dataset-creation.md |
| Iterate | workflows/iterative-training.md |
| Diagnose | workflows/diagnose-poor-results.md |
References
| Topic | File |
|---|---|
| SFT vs DPO vs RFT | references/training-types.md |
| Hyperparameters | references/hyperparameters.md |
| Data formats | references/dataset-formats.md |
| Grader design (RFT) | references/grader-design.md |
| Reward hacking | references/reward-hacking.md |
| Agentic RFT (tools) | references/agentic-rft.md |
| Deployment | references/deployment.md |
| Training curves | references/training-curves.md |
| Evaluation | references/evaluation.md |
| Vision fine-tuning | references/vision-fine-tuning.md |
| Large file uploads | references/large-file-uploads.md |
| Platform gotchas | references/platform-gotchas.md |
Scripts
| Script | Purpose |
|---|---|
scripts/submit_training.py | Submit SFT/DPO/RFT jobs |
scripts/monitor_training.py | Poll job until completion |
scripts/calibrate_grader.py | Find optimal RFT pass_threshold |
scripts/check_training.py | Analyze curves, list checkpoints |
scripts/deploy_model.py | Deploy via ARM REST API |
scripts/evaluate_model.py | LLM judge evaluation |
scripts/convert_dataset.py | Convert between SFT/DPO/RFT formats |
scripts/generate_distillation_data.py | Generate synthetic training data |
scripts/score_dataset.py | Quality scoring on training data |
scripts/cleanup.py | Delete old files and deployments |
scripts/validate/ | Data validators (SFT, DPO, RFT) + stats |
Rules
- Always baseline first — evaluate the base model before fine-tuning
- Validate data before submitting — run
scripts/validate/validate_sft.py - Calibrate RFT graders — target 25-50% failure rate on the base model
- Evaluate checkpoints — don't blindly deploy the final one
- Measure token cost alongside accuracy when comparing models
Quick Reference
| Task | Command |
|---|---|
| Validate SFT data | python scripts/validate/validate_sft.py data.jsonl |
| Submit SFT job | python scripts/submit_training.py --model gpt-4.1-mini --training-file train.jsonl --validation-file val.jsonl --type sft |
| Monitor job | python scripts/monitor_training.py --job-id ftjob-xxx |
| Analyze curves | python scripts/check_training.py --job-id ftjob-xxx |
| Deploy model | python scripts/deploy_model.py --model-id ft:gpt-4.1-mini:... --name my-eval |
| Evaluate model | python scripts/evaluate_model.py --deployment-name my-eval --test-file test.jsonl |
Error Handling
| Error | Cause | Fix |
|---|---|---|
| "API version not supported" | Older openai SDK on /v1/ endpoint | Upgrade to openai>=1.0 |
| "does not support fine-tuning with Standard TrainingType" | OSS model needs globalStandard | Use --use-rest flag or script auto-falls back |
| Job stuck in post-training eval | Under-provisioned tool endpoint (RFT) | Scale to S2+, enable Always On |
| "DeploymentNotReady" after ARM succeeds | ARM/data-plane race condition | Delete and recreate deployment, wait 5 min |
| Content safety block at deployment | PII-dense training data | Remove problematic document types |
Related skills
More from microsoft/azure-skills and the wider catalog.
azure-ai
Azure AI services skill for Search, Speech, OpenAI, and Document Intelligence in coding agents
azure-deploy
Execute Azure deployments for prepared applications with built-in error recovery and validation.
azure-diagnostics
Debug Azure production issues using AppLens, Azure Monitor, resource health, and systematic triage.
azure-prepare
Generate Azure deployment infrastructure (Bicep/Terraform, azure.yaml, Dockerfiles) for new or existing apps
azure-storage
Azure Storage skill: Blob, File Shares, Queue, Table, and Data Lake with access tier guidance and lifecycle management
azure-validate
Pre-deployment validation for Azure readiness with configuration, infrastructure, and RBAC checks.