aws-ai-ml
aws/agent-toolkit-for-aws
Fine-tune and deploy AI models on Amazon SageMaker with guided workflows for the full ML lifecycle.
What is aws-ai-ml?
This skill provides domain expertise for selecting, fine-tuning, and deploying models on Amazon SageMaker. It covers the complete model customization lifecycle—from planning and dataset preparation through training, evaluation, and production deployment—including support for multiple fine-tuning techniques (SFT, DPO, RLVR, RLAIF), endpoint diagnostics, and inference optimization.
- Select and evaluate base models from SageMaker Hub for fine-tuning or deployment
- Prepare, validate, and transform datasets for training jobs
- Fine-tune models using SFT, DPO, RLVR, or RLAIF techniques
- Evaluate and benchmark trained models using LLM-as-Judge and custom scorers
- Deploy models to SageMaker Real-Time Endpoints or Amazon Bedrock
- Optimize inference performance and diagnose endpoint health issues
How to install aws-ai-ml
npx skills add https://github.com/aws/agent-toolkit-for-aws --skill aws-ai-ml- AWS account with SageMaker access
- AWS CLI configured with appropriate credentials
- Python 3.8+ and boto3 SDK
- S3 bucket for storing training data and model artifacts
How to use aws-ai-ml
- 1.Run `npx skills add https://github.com/aws/agent-toolkit-for-aws --skill aws-ai-ml` to install the skill
- 2.Set up your AWS environment by routing to sdk-getting-started to configure IAM roles and S3 buckets
- 3.Define your use case and success criteria using the use-case-specification reference
- 4.Select a base model from SageMaker Hub using model-selection, filtered by your requirements
- 5.Validate and transform your dataset using dataset-evaluation and dataset-transformation
- 6.Choose a fine-tuning technique (SFT, DPO, RLVR, RLAIF) matched to your model and use case
- 7.Generate and run fine-tuning code via the finetuning reference to start training
- 8.Evaluate your trained model using model-evaluation with benchmarking and scoring
Use cases
- Fine-tune a foundation model on your proprietary data and deploy it to a SageMaker endpoint for production inference
- Evaluate multiple base models against your use-case requirements and select the best fit before training
- Validate dataset quality and format before starting a fine-tuning job to catch issues early
- Benchmark endpoint performance across instance types to find the most cost-effective deployment
- Diagnose why a deployed endpoint is experiencing high latency or inference errors
- ML engineers building custom models on AWS
- Data scientists preparing datasets and evaluating model quality
- DevOps engineers deploying and monitoring SageMaker endpoints
- AI teams planning end-to-end model customization projects
aws-ai-ml FAQ
SFT (Supervised Fine-Tuning), DPO (Direct Preference Optimization), RLVR (Reinforcement Learning from Verification Reward), and RLAIF (Reinforcement Learning from AI Feedback). The skill helps you choose the right technique based on your model and use case.
Yes. You can select a base model from SageMaker Hub and deploy it directly to a SageMaker endpoint or Bedrock without any training, using the model-selection and model-deployment references.
No. This skill does not cover Ground Truth labeling or Feature Store. It focuses on dataset validation, transformation, and preparation for training, but not on creating labeled datasets from scratch.
Use the endpoint-diagnostics reference to check endpoint status, view container logs, examine CloudWatch metrics, and diagnose the root cause of errors or performance issues.
Set up a SageMaker Managed MLflow app using the manage-mlflow reference to log parameters, metrics, and artifacts across fine-tuning and evaluation runs.
Full instructions (SKILL.md)
Source of truth, from aws/agent-toolkit-for-aws.
name: aws-ai-ml description: > Selects, deploys, and customizes AI models on Amazon SageMaker. Fine-tuning (SFT, DPO, RLVR, RLAIF), model selection, dataset preparation, evaluation, deployment to SageMaker endpoints or Bedrock, inference optimization and endpoint diagnostics. Covers the full lifecycle from planning through production. Use when fine-tuning models on SageMaker, choosing/selecting which base model to customize or fine-tune from SageMaker Hub, finding a model to deploy without fine-tuning, transforming datasets for training, checking data readiness, evaluating model quality, deploying to endpoints, benchmarking or optimizing inference, setting up IAM roles and S3 buckets for training jobs, or managing a SageMaker Managed MLflow app. Also use to check endpoint health, diagnose failures, debug latency or errors, or view container logs and CloudWatch metrics. Covers Serverless Model Customization, Nova and OSS deployment paths, and PySDK v3. NOT for Ground Truth labeling, Feature Store, or general-purpose AWS infrastructure. metadata: version: "4"
AWS AI/ML Model Customization
Domain expertise for fine-tuning and deploying models on Amazon SageMaker. Covers the full model customization lifecycle from planning through production deployment.
Routing
Match the user's intent to the appropriate reference folder and load only that content.
| User intent | Reference | When to use |
|---|---|---|
| Plan a model customization project, discover scope of work, resume or modify a plan | references/planning/ | User's request relates to model customization or deployment (fine-tuning, training, building, customizing, reviewing data, deploying or standing up a model — including selecting or deploying an off-the-shelf or base model with no training — or getting advice on approach). Always co-activate with other intents to discover full scope. Load this reference FIRST when the request matches multiple rows in this table — read its plan templates before routing to a single-action reference. |
| Define the business problem, success criteria, or use case spec | references/use-case-specification/ | User says "define my use case", "capture requirements", "what should I decide up front", or as default first step in any plan. Skip only if user explicitly declines. |
| Select or change a base model | references/model-selection/ | User asks which model to use, mentions a model name or family, or wants to evaluate what's available. Always activate model-selection even for known model names because the exact Hub model ID must be resolved. Recommended: route to use-case-specification first to capture requirements — this produces better filtering results. Routing to use-case-specification first is not required if user provides a specific model name/ID or declines. If intent is ambiguous (fine-tune vs deploy as-is), model-selection MUST confirm which path before proceeding. Base model filtering for deployment MUST go through select-for-deployment.md and its scripts for any final recommendation. |
| Choose a fine-tuning technique (SFT, DPO, RLVR, RLAIF) | references/finetuning-technique/ | User has decided to fine-tune and needs to choose a technique, or technique needs validation against the selected model's recipes. Requires a base model to be selected first. |
| Validate dataset quality and format | references/dataset-evaluation/ | User says "is my dataset okay", "check my training data", "I have my own data", or before starting any fine-tuning job. |
| Transform or convert a dataset between formats | references/dataset-transformation/ | User says "transform", "convert", "reformat", or dataset schema needs to change. Always use this rather than writing inline transformation code. |
| Generate fine-tuning code and start training | references/finetuning/ | User says "start training", "fine-tune my model", "I'm ready to train", or plan reaches the finetuning step. Supports SFT, DPO, RLVR, RLAIF trainers. |
| Evaluate or benchmark a trained model | references/model-evaluation/ | User says "evaluate my model", "run a benchmark", "test model performance", "compare models". Supports LLM-as-Judge and Custom Scorer. |
| Deploy, benchmark, or optimize a model on an endpoint or Bedrock | references/model-deployment/ | User says "deploy my model", "create an endpoint", or "make it available" (plain deploy) — or, for the inference-optimization sub-workflows on SageMaker Real-Time Endpoints only, "benchmark my endpoint" / "compare benchmark runs" (benchmarking), or states a performance/cost/latency/throughput goal for a new deployment such as "find the cheapest instance" (recommendations). Handles Nova vs OSS deployment pathways. |
| Set up IAM roles, S3 buckets, SDK configuration | references/sdk-getting-started/ | User says "set up", "getting started", "check my environment", "configure SDK", or as first step in any plan involving SageMaker training/evaluation/deployment. |
| Manage project directory and artifacts | references/directory-management/ | Starting a new project, resuming existing one, or when PLAN.md needs to be associated with a project directory. |
| Set up, update, or delete a SageMaker Managed MLflow app | references/manage-mlflow/ | User says "set up MLflow", "create MLflow app", "update my MLflow app", "delete my MLflow app", "I need an MLflow server", asks "what is SageMaker MLflow", or a workflow needs an MLflow backend and none is connected. |
| Diagnose a failing or unhealthy SageMaker endpoint | references/endpoint-diagnostics/ | User reports endpoint errors, latency, inference failures, or a deployment that failed. "What's the status of my endpoint?", "Is my endpoint erroring?", "My endpoint failed — why?", "How many instances are running behind my endpoint?", "Is the latency my model or SageMaker?", "Show me the container logs for my endpoint." NOT for training-job issues, endpoint deletion, scaling changes, or new deployments. |
Rules
- Progressive disclosure. Load only the reference folder relevant to the current user intent. Do not load all references at once.
- Best-effort help. If the user's request falls outside this skill's references, do not dead-end the conversation. Help them using general AWS knowledge and documentation, and inform the user that the guidance is not covered by this skill's validated workflows.
- Usage attribution. Before running any AWS CLI command or packaged script, set
export AWS_SDK_UA_APP_ID=AWSSkill-SageMaker.
Related skills
More from aws/agent-toolkit-for-aws and the wider catalog.

aws-amplify
Build and deploy full-stack web and mobile apps with AWS Amplify Gen2 using TypeScript code-first approach.

aws-auth
Amazon Cognito user authentication for web and mobile apps with sign-up, sign-in, MFA, OAuth, and token management.

aws-billing-and-cost-management
Analyze AWS costs, optimize spending, and manage budgets with expert domain knowledge.

aws-blocks
Infrastructure-from-Code framework for building full-stack AWS applications with pre-built Building Blocks.

aws-cdk
Author, deploy, and troubleshoot AWS infrastructure using CDK with TypeScript or Python.

aws-cleanrooms
Troubleshoot AWS Clean Rooms permission and logging issues for collaborations and ML jobs.