agents-optimize
aws/agent-toolkit-for-aws
Measure and improve agent quality through evaluation, monitoring, and observability.
What is agents-optimize?
Set up evaluators, quality gates, CI/CD monitoring, and observability for your AgentCore agent. Use this when you want to measure agent performance, understand behavior through logs and traces, or optimize costs—not for debugging broken agents.
- Configure evaluators (LLM-as-a-judge, code-based) and run quality assessments
- Set up online monitoring and CI/CD quality gates with pass/fail thresholds
- Create CloudWatch dashboards and X-Ray tracing for observability
- Monitor agent logs, metrics, and traces with ~10-second end-to-end latency
- Analyze and optimize agent costs and latency
How to install agents-optimize
npx skills add https://github.com/aws/agent-toolkit-for-aws --skill agents-optimize- AgentCore CLI v0.9.0 or later
- An existing AgentCore project (agentcore/agentcore.json)
How to use agents-optimize
- 1.Run `agentcore --version` to verify CLI is v0.9.0 or later
- 2.Read `agentcore/agentcore.json` to understand your current agent setup
- 3.Determine your workflow: quality measurement, observability setup, cost analysis, or a combination
- 4.Load the relevant reference file (evals.md, observability.md, or cost.md) and follow its step-by-step procedure
- 5.Configure evaluators, monitoring, or dashboards as specified in the reference
Use cases
- Add a quality gate to your CI/CD pipeline to block deployments below a quality threshold
- Set up continuous monitoring in production to track agent answer quality over time
- Create CloudWatch dashboards to visualize agent performance metrics and traces
- Run evaluators locally or online to measure if your agent gives good answers
- Understand agent behavior and latency by analyzing logs and X-Ray traces
- AgentCore developers measuring agent quality
- DevOps engineers setting up CI/CD quality gates
- Platform engineers building observability for production agents
- Teams optimizing agent performance and costs
agents-optimize FAQ
Use agents-optimize to measure quality and set up monitoring for slow-but-correct agents. Use agents-debug when your agent is broken, giving wrong answers, or crashing.
End-to-end latency is approximately 10 seconds from trace ingestion to query results—there is no separate indexing wait.
The skill supports it, but you should discuss the cost and performance impact before enabling 100% sampling in production.
No—you can do either first. If you want to understand and improve your agent holistically, start with observability, then add evals.
Use agents-debug to investigate the root cause of the agent's incorrect behavior.
Full instructions (SKILL.md)
Source of truth, from aws/agent-toolkit-for-aws.
name: agents-optimize description: > Use when measuring or improving agent quality and performance — set up evaluators, online monitoring, CI/CD quality gates, observability, or cost optimization. Triggers on: "evaluate my agent", "add evaluator", "measure quality", "quality gate", "run evals", "agent too slow", "why is it slow", "reduce latency", "set up observability", "CloudWatch dashboard", "how much does my agent cost", "cost optimization", "logs not showing up", "logs missing", "spans not found", "eval failing", "eval error", "dev traces", "local traces", "agentcore dev traces", "traces to CloudWatch". Not for debugging errors or crashes — use agents-debug. Slow but correct routes here; broken routes to debug. allowed-tools: Read Grep Glob Bash metadata: type: skill version: "1.0.0" author: aws-agentcore requires-cli: ">=0.9.0"
optimize
Measure and improve your AgentCore agent's quality through evaluation, monitoring, and observability.
When to use
- You want to know if your agent is giving good answers
- You want to set up continuous quality monitoring in production
- You want to add a quality gate to your CI/CD pipeline
- You want to understand agent behavior through logs, metrics, and traces
- You want to set up CloudWatch dashboards or X-Ray tracing
Do NOT use for:
- Debugging a specific broken agent (wrong answers, errors) → use
agents-debug - Production security hardening (IAM, auth) → use
agents-harden
Input
$ARGUMENTS can be:
- An eval goal: "add a quality gate", "set up monitoring"
- An observability goal: "set up CloudWatch dashboard", "understand my traces"
- A specific evaluator: "llm-as-a-judge", "code-based"
- Empty — the skill will guide based on project context
Process
Step 0: Verify CLI version
Run agentcore --version. This skill requires v0.9.0 or later.
Step 1: Read project context
Read agentcore/agentcore.json to understand existing evaluators, online eval configs, and agent setup.
If agentcore/agentcore.json is not found:
"This skill requires an AgentCore project. Use
agents-get-startedto create one."
Step 2: Determine the workflow
| Developer intent | Action |
|---|---|
| Measure quality, add evaluator, run eval, CI/CD gate, online monitoring | Load references/evals.md and follow its workflow |
| Set up observability, CloudWatch, X-Ray, logs, metrics, dashboards | Load references/observability.md and follow its workflow |
| Understand or reduce AgentCore costs | Load references/cost.md |
| Both — "I want to understand and improve my agent" | Start with observability setup, then add evals |
Step 3: Follow the loaded reference
The reference file contains the full procedure. Follow it step by step.
Cross-references
- After setting up evals, suggest
agents-hardenfor production readiness - If eval results reveal agent issues, suggest
agents-debugfor root cause analysis - If the developer needs to add capabilities first, suggest
agents-build
Output
Depends on the workflow — see the loaded reference for specific outputs.
Quality criteria
- Evaluator configuration uses only valid CLI flags
- Online eval sampling rate is appropriate (not 100% in production without discussion)
- CI/CD quality gate has a clear pass/fail threshold
- Observability setup includes both tracing and logging
- The developer understands the eval data delay: ~10 seconds put-to-get, end-to-end — one ingestion step covers both trace reads and eval queries; there is no separate indexing wait
Related skills
More from aws/agent-toolkit-for-aws and the wider catalog.

amazon-aurora-mysql
Create, modify, and advise on Amazon Aurora MySQL clusters with safety guardrails and cost optimization.

amazon-aurora-postgresql
Create, modify, and advise on Amazon Aurora PostgreSQL clusters with express configuration, serverless sizing, and upgrade planning.

amazon-bedrock
Build generative AI applications on Amazon Bedrock with model invocation, RAG, agents, and guardrails.

amazon-braket
Run quantum computing workflows on AWS Braket—submit circuits, manage devices, and control costs.

amazon-documentdb
End-to-end Amazon DocumentDB management: cluster setup, schema design, MongoDB migration, performance tuning, and Well-Architected reviews.

amazon-dynamodb
Design, review, and debug DynamoDB data layers with access-pattern axioms, cost estimation, and optional live validation.