resilience-hub-getting-started
aws/agent-toolkit-for-aws
Set up AWS Resilience Hub v2 from scratch: policies, systems, services, and failure mode assessments.
What is resilience-hub-getting-started?
Guides first-time setup of AWS Resilience Hub v2, including creating resilience policies with SLO targets, registering systems and user journeys, onboarding services with input sources, and running failure mode assessments. Use this when starting Resilience Hub v2 for the first time or onboarding a new service with specific availability, RTO, and RPO targets.
- Create resilience policies with SLO targets (availability, RTO, RPO) and DR approach
- Register systems, user journeys, and services in Resilience Hub v2
- Onboard services via input sources (CloudFormation, Terraform, EKS, resource tags)
- Run failure mode assessments to evaluate architecture against policy targets
- Troubleshoot common assessment issues (stuck progress, unachievable targets, missing resources)
How to install resilience-hub-getting-started
npx skills add https://github.com/aws/agent-toolkit-for-aws --skill resilience-hub-getting-started- AWS account with Resilience Hub v2 access
- IAM role with appropriate permissions (read-only discovery, assessment execution)
- Input sources ready: CloudFormation stack ARNs, Terraform state files, EKS cluster details, or resource tags
- AWS CLI or AWS MCP server for API calls
How to use resilience-hub-getting-started
- 1.Follow the setup procedure in references/setup-procedure.md exactly
- 2.Create a resilience policy with your SLO targets (availability %, RTO, RPO) and DR approach
- 3.Register your system and user journeys in Resilience Hub v2
- 4.Onboard services by specifying input sources (CloudFormation, Terraform, EKS, or tags)
- 5.Run the first failure mode assessment and review achievability results
- 6.Address any NOT_ACHIEVABLE findings by fixing infrastructure before running FIS experiments
Use cases
- Set up Resilience Hub v2 for the first time with a concrete policy and service
- Onboard a new service to an existing Resilience Hub v2 account with specific SLO targets
- Run a failure mode assessment to validate architecture meets availability and RTO/RPO requirements
- Diagnose why an assessment is stuck or shows unachievable targets
- AWS infrastructure and resilience engineers
- DevOps teams setting up disaster recovery policies
- Architects defining SLO targets and resilience strategies
- Teams onboarding services to Resilience Hub v2
resilience-hub-getting-started FAQ
Poll with `aws resiliencehubv2 list-failure-mode-assessments` and check the `errorCode` field. Common causes are INVALID_PERMISSIONS or CMK_ACCESS_DENIED on cross-account roles. If stuck >30 min, verify IAM permissions and KMS key access.
Your architecture cannot meet the policy targets. Fix infrastructure (add redundancy, reduce RTO/RPO, improve availability) before running FIS experiments — testing won't help if the architecture is fundamentally insufficient.
Verify input sources are correct: CloudFormation stack ARN exists, Terraform state file is accessible, EKS cluster is in the specified regions, or resource tags match actual resources.
No. The AWS MCP server is recommended for executing API calls, but all operations also work with the AWS CLI directly.
Use least privilege: scope the role to read-only discovery of resource types in your input sources, attach `AWSResilienceHubAsssessmentExecutionPolicy`, and add `aws:SourceAccount` and `aws:SourceArn` conditions to prevent confused-deputy attacks.
Full instructions (SKILL.md)
Source of truth, from aws/agent-toolkit-for-aws.
name: resilience-hub-getting-started description: > Sets up AWS Resilience Hub v2 from scratch: creates resilience policies with SLO targets, registers systems and user journeys, onboards services with input sources, and runs a first failure mode assessment. Applies when the user wants to get started with Resilience Hub v2, create a policy, onboard a service, or run an assessment — including creating one concrete policy with specific availability/RTO/RPO targets and a DR approach for a single service (even a tier-1 one). Does not apply to FIS experiments or ARC routing controls. version: 1
Getting Started with AWS Resilience Hub v2
Overview
Domain expertise for first-time Resilience Hub v2 setup: policies, systems, user journeys, services, input sources, and failure mode assessments.
The AWS MCP server is recommended for executing this skill's AWS API calls, but it is not required — all operations also work with the AWS CLI directly.
Guardrail — where this skill's own files live (MCP vs local install)
Before reading a reference file, determine how this skill was loaded:
- Loaded via the AWS MCP
retrieve_skilltool: the skill's reference files are not on the local filesystem. Fetch each one throughretrieve_skillwith thefileparameter (e.g.file="references/setup-procedure.md") — do NOTfile_readthese paths locally or search the filesystem for them. - Installed locally (e.g.
.kiro/skills/resilience-hub-getting-started/or~/.claude/skills/resilience-hub-getting-started/): read reference files from the local skill directory using the relative paths shown here.
This applies only to the skill's own reference files; always read and write user or session data in the working directory, never through retrieve_skill.
Set up Resilience Hub v2
To configure Resilience Hub v2 from scratch, follow the procedure exactly. See references/setup-procedure.md.
Troubleshooting
Assessment stuck in IN_PROGRESS
Poll with aws resiliencehubv2 list-failure-mode-assessments. If stuck >30 min, check
the errorCode field — common causes are INVALID_PERMISSIONS or CMK_ACCESS_DENIED on
cross-account roles.
Achievability shows NOT_ACHIEVABLE
Your architecture cannot meet the policy targets. Fix infrastructure before running FIS experiments — testing won't help if the architecture is fundamentally insufficient.
No resources discovered
Verify input sources are correct: CFN stack ARN exists, Terraform state file is accessible, EKS cluster is in the specified regions, or resource tags match actual resources.
Security Considerations
- Least privilege: scope the invoker role to read-only discovery of only the resource types in your input sources; attach the AWS managed
AWSResilienceHubAsssessmentExecutionPolicy(AWS spells it with three s's) or a tighter custom policy. - Encryption at rest / in transit: recommend S3 buckets for Terraform state and assessment reports use server-side encryption (SSE-KMS) and a bucket policy enforcing TLS via
aws:SecureTransport. - Condition keys (confused-deputy): add an
aws:SourceAccount(and ideallyaws:SourceArnscoped to the specific Resilience Hub service ARN) condition to the invoker role's trust policy so only your account's Resilience Hub can assume it. - Limit assessment exposure: restrict who can call
start-failure-mode-assessment(it reads infrastructure state) and who can read assessment findings and reports — these can contain sensitive architecture detail. - Further reading: see Security in AWS Resilience Hub and IAM Best Practices (including cross-service confused-deputy prevention).
Related skills
More from aws/agent-toolkit-for-aws and the wider catalog.

resilience-hub-multi-account
Set up AWS Resilience Hub v2 for multi-account resilience assessments across AWS Organizations.

resilience-program-design
Design org-wide resilience policies with tiered targets and activity cadence.

route53
Configure Amazon Route 53 DNS records, routing policies, health checks, DNS Firewall, and hybrid network resolution.

routing-traffic-with-route53-and-cloudfront
Configure Route 53 DNS routing to CloudFront distributions with custom domains and HTTPS.

running-release-tests
Run automated UI and API release tests via AWS DevOps Agent using pre-configured test profiles.

scanning-with-aws-security-agent
Run AWS Security Agent scans on your codebase to find vulnerabilities with ranked findings and remediation guidance.