PluginBench
Skill
Pass
Audit score 90

resilience-hub-getting-started

aws/agent-toolkit-for-aws

Set up AWS Resilience Hub v2 from scratch: policies, systems, services, and failure mode assessments.

What is resilience-hub-getting-started?

Guides first-time setup of AWS Resilience Hub v2, including creating resilience policies with SLO targets, registering systems and user journeys, onboarding services with input sources, and running failure mode assessments. Use this when starting Resilience Hub v2 for the first time or onboarding a new service with specific availability, RTO, and RPO targets.

  • Create resilience policies with SLO targets (availability, RTO, RPO) and DR approach
  • Register systems, user journeys, and services in Resilience Hub v2
  • Onboard services via input sources (CloudFormation, Terraform, EKS, resource tags)
  • Run failure mode assessments to evaluate architecture against policy targets
  • Troubleshoot common assessment issues (stuck progress, unachievable targets, missing resources)

How to install resilience-hub-getting-started

npx skills add https://github.com/aws/agent-toolkit-for-aws --skill resilience-hub-getting-started
Prerequisites
  • AWS account with Resilience Hub v2 access
  • IAM role with appropriate permissions (read-only discovery, assessment execution)
  • Input sources ready: CloudFormation stack ARNs, Terraform state files, EKS cluster details, or resource tags
  • AWS CLI or AWS MCP server for API calls
Claude Code
Cursor
Windsurf
Cline

How to use resilience-hub-getting-started

  1. 1.Follow the setup procedure in references/setup-procedure.md exactly
  2. 2.Create a resilience policy with your SLO targets (availability %, RTO, RPO) and DR approach
  3. 3.Register your system and user journeys in Resilience Hub v2
  4. 4.Onboard services by specifying input sources (CloudFormation, Terraform, EKS, or tags)
  5. 5.Run the first failure mode assessment and review achievability results
  6. 6.Address any NOT_ACHIEVABLE findings by fixing infrastructure before running FIS experiments

Use cases

Good for
  • Set up Resilience Hub v2 for the first time with a concrete policy and service
  • Onboard a new service to an existing Resilience Hub v2 account with specific SLO targets
  • Run a failure mode assessment to validate architecture meets availability and RTO/RPO requirements
  • Diagnose why an assessment is stuck or shows unachievable targets
Who it's for
  • AWS infrastructure and resilience engineers
  • DevOps teams setting up disaster recovery policies
  • Architects defining SLO targets and resilience strategies
  • Teams onboarding services to Resilience Hub v2

resilience-hub-getting-started FAQ

What if my failure mode assessment is stuck in IN_PROGRESS?

Poll with `aws resiliencehubv2 list-failure-mode-assessments` and check the `errorCode` field. Common causes are INVALID_PERMISSIONS or CMK_ACCESS_DENIED on cross-account roles. If stuck >30 min, verify IAM permissions and KMS key access.

What does NOT_ACHIEVABLE mean?

Your architecture cannot meet the policy targets. Fix infrastructure (add redundancy, reduce RTO/RPO, improve availability) before running FIS experiments — testing won't help if the architecture is fundamentally insufficient.

Why are no resources being discovered?

Verify input sources are correct: CloudFormation stack ARN exists, Terraform state file is accessible, EKS cluster is in the specified regions, or resource tags match actual resources.

Do I need the AWS MCP server to use this skill?

No. The AWS MCP server is recommended for executing API calls, but all operations also work with the AWS CLI directly.

What IAM permissions do I need?

Use least privilege: scope the role to read-only discovery of resource types in your input sources, attach `AWSResilienceHubAsssessmentExecutionPolicy`, and add `aws:SourceAccount` and `aws:SourceArn` conditions to prevent confused-deputy attacks.

Full instructions (SKILL.md)

Source of truth, from aws/agent-toolkit-for-aws.


name: resilience-hub-getting-started description: > Sets up AWS Resilience Hub v2 from scratch: creates resilience policies with SLO targets, registers systems and user journeys, onboards services with input sources, and runs a first failure mode assessment. Applies when the user wants to get started with Resilience Hub v2, create a policy, onboard a service, or run an assessment — including creating one concrete policy with specific availability/RTO/RPO targets and a DR approach for a single service (even a tier-1 one). Does not apply to FIS experiments or ARC routing controls. version: 1

Getting Started with AWS Resilience Hub v2

Overview

Domain expertise for first-time Resilience Hub v2 setup: policies, systems, user journeys, services, input sources, and failure mode assessments.

The AWS MCP server is recommended for executing this skill's AWS API calls, but it is not required — all operations also work with the AWS CLI directly.

Guardrail — where this skill's own files live (MCP vs local install)

Before reading a reference file, determine how this skill was loaded:

  • Loaded via the AWS MCP retrieve_skill tool: the skill's reference files are not on the local filesystem. Fetch each one through retrieve_skill with the file parameter (e.g. file="references/setup-procedure.md") — do NOT file_read these paths locally or search the filesystem for them.
  • Installed locally (e.g. .kiro/skills/resilience-hub-getting-started/ or ~/.claude/skills/resilience-hub-getting-started/): read reference files from the local skill directory using the relative paths shown here.

This applies only to the skill's own reference files; always read and write user or session data in the working directory, never through retrieve_skill.

Set up Resilience Hub v2

To configure Resilience Hub v2 from scratch, follow the procedure exactly. See references/setup-procedure.md.

Troubleshooting

Assessment stuck in IN_PROGRESS

Poll with aws resiliencehubv2 list-failure-mode-assessments. If stuck >30 min, check the errorCode field — common causes are INVALID_PERMISSIONS or CMK_ACCESS_DENIED on cross-account roles.

Achievability shows NOT_ACHIEVABLE

Your architecture cannot meet the policy targets. Fix infrastructure before running FIS experiments — testing won't help if the architecture is fundamentally insufficient.

No resources discovered

Verify input sources are correct: CFN stack ARN exists, Terraform state file is accessible, EKS cluster is in the specified regions, or resource tags match actual resources.

Security Considerations

  • Least privilege: scope the invoker role to read-only discovery of only the resource types in your input sources; attach the AWS managed AWSResilienceHubAsssessmentExecutionPolicy (AWS spells it with three s's) or a tighter custom policy.
  • Encryption at rest / in transit: recommend S3 buckets for Terraform state and assessment reports use server-side encryption (SSE-KMS) and a bucket policy enforcing TLS via aws:SecureTransport.
  • Condition keys (confused-deputy): add an aws:SourceAccount (and ideally aws:SourceArn scoped to the specific Resilience Hub service ARN) condition to the invoker role's trust policy so only your account's Resilience Hub can assume it.
  • Limit assessment exposure: restrict who can call start-failure-mode-assessment (it reads infrastructure state) and who can read assessment findings and reports — these can contain sensitive architecture detail.
  • Further reading: see Security in AWS Resilience Hub and IAM Best Practices (including cross-service confused-deputy prevention).