resilience-hub-failure-mode-assessment
aws/agent-toolkit-for-aws
Run and interpret AWS Resilience Hub v2 failure mode assessments to identify and remediate architectural weaknesses.
What is resilience-hub-failure-mode-assessment?
This skill enables you to execute Resilience Hub v2 failure mode assessments, understand findings by severity and category, triage by achievability, and drive remediation. Use it when you need to run an assessment, review findings, understand failure modes, or resolve specific findings in your AWS architecture.
- Start and run Resilience Hub v2 failure mode assessments
- Interpret findings with severity levels, categories, and recommendations
- Triage findings by achievability and priority
- Manage AI-generated service functions and resource assignments
- Resolve and remediate identified failure modes
- Generate and secure assessment reports
How to install resilience-hub-failure-mode-assessment
npx skills add https://github.com/aws/agent-toolkit-for-aws --skill resilience-hub-failure-mode-assessment- AWS Resilience Hub v2 service enabled in your AWS account
- IAM permissions to describe resources in configured regions (least-privilege invoker role recommended)
- AWS CLI or AWS MCP server for executing API calls
- S3 bucket configured for assessment report output (with encryption and public-access blocking recommended)
How to use resilience-hub-failure-mode-assessment
- 1.Follow the assessment workflow in references/assessment-workflow.md to set up your service and input sources
- 2.Start a failure mode assessment using the Resilience Hub API or CLI
- 3.Retrieve and review findings, sorted by severity (HIGH, MEDIUM, LOW)
- 4.Check achievability for each finding's policy component to determine if architecture changes are needed
- 5.Triage findings using the priority matrix: NOT_ACHIEVABLE findings require architecture changes; ACHIEVABLE findings can be validated with FIS experiments
- 6.Update service functions or resource assignments if AI-generated definitions are incorrect
- 7.Plan and execute remediation, then re-run assessments to validate fixes
Use cases
- Run a failure mode assessment on a multi-region application and prioritize HIGH-severity findings for immediate remediation
- Review assessment findings, filter by achievability, and plan sprint work for MEDIUM-severity issues
- Update AI-generated service functions to correct resource assignments and re-run assessments
- Validate architectural fixes using FIS experiments after addressing NOT_ACHIEVABLE findings
- Generate encrypted assessment reports for compliance and architecture review
- AWS Solutions Architects
- DevOps and SRE engineers
- Application resilience leads
- Cloud infrastructure teams
- AWS Well-Architected reviewers
resilience-hub-failure-mode-assessment FAQ
resilience-hub-getting-started covers initial Resilience Hub setup and configuration. This skill applies when you're running assessments, reviewing findings, or remediating specific failure modes.
The invoker role or cross-account roles lack access to resources in the configured regions. Verify IAM permissions to describe resources in all regions where your service is deployed.
Start with HIGH-severity findings first. For each, check achievability: NOT_ACHIEVABLE means the architecture must change; ACHIEVABLE means validate the fix with an FIS experiment. Plan MEDIUM findings this sprint; track LOW findings but don't block on them.
Yes. Use `aws resiliencehubv2 update-service-function` to rename or change criticality, and `create-service-function-resources` to reassign resources. There is no service-function 'type' parameter.
Store reports in S3 buckets with server-side encryption (SSE-S3 or SSE-KMS) and block public access. If granting Resilience Hub service principal write access, scope the bucket policy with aws:SourceArn and aws:SourceAccount conditions to prevent confused-deputy writes.
Full instructions (SKILL.md)
Source of truth, from aws/agent-toolkit-for-aws.
name: resilience-hub-failure-mode-assessment description: > Runs and interprets AWS Resilience Hub v2 failure mode assessments. Covers starting assessments, understanding findings (severity, categories, recommendations), triaging by achievability, working with AI-generated service functions, and resolving findings. Applies when the user wants to run an assessment, review findings, or understand failure modes, or has a specific finding and asks how to resolve, remediate, or fix it. Does not apply to initial setup (use resilience-hub-getting-started) or FIS experiments. version: 1
Failure Mode Assessment
Overview
Domain expertise for running Resilience Hub v2 failure mode assessments, interpreting findings, triaging by severity and achievability, and driving remediation.
The AWS MCP server is recommended for executing this skill's AWS API calls, but it is not required — all operations also work with the AWS CLI directly.
Guardrail — where this skill's own files live (MCP vs local install)
Before reading a reference file, determine how this skill was loaded:
- Loaded via the AWS MCP
retrieve_skilltool: the skill's reference files are not on the local filesystem. Fetch each one throughretrieve_skillwith thefileparameter (e.g.file="references/assessment-workflow.md") — do NOTfile_readthese paths locally or search the filesystem for them. - Installed locally (e.g.
.kiro/skills/resilience-hub-failure-mode-assessment/or~/.claude/skills/resilience-hub-failure-mode-assessment/): read reference files from the local skill directory using the relative paths shown here.
This applies only to the skill's own reference files; always read and write user or session data in the working directory, never through retrieve_skill.
Run and interpret assessments
To run assessments and triage findings, follow the procedure exactly. See references/assessment-workflow.md.
Troubleshooting
Assessment fails with INVALID_PERMISSIONS
The service's permission model (invokerRoleName / crossAccountRoles) doesn't have access to the resources. Verify the invoker role (and any cross-account roles) can describe resources in all configured regions.
Too many findings — where to start?
Prioritize by finding severity, highest first (HIGH, then MEDIUM, then LOW). For HIGH-severity
findings, check the service's achievability for the relevant policy component (from get-service
/ list-failure-mode-assessments): NOT_ACHIEVABLE means the architecture must change before
testing; ACHIEVABLE means validate the fix with an FIS experiment. MEDIUM findings: plan
remediation this sprint; LOW findings: track but don't block (see the priority matrix in
references/assessment-workflow.md Step 5).
AI-generated service functions are wrong
Update them: aws resiliencehubv2 update-service-function to rename or change criticality
(there is no service-function "type" parameter). Reassign resources by calling create-service-function-resources with the desired resource set (see references/assessment-workflow.md for the service-function operations).
Security Considerations
- Least privilege: the invoker role should be scoped to read-only discovery of only the resource types in the service's input sources; avoid granting access beyond what assessment needs.
- Encryption & access control: recommend that S3 buckets used for report output have server-side encryption (SSE-S3 or SSE-KMS) and block public access — assessment reports can contain sensitive architectural detail. If a bucket policy grants the Resilience Hub service principal write access, scope it with
aws:SourceArn/aws:SourceAccountcondition keys to prevent confused-deputy writes. - Further reading: see Security in AWS Resilience Hub and the AWS Well-Architected Security Pillar for securing assessment outputs and IAM configurations.
Related skills
More from aws/agent-toolkit-for-aws and the wider catalog.

resilience-hub-getting-started
Set up AWS Resilience Hub v2 from scratch: policies, systems, services, and failure mode assessments.

resilience-hub-multi-account
Set up AWS Resilience Hub v2 for multi-account resilience assessments across AWS Organizations.

resilience-program-design
Design org-wide resilience policies with tiered targets and activity cadence.

route53
Configure Amazon Route 53 DNS records, routing policies, health checks, DNS Firewall, and hybrid network resolution.

routing-traffic-with-route53-and-cloudfront
Configure Route 53 DNS routing to CloudFront distributions with custom domains and HTTPS.

running-release-tests
Run automated UI and API release tests via AWS DevOps Agent using pre-configured test profiles.