aws-resilience-lifecycle
aws/agent-toolkit-for-aws
Guide end-to-end AWS resilience: Define policies, Test with fault injection, Operate with controls.
What is aws-resilience-lifecycle?
Orchestrates the complete AWS resilience lifecycle across Resilience Hub v2, Fault Injection Service, and Application Recovery Controller. Use this when building a comprehensive resilience program, validating findings with experiments before marking them resolved, or connecting policy definitions to failure-mode testing to operational controls.
- Define resilience policies and register services in Resilience Hub v2 (NGRH)
- Run failure-mode assessments to identify gaps and generate findings
- Design and execute FIS experiments to validate remediation before resolving findings
- Operationalize resilience controls via Application Recovery Controller
- Connect findings to experiments to operational controls in a single workflow
- Validate that architectures recover within RTO/RPO targets under real failure conditions
How to install aws-resilience-lifecycle
npx skills add https://github.com/aws/agent-toolkit-for-aws --skill aws-resilience-lifecycle- AWS account with Resilience Hub v2 (NGRH), Fault Injection Service, and Application Recovery Controller enabled
- IAM roles for Resilience Hub invoker, FIS execution, and ARC operator with least-privilege permissions
- CloudWatch alarms configured as FIS stop conditions (bounded blast radius)
- Service registered in Resilience Hub v2 with baseline assessment completed
How to use aws-resilience-lifecycle
- 1.Create a resilience policy in Resilience Hub v2 defining RTO and RPO targets
- 2.Register your application or service and run an assessment to generate findings
- 3.For each finding, design a FIS experiment that reproduces the failure mode
- 4.Execute the experiment and confirm recovery within policy targets
- 5.Mark the finding resolved only after successful experiment validation
- 6.Configure Application Recovery Controller operational controls for the validated resilience posture
- 7.Monitor production using CloudWatch metrics and post-experiment analysis to detect cascading failures
Use cases
- Building a complete resilience strategy from policy through testing to production controls
- Validating that a remediation actually works before marking a finding resolved in Resilience Hub
- Running fault-injection experiments to prove recovery time objectives are met
- Designing multi-fault scenarios and blast-radius controls for production experiments
- Transitioning from paper compliance (resolved findings) to proven resilience (validated with FIS)
- Resilience engineers and architects designing end-to-end resilience programs
- DevOps and SRE teams validating application recovery under failure conditions
- AWS Well-Architected reviewers assessing resilience maturity
- Teams migrating from Resilience Hub v1 to v2 with integrated testing
- Organizations requiring change-management approval before production fault injection
aws-resilience-lifecycle FAQ
No. Marking resolved without FIS validation is paper compliance. You must run an experiment that reproduces the failure mode, confirm recovery within RTO/RPO, then mark resolved. Validating after marking resolved is the anti-pattern.
Experiments may not match real failure modes. Expand blast radius, add multi-fault scenarios, ensure stop conditions match production SLOs (not relaxed test thresholds), and verify observability signals are correct.
No. The AWS MCP server is recommended but not required — all operations work with the AWS CLI directly. If loaded via MCP retrieve_skill, fetch reference files through the tool; if installed locally, read them from the skill directory.
Use the companion AWS Observability skill for alarm and dashboard design. This skill focuses on how observability signals feed into the Define→Test→Operate workflow (e.g., alarms as FIS stop conditions).
Treat fault injection as a privileged operation requiring change-management authorization. Always bound blast radius with stop conditions, use least-privilege IAM roles, avoid embedding PII or secrets in experiment descriptions, and enforce encryption at rest and in transit for assessment reports.
Full instructions (SKILL.md)
Source of truth, from aws/agent-toolkit-for-aws.
name: aws-resilience-lifecycle description: > Guides the end-to-end AWS resilience lifecycle integrating Resilience Hub v2, Fault Injection Service, and Application Recovery Controller. Covers the Define → Test → Operate workflow: from policy creation through failure mode assessment, to FIS experiment validation, to ARC operational controls. Applicable when the user wants a complete resilience strategy, needs to connect findings to experiments to controls, or is planning a resilience program. Also applicable for the meta question of whether marking NGRH findings as resolved is enough, whether they are "done" after resolving findings, or how to validate findings before resolving them. Not applicable for resolving or remediating a specific individual finding (see resilience-hub-failure-mode-assessment), or when a single service is explicitly named (e.g. "what FIS experiment should I run"). version: 1
AWS Resilience Lifecycle
Overview
Domain expertise for the integrated resilience lifecycle across three AWS services: Define (Resilience Hub v2 — also called NGRH, New Generation Resilience Hub) → Test (FIS) → Operate (ARC).
Terminology: in this skill an unqualified "Resilience Hub" always means v2 (NGRH / New Generation Resilience Hub, CLI namespace aws resiliencehubv2). v1 (aws resiliencehub) is referenced only explicitly, and only for migration.
The AWS MCP server is recommended for executing this skill's AWS API calls, but it is not required — all operations also work with the AWS CLI directly.
Guardrail — where this skill's own files live (MCP vs local install)
Before reading a reference file, determine how this skill was loaded:
- Loaded via the AWS MCP
retrieve_skilltool: the skill's reference files are not on the local filesystem. Fetch each one throughretrieve_skillwith thefileparameter (e.g.file="references/lifecycle-workflow.md"orfile="references/api-reference.md") — do NOTfile_readthese paths locally or search the filesystem for them. - Installed locally (e.g.
.kiro/skills/aws-resilience-lifecycle/or~/.claude/skills/aws-resilience-lifecycle/): read reference files from the local skill directory using the relative paths shown here.
This applies only to the skill's own reference files; always read and write user or session data in the working directory, never through retrieve_skill.
Execute the full lifecycle
To implement end-to-end resilience across all three services, follow the procedure exactly. See references/lifecycle-workflow.md.
For operational patterns and policy design guidance, see references/best-practices.md.
Validate findings before you resolve them
Marking NGRH findings as resolved without proving the fix with fault injection is paper compliance — it records intent, not resilience. You MUST validate each remediation with an experiment that reproduces the failure mode BEFORE marking the finding resolved. Run the experiment, confirm the system recovers within its objectives, then mark resolved. Marking resolved first and validating "later" is the anti-pattern.
Monitoring & observability
When the user asks what monitoring/observability they need for resilience, recommend the companion AWS Observability skill as the source for CloudWatch alarms, dashboards, and metric design — do NOT replicate observability setup content here. Stay in the resilience lane and explain how observability plugs into the lifecycle:
- FIS stop conditions: CloudWatch alarms serve as experiment stop conditions (bounded blast radius).
- Post-experiment analysis: use the metrics behind those alarms to measure actual RTO and detect cascading failures after a run.
Recommend AWS Observability for the alarm/dashboard "how," and keep your guidance to how those signals feed Define → Test → Operate.
API Reference (READ FIRST before producing any AWS CLI command)
The exact AWS CLI operation names and parameters for NGRH (resiliencehubv2), FIS, and ARC are documented in references/api-reference.md. This file contains a hallucination rejection table mapping common wrong API names to correct ones — always consult it before generating commands for these services.
Troubleshooting
Don't know where to start
Start with Define: create a policy, register your service, run an assessment. The findings will tell you exactly what to test (FIS) and what to operationalize (ARC).
Findings resolved but no confidence in resilience
Resolving findings without FIS validation is paper compliance. Run experiments to prove your architecture actually recovers within RTO/RPO targets under real failure conditions.
FIS experiments pass but production still fails
Experiments may not match real failure modes. Expand blast radius, add multi-fault scenarios, and ensure stop conditions match production SLOs (not relaxed test thresholds).
Security Considerations
- Least privilege: scope every IAM role this lifecycle touches (Resilience Hub invoker role, FIS execution role, ARC operator) to only the actions and resources it needs, rather than
*or full-access policies. - Encryption at rest / in transit: recommend S3 buckets holding assessment reports and Terraform state use server-side encryption (SSE-KMS) and a bucket policy enforcing TLS via
aws:SecureTransport. - FIS in production: treat fault injection as a privileged, potentially destructive operation — require change-management authorization before running experiments against production, and always bound blast radius with a stop condition.
- Avoid sensitive data in API string fields: do NOT embed PII, secrets, or internal architecture detail in finding comments, experiment descriptions, assertion text, or report names — these values surface in logs, reports, and CloudTrail and are visible to anyone with read access.
- Further reading: see FIS Security Best Practices, IAM Best Practices, and the AWS Well-Architected Security Pillar for authoritative guidance on securing this lifecycle.
Related skills
More from aws/agent-toolkit-for-aws and the wider catalog.

aws-sdk-js-v3-usage
AWS SDK for JavaScript v3 development patterns and best practices.

aws-sdk-python-usage
AWS SDK for Python (boto3/botocore) development patterns and best practices.

aws-sdk-swift-usage
AWS SDK for Swift patterns and async client usage for S3, DynamoDB, CloudWatch, and other AWS services.

aws-secrets-manager
Safely use AWS Secrets Manager secrets in agent commands without exposing plaintext to the LLM context.

aws-security
Unified AWS security posture, threat detection, and compliance findings across Security Hub, GuardDuty, Inspector, Macie, and Detective.

aws-serverless
Build, deploy, and optimize serverless applications on AWS Lambda, API Gateway, Step Functions, and EventBridge.