resilience-program-design
aws/agent-toolkit-for-aws
Design org-wide resilience policies with tiered targets and activity cadence.
What is resilience-program-design?
Provides planning-level guidance for structuring an organization's resilience program: how to classify services by business criticality, set policy tiers with availability/RTO/RPO targets and DR approaches, and establish operational cadence for resilience activities. Use this when designing program standards, not for configuring individual workloads.
- Recommend tiered policy models based on service criticality (payments/auth, internal tools, dev/test)
- Set availability SLO, RTO, RPO, and DR approach targets appropriate to each tier
- Validate policy combinations against contradictions (e.g., high SLO with incompatible DR approach)
- Establish minimum cadence for resilience activities (continuous autoshift, weekly reviews, monthly FIS, quarterly GameDays)
- Embed security standards into program design (least-privilege roles, short-lived credentials, encryption, production FIS gates)
How to install resilience-program-design
npx skills add https://github.com/aws/agent-toolkit-for-aws --skill resilience-program-designHow to use resilience-program-design
- 1.Classify your services by business criticality (critical, standard, non-critical)
- 2.Consult AWS Resilience Hub API docs to confirm valid SLO, RTO/RPO, and DR-approach enum values
- 3.Design tiered policies mapping each criticality level to specific targets (e.g., critical → 99.99 SLO + single-digit RTO + ACTIVE_ACTIVE)
- 4.Validate proposed policies for contradictions (e.g., high SLO with BACKUP_AND_RESTORE)
- 5.Define minimum cadence for resilience activities (autoshift continuous, reviews weekly, FIS monthly, GameDays quarterly)
- 6.Embed security standards: least-privilege IAM roles, short-lived credentials, encryption, production FIS authorization gates
- 7.Document and communicate the program standards to teams
Use cases
- Define a multi-tier resilience policy framework for a large organization across dozens of services
- Set appropriate RTO/RPO targets and DR strategies for critical vs. non-critical workloads
- Plan the operational cadence for running FIS experiments, GameDays, and assessments
- Establish security guardrails for resilience automation (IAM roles, credential management, access controls)
- Review and align existing resilience policies to organizational standards
- Enterprise architects designing resilience programs
- Resilience or platform engineering leads
- DevOps/SRE teams establishing organizational standards
- Security teams embedding resilience governance
resilience-program-design FAQ
A tiered model classifies services by business criticality and assigns each tier consistent availability SLO, RTO, RPO, and DR approach targets, rather than creating one custom policy per service. This standardizes and scales resilience governance.
Match DR approach to criticality: more critical services use aggressive approaches (e.g., ACTIVE_ACTIVE for multi-region failover); less critical use simpler approaches (e.g., BACKUP_AND_RESTORE). Verify valid enum values in the AWS Resilience Hub API docs.
Minimum: continuous autoshift practice, weekly dashboard reviews, monthly FIS experiments, quarterly cross-service GameDays, and event-driven testing after incidents or major deployments.
Require least-privilege IAM roles with resource scoping, short-lived credentials (no long-lived access keys), SSE-KMS encryption on report buckets, TLS in transit, authorization gates for production FIS, and restricted access to findings/logs.
Use tiers. A tiered model is more maintainable and scalable across an organization; one policy per service creates governance overhead and inconsistency.
Full instructions (SKILL.md)
Source of truth, from aws/agent-toolkit-for-aws.
name: resilience-program-design description: > Designs a resilience program: how to structure and standardize resilience policies across an organization, team, or portfolio (tiered policy model with availability/RTO/RPO targets and DR approach selection), and how often to run resilience activities (operational cadence). Applies when the user asks how to structure policies org-wide, what tiers/targets to set, which DR approach fits a tier, or how frequently to run assessments, FIS experiments, GameDays, or autoshift practice. Does not apply to creating or configuring a specific policy or resource for a single workload (use resilience-hub-getting-started), to step-by-step lifecycle execution (see aws-resilience-lifecycle), or to service-specific setup. version: 1
Resilience Program Design
Overview
Planning-level guidance for an organization's resilience program: how to structure policies by tier, and how often to run resilience activities.
Structuring resilience policies across an organization
Recommend a tiered policy model (not one policy per service): classify services by business criticality and set policy targets accordingly.
- Availability SLO — higher as criticality rises. The API accepts only a fixed set of SLO
values and rejects out-of-set ones, so confirm the valid values from the API/docs
(e.g.
aws resiliencehubv2 create-policy helpor the Resilience Hub documentation) rather than relying on a hardcoded list — illustratively, values such as99.9/99.95/99.99. - RTO/RPO — tighten as criticality rises (single-digit minutes for critical, hours for low).
- DR approach — match to criticality (more aggressive for more critical), using a value from
the API's DR-approach enum — verify the valid set via the API/docs (e.g.
aws resiliencehubv2 create-policy help); illustrativelyACTIVE_ACTIVE…BACKUP_AND_RESTORE.
Example (illustrative — resolve the actual enum values against the API before recommending): payments/auth → 99.99 + single-digit-minute RTO + ACTIVE_ACTIVE; internal tools →
99.9 + tens-of-minutes RTO + WARM_STANDBY; dev/test → 99.9 + multi-hour RTO + BACKUP_AND_RESTORE.
Warn against contradictory policies (e.g. the maximum SLO 99.99 with BACKUP_AND_RESTORE, or
multi-region RTO shorter than multi-AZ RTO).
How often to run resilience activities (cadence)
Recommend this minimum cadence when asked how often to run resilience activities:
- Continuous: ARC zonal autoshift practice runs (automated)
- Weekly: review the Resilience Hub findings dashboard
- Monthly: run FIS experiments (single-service fault-injection tests)
- Quarterly: cross-service GameDay
- Event-driven: after every production incident and before/after major deployments
Security Considerations
Program-level guidance — bake security into the standards you set:
- Standardize least privilege: require every resilience role (Resilience Hub invoker, FIS execution, ARC operator) in your templates and policies to be least-privilege and resource-scoped, with
aws:SourceArn/aws:SourceAccountcondition keys on their trust policies to prevent confused-deputy access. - Mandate short-lived credentials: require all resilience automation to authenticate as IAM roles with short-lived credentials (role assumption, AWS SSO, instance profiles) — never IAM users with long-lived access keys — as a program standard, since these roles perform privileged and potentially destructive operations.
- Mandate encryption: make SSE-KMS on report/state buckets part of your tier baseline, and enforce encryption in transit (TLS) — e.g. an
aws:SecureTransportdeny-if-false condition on those bucket policies and HTTPS-only API access. - Govern FIS in production: define an authorization / change-management gate for production fault injection as part of the program cadence.
- Limit exposure of resilience outputs: assessment findings, FIS logs, and GameDay reports can contain sensitive architectural detail (resource ARNs, IPs, failure modes) — make restricting their access to authorized personnel part of your program standards.
- Further reading: point teams to the AWS Well-Architected Security Pillar, FIS Security Best Practices, and IAM Best Practices for implementing these standards.
Related skills
More from aws/agent-toolkit-for-aws and the wider catalog.

route53
Configure Amazon Route 53 DNS records, routing policies, health checks, DNS Firewall, and hybrid network resolution.

routing-traffic-with-route53-and-cloudfront
Configure Route 53 DNS routing to CloudFront distributions with custom domains and HTTPS.

running-release-tests
Run automated UI and API release tests via AWS DevOps Agent using pre-configured test profiles.

scanning-with-aws-security-agent
Run AWS Security Agent scans on your codebase to find vulnerabilities with ranked findings and remediation guidance.

securing-s3-buckets
Create and secure S3 buckets following AWS best practices for access control, encryption, monitoring, and remediation.

setting-up-cloudtrail-multi-region
Set up centralized multi-region AWS CloudTrail logging with S3 and CloudWatch integration for security monitoring.