arc-region-switch
aws/agent-toolkit-for-aws
Expert guidance on Amazon Application Recovery Controller (ARC) Region switch for cross-Region failover orchestration.
What is arc-region-switch?
Answers technical and positioning questions about ARC Region switch, the AWS service for orchestrating cross-Region workload failover and switchover. Use this when designing multi-Region recovery plans, troubleshooting execution, or preparing customer engagement on ARC adoption.
- Explains ARC Region switch architecture including plans, workflows, steps, and execution blocks
- Clarifies active/passive vs active/active deployment patterns and when to use each
- Describes execution modes (graceful, ungraceful, post-recovery) and their appropriate use cases
- Covers plan evaluation, automatic execution reports, and recovery time tracking
- Provides customer positioning language and SA engagement guidance
- Distinguishes Region switch from in-Region failover, data replication, and AZ-level resiliency
How to install arc-region-switch
npx skills add https://github.com/aws/agent-toolkit-for-aws --skill arc-region-switchHow to use arc-region-switch
- 1.Ask questions about ARC Region switch architecture, workflows, execution blocks, or execution modes
- 2.Provide context (e.g., active/passive vs active/active, RTO target, Region topology) for tailored guidance
- 3.Request customer positioning language or SA engagement talking points when preparing deals or reviews
- 4.Consult the skill for terminology guidance (e.g., 'Region switch' vs 'Routing Controls') to ensure accurate communication
- 5.Reference the skill's critical warnings (irreversible steps, regional endpoint requirements, monitoring during events) when designing or executing plans
Use cases
- Designing a multi-Region failover plan for an application already deployed across Regions
- Troubleshooting plan execution failures or understanding why a step did not run as expected
- Comparing active/passive vs active/active strategies for a customer's RTO/RPO requirements
- Preparing a sales or solutions architect engagement on ARC Region switch adoption
- Testing a Region switch plan execution and interpreting the recovery time results
- Solutions Architects designing multi-Region disaster recovery
- DevOps and platform engineers implementing ARC Region switch plans
- Technical account managers preparing customer engagement on ARC
- Support engineers troubleshooting plan execution issues
- Architects evaluating cross-Region failover strategies
arc-region-switch FAQ
Region switch is the plan-based orchestration service for cross-Region failover. Routing Controls is a legacy cluster-based approach or a specific execution block type within Region switch. Do not use 'Routing Controls' as a synonym for Region switch.
No. Region switch orchestrates failover of existing replicas and resources. Data replication is handled by separate AWS services (Aurora Global Database, DynamoDB Global Tables, S3 Cross-Region Replication, etc.). Region switch assumes your application is already deployed multi-Region.
Graceful execution runs all steps in orderly sequence and is preferred for planned switchovers and DR tests. Ungraceful execution skips or modifies certain blocks and is used only when the source Region is impaired and graceful execution is not possible.
Actual recovery time is the plan execution time plus the time for application health alarms to return to green. This is visible on the plan execution details page and can be compared against your Recovery Time Objective (RTO).
No. Plan executions can be paused or cancelled, but any step already started cannot be reversed without running another plan execution. Plan carefully and test thoroughly before automating with triggers.
Full instructions (SKILL.md)
Source of truth, from aws/agent-toolkit-for-aws.
name: arc-region-switch description: "Answers questions about Amazon Application Recovery Controller (ARC) Region switch including architecture, plans, execution blocks, workflows, triggers, active/active vs active/passive, cross-account support, recovery time, dashboards, and customer positioning. Applicable when users ask about ARC Region switch adoption, design, or troubleshooting." version: 1
ARC Region switch Expert
Overview
Makes the agent an expert on Amazon Application Recovery Controller (ARC) Region switch — the feature for orchestrating cross-Region workload failover and switchover. Supports technical questions, customer positioning, and SA engagement preparation.
Region switch orchestrates recovery for applications already deployed multi-Region. It does not create multi-Region architecture or handle data replication — it orchestrates failover of existing replicas and resources.
Guardrail — where this skill's own files live (MCP vs local install)
Before reading a reference file, determine how this skill was loaded:
- Loaded via the AWS MCP
retrieve_skilltool: the skill's reference files are not on the local filesystem. Fetch each one throughretrieve_skillwith thefileparameter (e.g.file="references/positioning.md"orfile="references/doc-links.md") — do NOTfile_readthese paths locally or search the filesystem for them. - Installed locally (e.g.
.kiro/skills/arc-region-switch/or~/.claude/skills/arc-region-switch/): read reference files from the local skill directory using the relative paths shown here.
This applies only to the skill's own reference files; always read and write user or session data in the working directory, never through retrieve_skill.
When Not to Use
- In-Region failover — Region switch is for cross-Region recovery only. Use AZ-level mechanisms (ALB, Auto Scaling) for in-Region resilience.
- Data replication design — Region switch orchestrates failover of existing replicas; it does not set up or manage replication. Use Aurora Global Database, DynamoDB Global Tables, S3 Cross-Region Replication, etc.
- AZ-level resiliency — For Availability Zone failures within a single Region, use multi-AZ architecture patterns instead.
Terminology Constraints
- ALWAYS use "Region switch" (lowercase 's') for the product name
- Use "Routing Controls" only when referring to the Routing Controls execution block or the legacy cluster-based approach — do not use it as a synonym for Region switch
- Do NOT conflate triggers (CloudWatch alarms that start execution) with application health alarms (measure actual recovery time)
- Region switch orchestrates failover — it does NOT replicate data
Critical Warnings
- Irreversible in-progress steps: Plan executions can be paused or cancelled, but any step that is already started cannot be reversed without another plan execution.
- Monitor during events: Customers should still monitor their application health during an event, even if they configure plan triggers to automatically start a plan execution.
- Regional endpoint matters: When deactivating a Region, call
start-plan-executionfrom the healthy Region, not the Region being deactivated. When activating a Region, call from the Region being activated. See StartPlanExecution API.
Workflow
- Classify the question: technical architecture, customer positioning, how-to, troubleshooting, or comparison
- Answer from embedded knowledge in this skill
- If the knowledge base is insufficient, search official AWS documentation (
docs.aws.amazon.com) - Format the response for the audience (engineer, SA, customer)
Always validate:
- Correct terminology (Region switch, not routing controls for plan-based features)
- Include doc links where helpful (see Documentation Links section)
- Use positioning language from the Positioning section
Architecture
Components
| Component | Description |
|---|---|
| Plan | Top-level resource scoped to a multi-Region application. Contains workflows. |
| Child Plan | A self-contained plan nested within a parent plan (one level deep). |
| Workflow | Ordered sequence of steps within a plan. Defines activation/deactivation logic. |
| Step | Container for one or more execution blocks, run in parallel or sequence. |
| Execution Block | Performs a specific recovery action (e.g., scale up, reroute traffic, failover DB). |
| Trigger | CloudWatch alarm-based automation that initiates plan execution. |
| Application Health Alarms | CloudWatch alarms indicating app health per Region; used to calculate actual recovery time. |
| Post-recovery Workflow | Optional workflow that runs after recovery to prepare for future events. |
| Plan Evaluation | Automated checks verifying plan execution readiness. Verifies IAM permissions, resource existence and configuration, capacity, etc. |
| Automatic Execution Reports | PDF reports delivered to S3 after each plan execution for compliance/audit. |
Execution Modes
Recommend using graceful execution unless not possible (e.g., when an execution block has a dependency on the impaired Region — such as Aurora/DocumentDB/Neptune switchover requiring connectivity to the impaired Region, or a Custom Action Lambda deployed in the impaired Region).
- Graceful: Runs all steps in orderly sequence. Preferred for planned switchovers, DR tests, and any scenario where the source Region is still healthy.
- Ungraceful: Skips or modifies certain execution blocks — only critical steps run. Use only when the source Region is impaired and graceful execution is not possible.
- Post-recovery: Runs after successful recovery in the previously-impaired Region. Requires both Regions to be healthy. Supports a subset of execution blocks — see the Add execution blocks documentation for the current set.
Active/Passive vs Active/Active
| Approach | Workflows Needed | Behavior |
|---|---|---|
| Active/Passive | 1 activation workflow (either Region) OR 2 separate activation workflows (one per Region) | Failover from primary to standby; failback when primary recovers |
| Active/Active | 1 activation workflow + 1 deactivation workflow per Region | Shift-away from impaired Region + return when healthy |
Supported Execution Blocks
Execution blocks are the individual step types a Region switch workflow is composed of — each performs one recovery action, spanning traffic/DNS rerouting, compute scaling, database failover, custom-action Lambdas, manual-approval gates, and nested child plans.
Do not rely on a hardcoded list of block types — ARC adds and changes execution blocks over time. Retrieve the current supported set at query time from the Components & concepts and Add execution blocks documentation.
Recovery Time Tracking
- Recovery Time Objective (RTO): Set when creating a plan
- Actual Recovery Time: Plan execution time + time for application health alarms to return to green
- Visible on plan execution details page for comparison against RTO
Plan Evaluation
- Validates: IAM permissions, resource configurations, running capacity
- Warnings surfaced in console, EventBridge, and API
- Passing evaluation alone is NOT sufficient — always test by executing plans
Automatic Execution Reports
- PDF reports generated after each plan execution
- Delivered to customer-specified S3 bucket (within ~30 min)
- Customers must configure the S3 bucket and update permissions for the PlanExecutionRole to enable reporting
- Contents: executive summary, plan config, execution timeline, resource states, alarm history, child plan details, glossary
- Useful for regulatory compliance and DR audit evidence
- See Security Considerations for encryption and access control guidance
Cross-Account Support
Plans can orchestrate resources across multiple AWS accounts via IAM roles with cross-account trust policies. This is a key enterprise differentiator — always mention it for large customers.
When configuring cross-account trust policies:
- Include condition keys (
aws:SourceArn,aws:SourceAccount,sts:ExternalId) to prevent confused deputy attacks - Scope IAM policies to least privilege — avoid
*resource wildcards andFullAccessmanaged policies - Scope permissions to only the specific resources (ASG ARNs, Aurora cluster ARNs, Route 53 health check ARNs, etc.) referenced in execution blocks
Regional Availability
Available in multiple commercial AWS Regions and AWS GovCloud (US) Regions — always verify the current list before stating availability to a customer, as Region coverage changes over time. Each Region has its own data-plane endpoint (arc-region-switch.<region>.api.aws), ensuring execution doesn't depend on the impaired Region.
Verify the complete list of available regions/endpoints at AWS Regions & endpoints.
Security Considerations
IAM Least Privilege
- Scope cross-account IAM roles to only the specific resources referenced in execution blocks (ASG ARNs, Aurora cluster ARNs, Route 53 health check ARNs, Lambda function ARNs, etc.)
- Avoid
*resource wildcards andFullAccessmanaged policies - Include condition keys (
aws:SourceArn,aws:SourceAccount,sts:ExternalId) in cross-account trust policies to prevent confused deputy attacks
Execution Reports S3 Bucket
- Enable default encryption (SSE-KMS preferred) on the reports S3 bucket
- Add a bucket policy denying requests where
aws:SecureTransportisfalse(enforce TLS) - Restrict bucket access to authorized personnel only — reports contain sensitive infrastructure details (plan config, execution timeline, resource states, alarm history)
- Enable S3 bucket versioning and MFA Delete for tamper protection
- Ensure the bucket is not publicly accessible
Custom Action Lambda Security
- Apply least-privilege execution roles to Custom Action Lambda functions
- Validate inputs within Lambda functions
- Do not embed secrets in Lambda environment variables — use Secrets Manager or Parameter Store
Notification & Event Targets
- Restrict EventBridge rule targets (SNS topics, Lambda functions, etc.) that receive plan-evaluation warnings and execution events to authorized recipients only
- Lock down SNS topic subscription policies and Lambda resource policies so sensitive infrastructure details (plan configuration, resource ARNs, execution state) are not exposed to unauthorized parties
Logging and Monitoring
- Enable AWS CloudTrail for auditing all ARC Region switch API calls
- Configure CloudWatch alarms for unexpected or unauthorized plan executions
- Enable S3 access logging on the execution reports bucket
- See Logging and monitoring for Region switch
Positioning
Customer-facing framing, the Region switch vs Routing Controls comparison, analyst talking points, and per-audience conversation guidance are maintained in Positioning. Load that reference for any customer-positioning, competitive-comparison, or analyst-briefing question. Key rules that always apply:
- Use "Region switch" (plan-based orchestration) framing; do NOT present legacy "routing controls" / "ARC clusters" language as the Region switch (plan-based) approach.
- Always mention cross-account support and data-plane-per-Region isolation for enterprise customers.
Documentation Links
The curated documentation index and the "when to link which doc" guidance live in Documentation Links. Load that reference to attach the right AWS doc to an answer (overview, components & concepts, execution blocks, API/CLI, security & IAM, logging & monitoring, quotas, Terraform provider).
Troubleshooting
Customer confuses triggers with health alarms
Triggers are CloudWatch alarms that start plan execution. Application health alarms measure when recovery is complete. They serve different purposes and are configured separately.
Customer assumes Region switch handles data sync
Clarify: Region switch orchestrates failover of existing replicas (e.g., Aurora Global DB promotion). The customer must set up multi-Region data replication independently.
Cross-account execution fails
Usually missing IAM permissions. Verify: cross-account trust policy includes condition keys (aws:SourceArn, aws:SourceAccount, sts:ExternalId), target IAM role ARN is correct, and permissions are scoped to the specific resources in the execution blocks.
Plan evaluation warnings
Warnings indicate IAM, resource, or capacity issues. Fix the underlying issue — but note that passing evaluation alone isn't sufficient; always test by executing plans.
Wrong Regional endpoint used
When deactivating a Region, start-plan-execution MUST be called from the healthy Region. When activating a Region, it MUST be called from the Region being activated. Using the wrong endpoint will fail or produce unexpected behavior. See StartPlanExecution API.
Related skills
More from aws/agent-toolkit-for-aws and the wider catalog.

aurora-dsql
Provision, manage, and query Aurora DSQL serverless distributed SQL with safe IAM auth and multi-tenant patterns.

aws-ai-ml
Fine-tune and deploy AI models on Amazon SageMaker with guided workflows for the full ML lifecycle.

aws-amplify
Build and deploy full-stack web and mobile apps with AWS Amplify Gen2 using TypeScript code-first approach.

aws-auth
Amazon Cognito user authentication for web and mobile apps with sign-up, sign-in, MFA, OAuth, and token management.

aws-billing-and-cost-management
Analyze AWS costs, optimize spending, and manage budgets with expert domain knowledge.

aws-blocks
Infrastructure-from-Code framework for building full-stack AWS applications with pre-built Building Blocks.