To validate your multi-AZ architecture’s resilience, run a 48-72 hour Availability Zone evacuation drill using ARC Zonal Shift to shift traffic away from an impaired AZ across Amazon ECS, EKS, RDS for PostgreSQL, and Aurora PostgreSQL, then monitor failover and restore procedures.


Prerequisites and Setup for Zonal Shift Drill
Ensure you have ARC Zonal Shift enabled in your AWS account and the necessary IAM permissions to create and manage zonal shifts. Your workloads must be deployed across at least two Availability Zones using Amazon ECS services, EKS clusters, RDS for PostgreSQL instances, and Aurora PostgreSQL clusters with multi-AZ configurations. Confirm that your application traffic is distributed via Application Load Balancers or equivalent routing mechanisms that respect zonal shift controls.
Identify the target Availability Zone to simulate impairment and verify that your observability stack includes CloudWatch metrics for request latency, error rates, and healthy host counts per AZ. Tag your resources consistently to enable filtering during the drill. Pre-stage the AWS CLI with the arc-zonal-shift profile configured and test read-only access to shift status before initiating any traffic changes.
Document your baseline performance metrics including normal request throughput, database replication lag, and container restart rates. Establish a rollback plan that outlines the exact CLI commands to cancel the zonal shift and restore traffic to the impaired AZ. Notify stakeholders of the drill window and establish a communication protocol for incident updates during the 48-72 hour window.
Executing the Multi-Day Zonal Shift
Initiate the zonal shift using the AWS CLI command: aws arc-zonal-shift start-zonal-shift --resource-identifier <your-resource-arn> --away-from <availability-zone-id> --expires-in 72h. Replace <your-resource-arn> with the ARN of your load balancer, service, or cluster and <availability-zone-id> with the target AZ (e.g., useast1-az1). The shift begins immediately and prevents new traffic from entering the impaired AZ while allowing existing connections to drain.
Monitor the shift status every 15 minutes using aws arc-zonal-shift get-zonal-shift --resource-identifier <your-resource-arn> to confirm the shift remains active and track the expiration time. Observe application metrics: verify that request rates drop to zero in the impaired AZ and increase correspondingly in the healthy AZ. Check ECS task placement, EKS pod scheduling, and RDS/Aurora failover events to ensure workloads have successfully rebalanced.
Simulate realistic failure conditions by stopping network traffic or disabling health checks in the impaired AZ (without terminating resources) to test end-to-end failover. Log application-level metrics such as login success rates, transaction completion times, and cache hit ratios. If using Aurora PostgreSQL, confirm that read replicas remain promoted and that write traffic is directed to the new primary instance in the healthy AZ.
Observability, Validation, and Restore Procedures
Track key performance indicators during the drill: target 99.9% request success rate, less than 100ms p95 latency increase, and zero failed health checks in the receiving AZ. Use CloudWatch Logs Insights to correlate errors with AZ-specific dimensions and validate that no critical alerts are firing due to the shift itself. For databases, monitor replication lag (should remain under 1 second for Aurora) and verify that backup and snapshot processes continue uninterrupted.
To restore traffic, execute: aws arc-zonal-shift cancel-zonal-shift --resource-identifier <your-resource-arn>. This immediately allows new traffic to return to the previously impaired AZ. Monitor rebalancing as tasks, pods, and connections gradually redistribute. Confirm that database write traffic resumes in the original AZ only after explicit failback procedures (if using Aurora global databases or manual promotion).
After full restoration, conduct a 15-minute stability window to ensure no oscillations or thrashing occur. Generate a post-drill report comparing baseline vs. drill metrics, including any observed delays in failover, connection draining times, or metric gaps. Use findings to adjust auto-scaling thresholds, health check grace periods, or routing weights for future resilience improvements.
What to do next
After completing the drill, review your observability dashboards to confirm that failover occurred within expected timeframes and that no data loss or prolonged degradation was observed. Update your runbooks with any lessons learned, particularly around connection draining delays or database failover timing. Schedule quarterly zonal shift drills to maintain readiness for real AZ impairments.
FAQ
How do I ensure my RDS for PostgreSQL instance fails over correctly during a zonal shift?
Confirm that your RDS instance has multi-AZ enabled and that the primary and standby are in different AZs. During the shift, monitor the AWS/RDS metric 'DatabaseConnections' to see connections drop in the impaired AZ and rise in the standby AZ. Verify that failover completes within 30 seconds using the 'ReplicaLag' metric, which should spike then return to near zero.
Can I run a zonal shift drill on an Aurora PostgreSQL cluster without disrupting backups?
Yes, Aurora backups and snapshots continue unaffected during a zonal shift because they are managed at the storage layer. Monitor the 'BackupRetentionPeriod' and 'SnapshotCreateTime' CloudWatch metrics to confirm backup processes remain uninterrupted while traffic is shifted away from the impaired AZ.
Source: Running multi-day AZ evacuation drills with ARC Zonal Shift (AWS).



