AWS Lesson 89 of 123

AWS Enterprise Architecture: Disaster Recovery Strategies

In a nutshell

Disaster recovery (DR) is your plan for the big bad day — not a single server dying (that is everyday high availability), but a whole AWS Region going dark, ransomware encrypting everything it can reach, or a bad change that corrupts your data and faithfully replicates the corruption everywhere. DR asks two blunt questions and demands numbers, not adjectives: how long can you be down (your RTO — Recovery Time Objective) and how much recent data can you afford to lose (your RPO — Recovery Point Objective). “We have backups” is not an answer to either.

The mental model is insurance, and it comes in four policy tiers. Think of getting back on the road after your car dies. Backup & Restore is a spare tire in the trunk: cheap to carry, but you have to stop, jack the car up, and fit it — slow (hours). Pilot Light is a second car in the garage, fuelled and maintained but engine cold: the important part (the data) is ready, you just have to start it and let it warm up (tens of minutes). Warm Standby is that second car already idling in the driveway with a smaller engine — hop in, press the accelerator, go (minutes). Multi-Site Active-Active is two cars already driving the route side by side: lose one and the other simply carries everyone, no stopping at all (near-zero) — and you pay for two cars every single day. You do not buy one policy for the whole company; you buy the cheapest tier that meets each workload’s agreed RTO/RPO, and — the step everyone skips — you actually test that the spare works before you need it.

Level: Advanced · Time: ~50 min

Before this lesson, you should be comfortable with: AWS Regions vs Availability Zones, the basics of EC2 / RDS / S3, and ideally single-Region high availability (Multi-AZ). It helps to have met Aurora Global Database, AWS Backup with Vault Lock, and the companion multi-Region / active-active lesson, though this article reintroduces each as it needs them.

After this lesson you will be able to:

Every disaster recovery conversation that goes wrong starts the same way: someone asks “are we covered if a region goes down?” and someone else answers “yes, we have backups.” Those two sentences are about completely different things, and the gap between them is where outages turn into resignations. “Covered” is a question about how much data you can afford to lose and how long you can afford to be down — RPO and RTO — and the honest answer is a number with a price tag, not a yes. Disaster recovery on AWS is not one architecture; it is a spectrum of four, each buying you a tighter RTO/RPO for more money and more operational discipline. The art is choosing the cheapest one that meets your actual recovery objectives, building it so it genuinely works, and — the part everyone skips — proving it works on a schedule.

This article lays out the four canonical AWS DR strategies the way the AWS Disaster Recovery of Workloads on AWS guidance frames them — Backup & Restore, Pilot Light, Warm Standby, and Multi-Site Active-Active — but as a single decision framework rather than four disconnected diagrams. We will anchor everything to RTO (recovery time objective — how long until you are serving again) and RPO (recovery point objective — how much recent data you lose), because those two numbers are the only honest way to compare them. Along the way: the AWS services that implement each tier (AWS Backup, S3 Cross-Region Replication, Aurora Global Database, DynamoDB Global Tables, AWS Elastic Disaster Recovery, Route 53 Application Recovery Controller), how to express them in Terraform, how to orchestrate an actual failover, and how to keep the bill from quietly turning a passive standby into an expensive insurance policy nobody has ever cashed.

The business scenario

The thing that makes DR universally relevant is that the driver is never “we want DR” — it is a specific loss the business cannot absorb, and that loss looks different at every size of company while the underlying maths is identical.

What unites all three is the same uncomfortable trio of facts. First, the failures that actually require DR are correlated and large — an AZ-spanning power or network event, a regional control-plane degradation, a bad region-wide config change, ransomware that encrypts your primary and anything it can reach, or a data-corruption bug that faithfully replicates to your replica. Multi-AZ (which you should already have) handles the small ones; DR is about the big ones, including the ones that come from inside the house. Second, the cost of recovery is dominated by what you keep running while nothing is wrong — a hot standby costs the same on a quiet Tuesday as during a disaster, so over-provisioning DR is a tax you pay every single day for an event that may never come. Third, an untested recovery plan is not a plan — it is a document, and documents do not fail over.

The problem this architecture solves, stated precisely: map each workload to a recovery objective (RTO/RPO) the business has actually signed off on, implement the cheapest DR strategy that meets it using native AWS services, orchestrate the failover so it is deterministic rather than heroic, and verify it on a recurring schedule so the RTO/RPO you claim is the RTO/RPO you can deliver. The rest of this article is the map from “RTO/RPO target” to “which of the four, built how.”

Architecture overview

AWS disaster recovery reference architecture showing a primary Region (us-east-1) replicating to a recovery Region (us-west-2) across four strategy lanes — Backup & Restore, Pilot Light, Warm Standby, and Multi-Site Active-Active — under a shared Route 53 / ARC / Global Accelerator failover control plane, with an immutable cross-account backup vault as the ransomware backstop.

The four strategies are best understood as one axis — increasing investment in pre-provisioned, pre-warmed, pre-replicated infrastructure in a recovery Region — sliding RTO and RPO from “hours/hours” down toward “near-zero/near-zero.” Picture a primary Region (say us-east-1) serving production, and a recovery Region (say us-west-2). The difference between the four strategies is entirely about what is already running, and how current the data is, in us-west-2 before disaster strikes.

Strategy 1 — Backup & Restore (RPO: hours; RTO: hours). Nothing runs in the recovery Region in steady state. AWS Backup takes scheduled snapshots of EBS volumes, RDS/Aurora clusters, DynamoDB tables, EFS, and more, and copies them cross-Region (and ideally cross-account, into a locked vault) on a schedule. Application artifacts live as container images in ECR with cross-Region replication and IaC in version control. When disaster strikes, you build the recovery environment from IaC and restore data from the latest copied backups. The data path in steady state is just “snapshot → copy to Region B → sit in a vault.” The recovery path is “terraform apply in Region B, restore snapshots, repoint DNS.” It is the cheapest by far and the slowest by far.

Strategy 2 — Pilot Light (RPO: minutes; RTO: tens of minutes). The data layer is live and continuously replicating to the recovery Region, but compute is switched off or minimal. An Aurora Global Database secondary cluster sits in us-west-2 receiving storage-level replication (sub-second lag); DynamoDB Global Tables replicate continuously; S3 Cross-Region Replication mirrors objects. The “pilot light” is exactly this always-warm data plus the core scaffolding — the VPC, subnets, security groups, and a scaled-to-zero (or minimal) compute definition — kept current but not serving. On failover you promote the Aurora secondary, scale the compute up from zero to production size, and repoint traffic. You pay for replicated storage and data transfer, but almost nothing for idle compute.

Strategy 3 — Warm Standby (RPO: seconds; RTO: minutes). A fully functional but under-scaled copy of the workload runs in the recovery Region all the time. The data layer replicates as in Pilot Light, and a smaller version of the compute fleet — fewer/smaller ECS tasks or EC2 instances, a smaller Aurora reader — is actually running and could serve traffic right now, just not at full capacity. Failover is promote the database, scale the already-running fleet up, and shift traffic — faster than Pilot Light because nothing has to cold-start from zero. You pay for a continuously-running (if minimal) second environment.

Strategy 4 — Multi-Site Active-Active (RPO: near-zero; RTO: near-zero). Both Regions serve live production traffic simultaneously. There is no “failover” of compute because both fleets are already hot at full-ish capacity; DynamoDB Global Tables are multi-active (both Regions write locally); Aurora Global Database has one writer but a promotable secondary. A regional loss is a capacity event, not a recovery event. This is the most expensive and the most complex, and it gets its own deep treatment in the companion Active-Active Multi-Region article — here it is the top rung of the ladder.

The common control plane across all four: traffic steering and failover orchestration are the same machinery regardless of strategy. Amazon Route 53 provides DNS with health-check-based failover records; Route 53 Application Recovery Controller (ARC) provides routing controls (deterministic on/off switches you flip to redirect traffic, instead of hoping health checks fire correctly) and readiness checks (continuous verification that the standby is actually capable of taking load). For non-cacheable latency-critical APIs, AWS Global Accelerator fails over at the network layer in seconds rather than waiting on DNS TTLs. And for server-based workloads that are hard to re-platform, AWS Elastic Disaster Recovery (DRS) continuously block-level-replicates whole servers into a low-cost staging area in the recovery Region and launches them on demand — effectively a managed Pilot Light / Warm Standby for lift-and-shift estates.

The single most important thing this overview should make obvious: you do not pick one strategy for the company — you pick one per workload, based on its RTO/RPO tier. A mature enterprise runs all four at once: active-active for the payment rail, warm standby for the order system, pilot light for the reporting platform, and backup & restore for the internal wiki — each priced to its actual cost of downtime.

Component breakdown

Component AWS service Role in DR Key configuration choices
Cross-Region backup AWS Backup The foundation of Backup & Restore; defence against corruption/ransomware for all tiers Backup plans with cross-Region copy + cross-account copy to an isolated account; Vault Lock (compliance mode) for immutability; lifecycle to cold storage; backup of RDS, EBS, DynamoDB, EFS, S3, Aurora
Relational replication Aurora Global Database Live cross-Region data plane for Pilot Light / Warm Standby / Active-Active One global cluster; primary writer + ≥1 secondary Region; storage-layer replication (typically <1s lag); managed planned failover for drills; target RPO ~1s
Relational (RDS, non-Aurora) RDS cross-Region read replica Same idea for MySQL/PostgreSQL/MariaDB on RDS Cross-Region read replica that can be promoted to standalone; async replication (RPO = replication lag, seconds–minutes)
NoSQL replication DynamoDB Global Tables Multi-active key/document data; zero-failover for the tiers that use it Global Tables v2 (2019.11.21); PITR enabled; design for last-writer-wins; watch ReplicationLatency
Object replication S3 Cross-Region Replication (CRR) Mirror uploads, exports, static origins, and backups to Region B Versioning on; CRR rules (optionally RTC for a 15-min replication SLA); replication metrics; bidirectional where both Regions write
Server-based DR AWS Elastic Disaster Recovery (DRS) Continuous block replication of whole EC2/on-prem servers; managed pilot-light/warm-standby for lift-and-shift Low-cost staging area with cheap instances + EBS; point-in-time recovery snapshots; launch templates for recovery; supports drill launches into an isolated subnet
Container artifacts Amazon ECR (cross-Region replication) Make the same image digest available in Region B for fast rebuild/scale-up Registry replication rules to Region B; deploy by digest, not tag
DNS failover Route 53 Steer traffic to the healthy Region Failover or latency routing; health checks on a deep /health endpoint; low TTL (30–60s); Evaluate Target Health on alias records
Failover orchestration Route 53 Application Recovery Controller (ARC) Deterministic, audited Region switch + standby readiness assurance Routing controls (manual/automated on-off) gated by safety rules; readiness checks per resource type; multi-Region cluster of 5 endpoints
Network-layer failover AWS Global Accelerator Sub-DNS failover for non-cacheable latency-critical APIs Two anycast IPs; endpoint groups per Region; traffic dials; health checks at the network layer
Keys & secrets KMS (multi-Region keys) + Secrets Manager (replica secrets) Ensure encrypted data and credentials are usable in Region B Multi-Region KMS keys so replicated data decrypts locally; replica secrets so Region B reads DB creds without a cross-Region call
IaC + pipeline Terraform + CodePipeline/GitHub Actions Recreate compute/networking in Region B on demand (Backup & Restore / Pilot Light) Region-parameterised modules; recovery environment is terraform apply, not click-ops; pipeline can deploy Region B independently

A few component choices deserve their why, not just their what:

Why AWS Backup with cross-Region and cross-account copy and Vault Lock, not just snapshots? Because the disaster that most often actually happens is not a region vanishing — it is deletion or encryption of your data, whether by a bad actor, ransomware, or a mistake. A snapshot in the same account that the same compromised credentials can delete is not a backup; it is a hostage. Copying backups into a separate, locked account with Vault Lock in compliance mode (which even the root user cannot shorten or delete before the retention period) is what turns “we have snapshots” into “we can actually recover from ransomware.” This single control protects every tier, including active-active stacks whose live replication would have dutifully copied the corruption to the other Region.

Why Aurora Global Database over RDS cross-Region read replicas for the higher tiers? Both give you a promotable copy in Region B, but Aurora Global Database replicates at the storage layer with typically sub-second lag and far lower RPO, supports managed planned failover (a clean, ~1-minute switch you can rehearse), and decouples replication from the database engine’s own load. RDS cross-Region read replicas use engine-level async replication — perfectly fine for Backup & Restore / Pilot Light on smaller MySQL/PostgreSQL workloads, but with looser, more variable lag. Choose Aurora Global DB when your RPO budget is sub-second and you intend to rehearse failover regularly.

Why Elastic Disaster Recovery (DRS) instead of re-architecting? Because a huge share of enterprise estates is not cloud-native — it is lift-and-shift EC2 (or still on-prem) running software that nobody is going to re-platform onto Fargate and Aurora just to get DR. DRS continuously replicates the entire server (OS, app, data) as block-level changes into a cheap staging area in Region B, and on failover (or a non-disruptive drill) launches production-sized instances from the latest point-in-time. It is the pragmatic way to put a Pilot-Light-grade RTO/RPO on workloads you cannot or will not refactor — and it is far better than a quarterly AMI copy.

Why Route 53 ARC and not just health-check failover? Because DNS health checks fail in the messy middle — a partial brownout where the primary is sick enough to lose data but healthy enough to pass a shallow health check, or healthy enough that flapping health checks bounce traffic back into a degraded Region. ARC routing controls are deterministic switches you (or an automated runbook) flip with intent, protected by safety rules (e.g. “never turn both Regions off,” “always keep at least one on”). ARC readiness checks continuously answer the question that actually matters before you fail over: is the standby genuinely ready to take this load right now? For tier-1 systems, that determinism is worth the added moving part.

A worked RTO/RPO reckoning

The component table tells you what the services are; this section shows the arithmetic that turns “we should do DR” into “we chose Warm Standby for this workload, and here is why.” Do this reckoning once per workload and the strategy usually chooses itself. We will keep the mental picture from the primer — spare tire, cold car, idling car, two cars — and attach real numbers to it.

The workload. Northwind Retail runs a checkout service on Aurora PostgreSQL in us-east-1: the write path that takes orders and payments. Two business inputs come from the people who own the money, not from engineering:

Step 1 — turn the business inputs into two numbers. Downtime tolerance becomes the RTO; data-loss tolerance becomes the RPO. Here the business signs off on RTO ≤ 15 minutes and RPO ≤ 10 seconds. Write them down and get them approved — every later decision is measured against these two numbers, and “approved by the business” is what makes them real rather than an engineer’s guess.

Step 2 — walk up the ladder and stop at the first tier that meets both. The rule is to buy the cheapest policy that clears the bar, not the fanciest.

Tier RPO you would actually get RTO you would actually get Meets 15 min / 10 s?
Backup & Restore Time since last recoverable copy. Nightly snapshot at 02:00 copied cross-Region → a Region loss at 14:00 loses ~12 hours of orders Rebuild Region B from IaC + restore snapshot ≈ 2–4 hours ❌ fails both
Pilot Light Aurora Global Database replicating continuously → ~1 s ✔ Promote the secondary (~1 min) + scale compute from zero (cold-start ECS/EC2, warm the fleet) ≈ 20–35 min ✔ RPO / ❌ RTO
Warm Standby Same live replication → seconds ✔ Promote the secondary + scale an already-running minimal fleet up ≈ 3–6 min ✔ both — stop here
Multi-Site Active-Active Near-zero ✔ Near-zero (capacity event, not a failover) ✔ ✔ both, but ~2× cost + write-conflict complexity you do not need

The first tier that clears both numbers is Warm Standby, so that is the answer. Active-Active would also clear the bar, but it costs roughly twice as much every day and adds multi-writer conflict handling — you only climb that last rung when there is a separate reason (users on several continents needing low write latency), not merely a failover need.

Step 3 — why each rejected tier fails, in one line each (this is the teaching part).

Step 4 — sanity-check the money. The tier you pick is a bet: you pay a steady monthly premium to cap a per-event loss. Put both sides on the table.

The two formulas worth memorising. For backup-based tiers, RPO ≈ time since last successful, *copied*, restorable backup — worst case, the full backup-plus-copy interval (a copy that silently failed makes it worse, which is why copy-job success is a paging alarm). For replication-based tiers, RPO ≈ current replication lag at the instant of loss — which is not constant, so you alarm on the lag metric (AuroraGlobalDBRPOLag, DynamoDB ReplicationLatency, S3 CRR metrics) rather than assuming yesterday’s number. And always quote the unplanned RPO in your SLA: a planned failover drains in-flight writes and can reach ~0, but a real Region loss keeps whatever was in flight, so the number you promise the business is the unplanned one.

Implementation guidance

Infrastructure as Code is the load-bearing wall of Backup & Restore and Pilot Light, because in those strategies the recovery environment’s compute and networking do not exist (or barely exist) until you create them. If “recovery” means an engineer hand-clicking a VPC together under pressure at 3am, your real RTO is “however long that takes, plus the mistakes.” Express the workload as a region-parameterised Terraform module so that standing up Region B is terraform apply against a second provider alias — not archaeology.

The clean structure:

Terraform shape for the AWS Backup plan with cross-Region copy and an immutable vault (illustrative):

resource "aws_backup_vault" "dr" {
  provider = aws.usw2                       # vault in the recovery Region
  name     = "dr-locked-vault"
  kms_key_arn = aws_kms_key.backup_usw2.arn
}

# Compliance-mode lock: cannot be deleted/shortened before retention elapses
resource "aws_backup_vault_lock_configuration" "dr" {
  provider            = aws.usw2
  backup_vault_name   = aws_backup_vault.dr.name
  min_retention_days  = 35
  changeable_for_days = 3                    # cooling-off before lock is permanent
}

resource "aws_backup_plan" "core" {
  name = "core-cross-region"

  rule {
    rule_name         = "daily-35d"
    target_vault_name = aws_backup_vault.primary.name
    schedule          = "cron(0 5 * * ? *)"  # 05:00 UTC daily
    start_window      = 60
    completion_window = 180
    lifecycle { delete_after = 35 }

    copy_action {                            # the DR-critical part
      destination_vault_arn = aws_backup_vault.dr.arn   # different Region + account
      lifecycle { delete_after = 35 }
    }
  }
}

Terraform shape for the Aurora Global Database data plane (the Pilot Light / Warm Standby relational tier):

resource "aws_rds_global_cluster" "this" {
  global_cluster_identifier = "ord-global"
  engine                    = "aurora-postgresql"
  engine_version            = "16.4"
  storage_encrypted         = true
}

resource "aws_rds_cluster" "primary" {          # writer — us-east-1
  provider                    = aws.use1
  cluster_identifier          = "ord-use1"
  engine                      = aws_rds_global_cluster.this.engine
  engine_version              = aws_rds_global_cluster.this.engine_version
  global_cluster_identifier   = aws_rds_global_cluster.this.id
  master_username             = var.db_user
  manage_master_user_password = true            # secret in Secrets Manager, not state
  kms_key_id                  = aws_kms_key.use1.arn
  db_subnet_group_name        = module.region_use1.db_subnet_group
}

resource "aws_rds_cluster" "secondary" {        # promotable reader — us-west-2
  provider                  = aws.usw2
  cluster_identifier        = "ord-usw2"
  engine                    = aws_rds_global_cluster.this.engine
  engine_version            = aws_rds_global_cluster.this.engine_version
  global_cluster_identifier = aws_rds_global_cluster.this.id
  kms_key_id                = aws_kms_key.usw2.arn
  db_subnet_group_name      = module.region_usw2.db_subnet_group
  depends_on                = [aws_rds_cluster.primary]
}

And the Region B compute as a single variable away from Pilot Light vs Warm Standby:

module "region_usw2" {
  source    = "./modules/region"
  providers = { aws = aws.usw2 }

  # Pilot Light: desired_count = 0  (scaffolding only, scaled to zero)
  # Warm Standby: desired_count = 2 (minimal live fleet, ready to scale)
  ecs_desired_count = var.dr_warm ? 2 : 0
  ecs_max_count     = 40            # full production ceiling on failover
}

Networking and identity. Keep request handling in-Region; only data replication should cross Regions, and it travels the AWS backbone natively for Aurora, DynamoDB, and S3 — you do not need a hot-path VPC peering for the user flow. Put Gateway VPC endpoints for S3 and DynamoDB and Interface endpoints for the rest in each Region so data-plane traffic stays off NAT and the internet. Use multi-Region KMS keys so an encrypted snapshot or replicated object decrypts in Region B under the local key replica — a single-Region key is a silent way to make your “recovered” data unreadable. Put DB credentials in Secrets Manager replica secrets so Region B never makes a cross-Region Secrets Manager call on the recovery path. For human and machine identity, one AWS Organization with IAM Identity Center for SSO and per-Region IAM roles scoped to that Region’s resource ARNs (IRSA on EKS) — a compromised task in Region A should have no standing path to Region B beyond what replication already grants.

The failover runbook itself must be code, not prose. Whatever the strategy, the switch should be an executable sequence — a Systems Manager Automation document or a Step Functions state machine — that: (1) confirms the standby’s readiness (ARC readiness check), (2) promotes the data tier (failover-global-cluster for a planned switch, or promote-on-loss for unplanned), (3) scales Region B compute to production size, (4) flips the ARC routing control (and/or Global Accelerator traffic dial) to send traffic to Region B, and (5) verifies synthetic transactions succeed before declaring victory. Backup & Restore adds a step zero: terraform apply the Region B environment and restore the latest copied backups. Encoding this is what collapses RTO from “however long the on-call figures it out” to a predictable number.

Enterprise considerations

Security and Zero Trust. DR widens your attack surface — there is now a second copy of everything — so the recovery Region must be held to the same standard, not a relaxed one. Encrypt every backup and replica with multi-Region KMS; keep the immutable backup copy in a separate, least-privilege account so the credentials that run production cannot delete your last line of defence. Treat ransomware as a first-class DR scenario: your live cross-Region replication will faithfully copy encrypted data to the other Region, so the only recovery is the immutable, point-in-time backup — design retention and Vault Lock accordingly, and rehearse a restore from the locked vault specifically. Enable GuardDuty, Security Hub, and an organization-wide multi-Region CloudTrail so detection is symmetric across both Regions. Scope IAM per Region; a breach in the primary should not hand the attacker the recovery Region for free.

Cost optimization. This is where DR strategy selection literally is the cost decision, because the four strategies are a price ladder and the steady-state spend is dominated by what you keep running while nothing is wrong:

Strategy Steady-state cost driver Rough relative cost What you’re paying for
Backup & Restore Snapshot storage + cross-Region copy transfer $ (lowest) Just durable, replicated backups; zero idle compute
Pilot Light Replicated DB storage + transfer; minimal scaffolding $$ Live data plane; compute scaled to ~zero
Warm Standby Above + a small always-running compute fleet $$$ A real (if minimal) second environment, always on
Multi-Site Active-Active A near-full second environment + cross-Region transfer $$$$ Two live Regions; capacity for either to take 100%

The discipline is matching the strategy to the cost of downtime per workload, not to anxiety. A tier-3 internal tool on active-active is pure waste; a payment rail on backup-and-restore is negligence. Concrete levers: don’t replicate data that doesn’t need it (DynamoDB Global Tables charge replicated write capacity, Aurora Global DB charges cross-Region transfer — keep purely-regional data single-Region); run the Region B Aurora secondary smaller and scale it up as a step in the failover runbook if your RTO budget allows the extra minute; use AWS Backup lifecycle to cold storage for long-retention copies; and for Warm Standby, size Region B to the minimum that can survive the first few minutes while Auto Scaling ramps, not to full production.

Scalability. The scaling question in DR is specifically “can Region B actually absorb production load when it has to?” — and the failure mode is a Pilot Light or Warm Standby that looks ready but cannot scale fast enough, hitting service quotas (Region-specific limits on EC2 vCPUs, Elastic IPs, Lambda concurrency) or cold-start cliffs at the worst moment. Pre-raise quotas in Region B to production levels now, not during the incident. For Pilot Light especially, validate that scale-from-zero actually reaches capacity within your RTO — a fleet that takes 20 minutes to warm up turns a “10-minute RTO” into fiction.

Reliability and DR (RTO/RPO). This is the headline the whole article serves — the explicit mapping:

Strategy RPO (data loss) RTO (time to recover) How it’s achieved
Backup & Restore Hours (since last backup/copy) Hours (build + restore) AWS Backup cross-Region copies; IaC rebuild; snapshot restore
Pilot Light Minutes (replication lag) Tens of minutes (promote DB + scale compute from ~0) Live data replication; scaffolding ready; cold compute warms on failover
Warm Standby Seconds (replication lag) Minutes (promote DB + scale up an already-running fleet) Live data + minimal live compute; scale, don’t cold-start
Multi-Site Active-Active Near-zero Near-zero (capacity event, not failover) Both Regions hot; multi-active data; promote-only relational writer

Two practices make these numbers real rather than aspirational. First: rehearse on a schedule. Run Aurora managed planned failover and an ARC routing-control flip in a monthly/quarterly GameDay; launch DRS drills into an isolated subnet without disrupting production. A failover path you have never executed is an RTO you cannot honestly claim. Second: watch the leading indicators. Rising Aurora AuroraGlobalDBRPOLag or DynamoDB ReplicationLatency means your RPO promise is silently degrading before any outage — alarm on them. And distinguish planned (clean, replication caught up, RPO ~0) from unplanned (region truly lost, RPO = whatever was in-flight) failover in your runbooks and your SLA claims; they are different numbers.

Observability. Emit metrics, logs, and traces per Region (CloudWatch, X-Ray/OpenTelemetry) and aggregate into a single pane (cross-account/cross-Region CloudWatch dashboards or Datadog/Grafana). The DR-specific signals to watch: replication lag (Aurora RPO lag, DynamoDB replication latency, S3 CRR metrics), backup job success/failure and copy-job completion (a silently-failing cross-Region copy is a DR outage you discover at the worst time), ARC readiness-check status, Route 53 health-check status, and recovery-Region service-quota headroom. Build a DR readiness dashboard that answers “if we had to fail over in the next five minutes, would it work?” — and put backup-copy failures and readiness-check regressions on the on-call pager, because they are the failures that bite you precisely when you reach for the parachute.

Governance. Enforce with Service Control Policies (deny resource creation outside sanctioned Regions to stop shadow expansion; deny deletion of backup vaults), AWS Config conformance packs verifying that critical resources are actually covered by a backup plan and that replication is enabled, and a data-residency tagging scheme so a future engineer cannot accidentally replicate EU-resident personal data into a non-permitted recovery Region — a genuine compliance trap in any cross-Region design. Maintain a per-workload DR register: each workload’s tier, its agreed RTO/RPO, its chosen strategy, the date of its last successful failover test, and the owner who signed off the objectives. That register is your audit evidence and your honesty check in one document.

Reference enterprise example

Meridian Logistics, a fictional mid-market freight and supply-chain platform, runs three customer-facing systems in us-east-1 for ~450 enterprise shippers: a shipment-tracking API and portal, an order/billing system (the financial system of record), and an analytics & reporting platform. After a six-hour us-east-1 AZ event cost them a day of degraded service and a near-miss on a major customer’s renewal, their board mandated a DR program — but their CFO refused to “build everything active-active” after seeing the quote. The CTO’s mandate became the right one: tier each system and spend per tier.

What they decided (one strategy per workload):

System Criticality Agreed RTO / RPO Strategy chosen Why
Order / Billing Tier-1 (financial SoR) RTO 15 min / RPO ~5 s Warm Standby Money can’t be lost or down for long; needs a fast, rehearsable failover but not full active-active
Shipment Tracking API/Portal Tier-2 (customer-facing) RTO 30 min / RPO 5 min Pilot Light Customers tolerate a short gap; data must be current; idle compute is wasteful
Analytics / Reporting Tier-3 (internal + batch) RTO 8 h / RPO 24 h Backup & Restore Re-runs nightly; a day-old report is fine; cheapest is correct here

How each was built:

Numbers:

The test that mattered: in their first quarterly GameDay they ran a managed planned Aurora failover for the billing system from us-east-1 to us-west-2 in a low-traffic window. Writes were serving from us-west-2 in ~75 seconds; the warm fleet scaled to full in ~3 minutes; the ARC routing control flipped traffic deterministically; synthetic billing transactions passed — total measured RTO under 6 minutes, comfortably inside the 15-minute commitment. The tracking-system Pilot Light test came in at ~22 minutes (the scale-from-zero compile being the long pole), inside its 30-minute target. And restoring an analytics snapshot from the locked vault into us-west-2 took ~3 hours, well inside the 8-hour budget.

One scar they earned: their first Pilot Light test for tracking failed — Region B hit the default EC2 vCPU service quota while scaling from zero and stalled at half capacity, blowing the RTO. The fix was to pre-raise service quotas in us-west-2 to production levels as a standing item in the DR register, and to add an ARC readiness check that flags quota headroom. It is the canonical Pilot Light lesson: a standby that exists is not a standby that can scale — and you find that out in a drill or in a disaster, so make it the drill.

When to use it

The decision is never “should we do DR” — it is “which strategy, for which workload, at what cost.” The clean rule: start from the agreed RTO/RPO and walk up the ladder only as far as those numbers force you.

Choose Backup & Restore when:

Choose Pilot Light when:

Choose Warm Standby when:

Choose Multi-Site Active-Active when:

Anti-patterns to avoid:

Alternatives and adjacent choices:

The honest framing for any DR review: the RTO/RPO numbers are a business decision and a price tag, not a technical preference. Get those numbers signed off per workload, build the cheapest strategy that meets them, encode the failover so it is deterministic, and prove it on a schedule. A DR plan you have tested at the lowest tier that meets your objectives beats a gold-plated one you have never run — because the parachute you have actually deployed is worth more than the one you only bought.

Going deeper

Everything above is the architecture a senior engineer signs off. This section is the set of things that bite you after you have signed — the internals, the failure modes, and the version caveats that separate a DR plan that works in the drill from one that works in the disaster.

The RPO you get is not the RPO you think you have

RPO in a replication-based tier is a live measurement, not a design constant. Replication lag is quiet at rest and spikes exactly when you least want it: a write burst, a large UPDATE, an index rebuild, a cross-Region network wobble, or an undersized secondary all push it up. Your unplanned RPO — the number that matters — is whatever the lag is at the instant the primary is lost, because those in-flight, un-replicated writes are the ones you never get back. This is why “we have Aurora Global Database, so RPO is ~1 second” is a claim with an expiry date; it is only true while the lag metric says so.

Signal Service What it tells you Turn it into
AuroraGlobalDBRPOLag / AuroraGlobalDBReplicationLag Aurora Global Database How many ms/s of writes the secondary is behind A CloudWatch alarm; a paging alert if it exceeds your RPO budget
ReplicationLatency DynamoDB Global Tables Per-Region replication delay Same — trend it, do not sample it once
CRR bytes/operations pending; RTC 15-min SLA S3 Cross-Region Replication Whether objects are actually landing in Region B Enable Replication Time Control for a contractual 15-min replication SLA + metrics
Backup copy-job success/failure AWS Backup Whether the cross-Region/-account copy actually completed A pager alarm — a silently failing copy is a DR outage you discover at 3am

The distinction to internalise: planned vs unplanned failover are different RPO numbers. A managed planned failover (healthy primary, replication caught up, you initiate it) drains in-flight writes and reaches RPO ~0. An unplanned failover (the Region is genuinely gone) keeps only what had replicated. Your runbooks, dashboards, and SLA language must name which one they mean.

Aurora Global Database internals and its two failover modes

Aurora Global Database replicates at the storage layer: the primary’s distributed storage volume ships changes to the secondary Region’s storage, independent of the database engine’s own CPU, which is why lag is typically sub-second and why replication does not steal capacity from your writers. The secondary is read-only until promoted, and there are two very different ways to promote it — do not confuse them:

Two more caveats. Write forwarding lets a secondary accept writes and forward them to the primary — convenient for read-mostly apps that occasionally write, but it adds cross-Region latency and is not a substitute for a local writer; it does not change who the single writer is. And engine versions must match across the global cluster: upgrades are choreographed across Regions, so a Global Database constrains your patch cadence in a way a standalone cluster does not.

DynamoDB Global Tables: multi-active, but last-writer-wins

Global Tables (use v2, the 2019.11.21 version) are genuinely multi-active — every Region takes local writes and replicates to the others, which is why the tiers that use DynamoDB often have zero database failover to perform. The catch is the conflict rule: last-writer-wins by wall-clock timestamp, per item. If two Regions write the same item concurrently, the write with the later timestamp wins and the other is silently discarded — no error, no merge. That makes Global Tables superb for Region-owned or idempotent data (session state keyed by Region, idempotency keys, append-mostly event records) and dangerous for a shared strongly-consistent counter or a bank balance updated from both Regions. Design for commutative or Region-partitioned writes. PITR (point-in-time recovery) is configured per Region, and replicated writes are billed as replicated write capacity units (rWCU) — roughly, you pay for the write again in every Region it lands in.

Route 53 ARC internals, and why it beats bare health-check failover

Application Recovery Controller’s whole value is determinism during chaos. Its control lives in a cluster of five Regional endpoints; you flip a routing control by talking to any one of them, so you can still steer traffic even when one Region — possibly the one you are failing away from — is impaired. Routing controls are protected by safety rules: assertion rules (“routing control X must be ON”) and gating rules (“never turn both Regions OFF”, “at least one must stay ON”) that stop a panicked operator or a buggy automation from black-holing all traffic. Readiness checks continuously compare the standby’s capacity and configuration against the primary and answer the only question that matters right before you commit: is Region B genuinely able to take this load right now? ARC also offers zonal shift / zonal autoshift — evacuating a single impaired Availability Zone, a finer granularity than a whole-Region failover, and often the first thing to reach for in an AZ-scoped event. ARC clusters are billed hourly, so they are a real line item, justified for tier-1 systems by the determinism they buy.

Bare Route 53 health-check failover fails in the messy middle: a partial brownout where the primary is sick enough to lose data but healthy enough to pass a shallow /health probe, or health checks that flap and bounce users back into a degraded Region. ARC replaces “hope the health check fires correctly” with “an operator or runbook flips a switch, and safety rules forbid the dangerous states.”

Failover speed: DNS is TTL-bound; Global Accelerator is not

Even with a 30–60 s TTL on your Route 53 records, real-world DNS failover is minutes-slow because many resolvers and clients ignore low TTLs and cache aggressively. For non-cacheable, latency-critical APIs, AWS Global Accelerator sidesteps this: it hands clients two static anycast IPs that never change, and fails over by moving the endpoint set at the AWS network edge — clients keep the same IP and are redirected in seconds, with traffic dials and endpoint weights for gradual shifts. You do not need it for browser apps already behind CloudFront (the edge handles origin failover there); you reach for it precisely on the raw API paths where DNS caching would otherwise be your long pole.

DRS mechanics: continuous block replication and a cheap staging area

For lift-and-shift estates you will never re-platform, AWS Elastic Disaster Recovery (DRS) — the successor to CloudEndure Disaster Recovery — is the pragmatic Pilot-Light/Warm-Standby-in-a-box. A lightweight agent on each source server streams continuous, asynchronous, block-level changes to a low-cost staging area in the recovery Region: cheap instances plus low-cost EBS, so you hold a current copy of whole servers for a fraction of running them. Two features make it real DR rather than a fancy copy: point-in-time recovery (recover to a moment before a ransomware or corruption event, since plain replication would have copied the damage) and non-disruptive drills that launch recovered instances into an isolated subnet from a chosen point-in-time without stopping replication or touching production. On a real failover, launch settings/templates control the production-sized shape the recovered instances take. See configuring DRS cross-Region failover for the hands-on build.

Ransomware is an integrity problem, and replication cannot solve it

The single most important distinction in this whole topic: live cross-Region replication is an availability control, not an integrity control. It copies a DROP TABLE, a bad migration, or a ransomware encryption to the other Region in seconds, exactly as designed. The only defence against that class of disaster is a copy that can travel back in time and that the attacker (or the mistake) cannot alter:

Correlated failure and the control-plane trap

Choose a recovery Region far enough that a shared power grid, natural disaster, or fibre cut cannot take both — but then remember that some AWS control-plane operations have a home Region (historically us-east-1 for parts of IAM, Route 53’s control plane, CloudFront, and the global STS endpoint). If your failover depends on creating or mutating global resources during the incident, a control-plane brownout in that home Region can block you in the exact event you are recovering from. The design response is to lean on data-plane operations that are already provisioned and independent: ARC’s highly-available data plane to flip routing controls, Global Accelerator’s edge, pre-created resources, and regional STS endpoints (sts.<region>.amazonaws.com) instead of the global one. The best failover path performs no control-plane creation on the critical path — everything it needs already exists, and it only flips switches.

Testing depth: drills validate the mechanism, GameDays validate the numbers

A drill is non-disruptive and isolated (an Aurora managed planned failover in a quiet window, a DRS launch into a sandbox subnet, an ARC readiness check) — it proves the mechanism works. A GameDay is a planned exercise against real (low) traffic that proves the organisation and the numbers work: does the on-call know the runbook, does the measured RTO match the claim, do the synthetics pass? AWS Fault Injection Service (FIS) lets you script the failure rather than wait for it — AZ-level impairment experiments and integration with ARC zonal shift turn “we think it fails over” into “we watched it fail over.” Whatever you measure, write it into the per-workload DR register next to the claimed RTO/RPO and diff them; the day those two diverge is the day your SLA became fiction, and you want to learn it in a GameDay, not from a customer.

Cost internals worth knowing before the bill arrives

The steady-state premium of each tier hides in a few specific meters: cross-Region data transfer is billed per GB (it is not free the way some intra-AZ traffic is); Aurora Global Database adds replicated write I/O plus that cross-Region transfer; DynamoDB Global Tables bill replicated writes as rWCU in every Region; S3 CRR means paying for storage in both Regions plus transfer plus replication PUT requests; ARC clusters bill hourly; and DRS staging is deliberately cheap, but a real recovery launch spins full-price production instances. Warm Standby’s dominant cost is simply the always-on minimal fleet — right-size it to “survive the first few minutes while Auto Scaling ramps,” not to full production, and let the runbook do the scaling.

Practice challenges

Work these in order — they climb from “read the numbers” to “design the failover.” Try each before opening the solution.

1. (Beginner) Read an RPO off a schedule. A workload is protected only by AWS Backup, taking one snapshot nightly at 02:00 UTC and copying it cross-Region by 03:00. The primary Region is lost at 14:30 UTC. What is the worst-case RPO, and which single change most reduces it?

<details><summary>Solution</summary>

Worst-case RPO ≈ ~12.5 hours — the time since the last restorable, copied backup (the 02:00 snapshot). All writes since 02:00 are gone. The single most effective change is to shorten the backup + copy interval (e.g. every 4 hours) — but you can never reach seconds with snapshots; that needs live replication (Pilot Light or higher).

Why: For backup-based tiers, RPO ≈ time since the last successful copied backup, bounded below by the backup interval. </details>

2. (Beginner) Match strategy to band. Without looking back, match each strategy to its rough RTO/RPO band: hours/hours, tens-of-minutes/minutes, minutes/seconds, near-zero/near-zero.

<details><summary>Solution</summary>

Why: The four strategies are one axis — each rung pre-provisions more warm infrastructure, buying tighter RTO/RPO for more steady-state cost. </details>

3. (Intermediate) Pick a tier from the numbers. A PostgreSQL-on-RDS (non-Aurora) workload has an agreed RTO of 20 minutes and RPO of 30 seconds. Which strategy, and which replication mechanism, and what must you validate?

<details><summary>Solution</summary>

RPO of 30 s rules out backups → you need live replication. Options: an RDS cross-Region read replica (async, promote on failover — RPO = replication lag, which must stay comfortably under 30 s) or migrate to Aurora Global Database for tighter, managed failover. RTO of 20 min fits Pilot Light if you can prove scale-from-zero completes in time; if that is tight, use Warm Standby (already-running minimal fleet). Validate: (a) the read-replica lag stays under 30 s under peak write load, and (b) promote-plus-scale actually completes inside 20 minutes in a drill.

Why: Seconds-grade RPO demands replication, not snapshots; the RTO choice hinges on whether scale-from-zero fits — and you measure that, you do not assume it. </details>

4. (Intermediate) Harden the backup against ransomware. You already copy backups cross-Region within one account. Explain what is still missing for a genuine ransomware defence, and sketch the two Terraform changes.

<details><summary>Solution</summary>

Same-account copies can be deleted by the same compromised credentials, so add (a) cross-account copy into an isolated backup account, and (b) Vault Lock in compliance mode so retention cannot be shortened even by root:

resource "aws_backup_vault_lock_configuration" "dr" {
  provider           = aws.usw2
  backup_vault_name  = aws_backup_vault.dr.name
  min_retention_days = 35
  changeable_for_days = 3     # cooling-off, then the lock is permanent
}

# copy_action destination points at a vault in a DIFFERENT account + Region
copy_action {
  destination_vault_arn = "arn:aws:backup:us-west-2:210987654321:backup-vault:dr-locked-vault"
}

Why: Ransomware/insiders delete same-account snapshots; cross-account + compliance-mode immutability turns “we have snapshots” into “we can actually recover.” </details>

5. (Advanced) Order the failover runbook. For a Warm Standby workload, list the executable steps of the failover in the correct order and explain why the order matters.

<details><summary>Solution</summary>

  1. Confirm standby readiness (ARC readiness check + replication lag within budget).
  2. Promote the data tier (aws rds failover-global-cluster for a planned switch, or detach-and-promote for a true loss).
  3. Scale Region B compute up to production size (the fleet is already running, so this is scale-up, not cold-start).
  4. Flip the ARC routing control (and/or Global Accelerator traffic dial) to send traffic to Region B.
  5. Verify synthetic transactions succeed before declaring the failover complete.

Order matters because you must not shift traffic (step 4) before the database is writable (step 2) and the fleet can serve (step 3) — do it early and you route live users straight into errors. Readiness (step 1) gates the whole thing; verification (step 5) is what lets you claim success rather than hope it.

Why: Failover is a dependency chain — data-writable → compute-ready → traffic-shifted → verified; reordering it turns a recovery into a second outage. </details>

6. (Advanced) Diagnose a creeping RPO. Over a week with no incident, your AuroraGlobalDBRPOLag trends from 200 ms to 4 s. Nothing is “down.” What is actually happening, what would break, and what do you do?

<details><summary>Solution</summary>

Your unplanned RPO is silently degrading: if the primary Region were lost now, you would lose ~4 s of writes, not 0.2 s — your SLA promise is quietly becoming false with no outage to warn you. Likely causes: rising write throughput, an undersized secondary, cross-Region network pressure, or large transactions/index builds on the primary. Actions: alarm on the metric (page if it exceeds the RPO budget), scale the secondary, investigate the write pattern, and re-baseline the SLA to the number you can actually meet.

Why: For replication-based tiers, RPO ≈ current lag — the lag metric is your early-warning system before a real failover exposes the gap. </details>

Common beginner mistakes

These are mental-model errors — the intuitions that read as obviously true and quietly lead you to a DR plan that fails in the disaster. (They are distinct from the architectural anti-patterns in When to use it; these are the misconceptions that get you there.)

“RTO and RPO are basically the same number.” They are independent axes. RTO is time — how long until you are serving again. RPO is data — how much recent work you lose. A system can have a great RTO and a terrible RPO (fails over in 60 seconds, but to a day-old backup) or the reverse. Right model: set both numbers separately for each workload, because different strategies and services move them independently — replication drives RPO, warm compute drives RTO.

“We’re Multi-AZ, so we’re covered for a disaster.” Multi-AZ is high availability within one Region — it survives an Availability Zone or a single-instance failure automatically, and you should absolutely have it. It does nothing for a whole-Region outage, a region-wide bad config change, or ransomware. Right model: Multi-AZ handles the small, uncorrelated failures; DR is a second Region for the large, correlated ones. They are different tools solving different problems, and you need both.

“A cross-Region read replica (or a Global Table) is our backup.” Replication is an availability control, not an integrity one — it copies your mistakes at wire speed. A DROP TABLE, a bad migration, or an encryption event replicates to Region B in seconds, corruption and all. Right model: replicas protect against losing the infrastructure; only immutable, point-in-time backups (Vault Lock, PITR) protect against corrupting the data. You need both, and they are not interchangeable.

“Tighter RTO/RPO is always better, so let’s build the best tier.” Every rung down the ladder costs more every single day, for an event that may never come — a hot standby bills the same on a quiet Tuesday as during a disaster. Gold-plating a low-criticality system is how DR budgets get wasted and how the crown jewels end up under-funded. Right model: pick the cheapest tier that meets the business-agreed number, per workload — active-active for the payment rail, backup-and-restore for the internal wiki.

“We wrote the runbook, so we’re ready.” A runbook you have never executed is a hypothesis, not a capability — the first real failover should never be the first failover. Untested plans fail on the boring things: a stale IAM permission, a service quota, a single-Region KMS key, a script that assumes the primary is reachable. Right model: the drill is the deliverable. Rehearse on a schedule (Aurora planned failover, DRS drills, ARC flips, GameDays), measure the real RTO/RPO, and record it — an unrehearsed RTO is a number you cannot honestly claim.

“The data is replicated to Region B, so our app runs there.” Data is only one of the ingredients. Without pre-provisioned or one-command-deployable compute, networking (VPC/subnets/security groups/ALB), DNS failover, multi-Region KMS keys, and replica secrets in Region B, your perfectly-replicated database is unreachable. Right model: a recovery Region needs the whole environment ready (or rebuildable by IaC), not just the bytes — “recovered data” you cannot connect to, decrypt, or route traffic to is not a recovery.

Glossary

AWSArchitectureEnterpriseReference Architecture
Need this built for real?

Vinod is a Senior Cloud Architect (22+ yrs) — available for Azure / AWS / GCP architecture, landing zones, and migrations.

Work with me

Comments