AWS Lesson 36 of 123

Amazon Route 53, In Depth: Hosted Zones, Records, Routing Policies & Health Checks

In a nutshell

Picture the internet as a vast city where every server is a building with a numeric street address — its IP address. Nobody memorises street addresses; we remember names (“the coffee shop on Main Street”). DNS (the Domain Name System) is the city-wide directory that turns a name you can remember into the address a computer actually needs, and Amazon Route 53 is AWS’s version of that directory — with a smart traffic controller bolted on top.

Route 53 wears three hats, and beginners often blur them. It is a registrar (you can buy a name like example.com from it), it is the authoritative directory for that name (the single source of truth that says “example.com lives at this address”), and it is a traffic director that can hand back different addresses to different people — the nearest datacentre, a healthy server instead of a broken one, 5% of users to a new version for testing. That third hat is what lifts Route 53 above a plain phone book.

Why should a beginner care? Because Route 53 is the first thing every request touches, before any server you pay for. A small mistake here — the wrong record type, or a cache told to last a day when you need to fail over in a minute — becomes an outage that the rest of your perfectly healthy architecture cannot fix. Learn DNS and you can build a front door that is fast, resilient, and fails over in seconds.

The lesson goes deep, but you can read it in layers: the plain-language core first, then the Going deeper and Practice challenges sections when you want to push further.

Level: Intermediate · Time: ~52 min · You need: basic AWS Console/CLI comfort and a rough idea of what a load balancer does; every DNS concept itself is built up from scratch here.

Every connection on the internet begins with a question: what is the IP address for this name? The Domain Name System (DNS) answers it, and on AWS that answer is served by Amazon Route 53 — a highly available, globally distributed authoritative DNS service named after the port DNS runs on, port 53. Route 53 does three distinct jobs that people often conflate: it registers domain names (you can buy example.com through it), it hosts the authoritative records for a domain (it is the source of truth resolvers ask), and it routes traffic intelligently using policies that go far beyond plain DNS — weighting answers for canary releases, sending users to the lowest-latency Region, failing over to a backup when a health check goes red, and answering differently based on the user’s geography.

Route 53 sits at the very front of almost every architecture. It is the first thing a client touches, before the load balancer, before CloudFront, before any compute you pay for. Get it right and you have a resilient, fast-resolving front door with sub-second failover. Get it wrong — a stray CNAME at the apex, a TTL of 86,400 seconds on a record you need to fail over, a health check pointed at the wrong port — and you have an outage that DNS caches will keep serving for hours. This lesson is the exhaustive version: every record type, the Alias-versus-CNAME distinction that trips up nearly everyone, all seven routing policies with explicit when-to-use guidance, the three kinds of health check, and how TTL governs how fast the world sees your changes.

Learning objectives

By the end of this lesson you will be able to:

Prerequisites & where this fits

You should be comfortable with the AWS Console and the AWS CLI (see AWS Hands-On First Steps: Console, CLI, CloudShell, SDKs & Access Keys), and you should have met load balancers in AWS Elastic Load Balancing, In Depth: ALB, NLB, GWLB & Target Groups — Route 53 most often points at an ALB or CloudFront distribution. A working mental model of how a domain delegates to name servers helps but is not assumed; we build it here. This lesson belongs in the Networking module of the AWS Zero-to-Hero course, immediately after load balancing, because the natural progression is: distribute traffic within a Region (ELB), then distribute and fail over traffic across Regions and endpoints (Route 53). It also feeds directly into the edge and DNS-security lessons cross-linked at the end.

Core concepts

Before the settings, the mental model. DNS is a hierarchical, distributed database that maps human-readable domain names to machine-usable data, most commonly IP addresses. The hierarchy reads right-to-left: the root (.), then the top-level domain (TLD — .com, .org, .co.uk), then your registered domain (example.com), then any subdomains (api.example.com).

When a user’s machine needs to resolve api.example.com, it does not ask Route 53 directly. It asks a recursive resolver (typically run by the user’s ISP, or a public one like 8.8.8.8). That resolver walks the hierarchy: it asks a root server “who handles .com?”, asks the .com TLD servers “who is authoritative for example.com?”, and they answer with the name servers for your domain — which, if you host the zone in Route 53, are four Route 53 name servers. The resolver then asks one of those Route 53 name servers for api.example.com, gets the record, caches it for the TTL, and returns it to the user. Route 53 is the authoritative server at the end of that chain — it holds the truth. Everything else is asking and caching.

A few load-bearing terms:

With that, the settings.

Hosted zones: every setting

A hosted zone is where you start. There are two kinds, and the difference is who can see the records.

Public hosted zones

A public hosted zone is authoritative on the public internet. Any resolver in the world can query it. You create one for a domain you intend to serve publicly (example.com). On creation Route 53 assigns four name servers (an NS record) and a SOA record, both created automatically — never delete them. To make the zone live, you must delegate to those four name servers from the parent: if you registered the domain through Route 53, you point the registered domain at the zone’s name servers (Route 53 can do this automatically); if you registered elsewhere, you copy the four name servers into your registrar’s control panel.

Setting What it does Choices / default When to change / gotcha
Domain name The apex/root the zone is authoritative for Any DNS name; immutable after creation Must match the registered domain. Typo means nothing resolves — delete and recreate.
Type Public vs private Public (default) Public = internet-visible. Cannot convert a zone’s type after creation.
Comment Free-text description Optional Use it to record owner/ticket; editable later.
Name servers (NS) The four authoritative servers Route 53 assigns Auto-generated, 4 per zone Copy exactly these into your registrar. Two zones for the same name get different name servers — only the delegated set is live.
SOA Start-of-authority metadata (primary NS, contact, serial, timers) Auto-generated Rarely edited; the minimum TTL field affects negative-caching. Do not delete.

A subtle and very common mistake: deleting a public hosted zone and recreating it gives you a brand-new set of four name servers. The old delegation at the registrar now points at name servers that no longer host the zone, and the domain goes dark until you update the registrar. Treat hosted-zone name servers as something to wire once and leave alone.

Private hosted zones

A private hosted zone is authoritative only inside one or more VPCs you associate with it. Queries from those VPCs (via the Route 53 Resolver at the VPC .2 address) get the private answers; the rest of the internet cannot see the zone at all. This is how you give internal services friendly names — db.internal.example.com resolving to a private IP — without exposing topology publicly.

Setting What it does Choices / default When to change / gotcha
Domain name The zone the records cover Any name (often a real domain or an internal-only one) You can run a private zone for the same name as a public zone — this is split-horizon DNS (below).
VPCs to associate Which VPCs see these records One or more, any Region/account The VPC must have enableDnsHostnames and enableDnsSupport both on, or resolution fails silently. Cross-account association needs a CLI authorisation step.
Region Where the zone is created Any; association can span Regions Private zones are global objects but associations are per-VPC.

Split-horizon (split-view) DNS is the headline use case: associate a private zone for example.com with your VPCs and keep a public zone for example.com on the internet. Inside the VPC, app.example.com resolves to a private ALB; outside, the same name resolves to a public CloudFront distribution. Route 53 evaluates the most specific matching private zone first for queries from an associated VPC, falling back to public resolution only if no private zone covers the name. The classic gotcha: if a private zone for example.com exists but lacks a record for legacy.example.com, queries from the VPC get NXDOMAIN rather than falling through to the public zone — the private zone is authoritative for the whole name space it covers.

Record types: every type you will meet

A record (resource record set) has a name, a type, a TTL (except Alias), and a value. The type tells resolvers what kind of data to expect. Route 53 supports the full standard set; these are the ones you will actually configure.

Type What it holds Typical value Notes & gotchas
A IPv4 address 203.0.113.10 The workhorse. Can be an Alias (see below) instead of a literal IP.
AAAA IPv6 address 2001:db8::1 The IPv6 equivalent of A; also Alias-capable. Add it whenever you serve IPv6.
CNAME Canonical name (an alias to another name) lb-123.eu-west-1.elb.amazonaws.com Returns a name, not an IP; the resolver must look that up too. Forbidden at the zone apex and must be the only record at its name.
NS Name servers for a zone or delegated subdomain four ns-xxx.awsdns-xx.* Created automatically for the zone apex. Add your own NS records to delegate a subdomain to a different zone/provider.
SOA Start of authority — zone metadata and timers ns-... hostmaster... serial refresh retry expire minTTL One per zone, auto-created. The last field sets negative caching (how long NXDOMAIN is cached).
MX Mail exchanger — where email for the domain goes 10 mail.example.com The number is priority (lower = preferred). Required for receiving email.
TXT Arbitrary text "v=spf1 include:_spf.google.com ~all" Used for SPF, DKIM, DMARC, and domain-ownership verification. Quote each string; 255-char chunks.
SRV Service location (host + port + priority + weight) 10 60 5060 sip.example.com For protocols that advertise host and port (SIP, LDAP, some game and chat services).
CAA Which Certificate Authorities may issue certs for the domain 0 issue "amazon.com" A security control — stops a rogue CA issuing a cert for your domain. Add amazon.com so ACM can issue.
PTR Reverse DNS (IP → name) host.example.com Lives in special in-addr.arpa / ip6.arpa zones; used for reverse lookups and mail-server reputation.
NAPTR / DS / SPF / others Telephony rewriting, DNSSEC delegation signer, legacy SPF varies Less common; Route 53 supports them. SPF-the-type is deprecated — use TXT for SPF.

Two practical rules that catch people: a CNAME must be alone at its name (you cannot have a CNAME and an A for www.example.com), and you cannot put a CNAME at the apex (example.com itself), because the apex must also carry the NS and SOA records and the DNS spec forbids a CNAME coexisting with other records. The fix for the apex is the Alias record.

Alias vs CNAME: the distinction that trips everyone

This is the single most-asked Route 53 interview question, so understand it cold.

A CNAME is standard DNS: it says “this name is really that name; go look that up.” It works for any target, AWS or not, but it returns a name, forcing the resolver to do a second lookup, and it is forbidden at the apex and must be alone at its record name.

An Alias record is a Route 53 extension (not standard DNS). It points an A or AAAA record directly at a supported AWS resource — and at resolution time Route 53 substitutes that resource’s current IP address(es) into the answer. To the resolver it looks like a normal A/AAAA answer (it gets IPs, not a name), so there is no second lookup and no charge for Alias queries to AWS resources. Crucially, an Alias works at the zone apex, which is why example.com → CloudFront is always an Alias, never a CNAME.

Dimension CNAME Alias
Standard? Yes (RFC) No — Route 53 only
Returns A name (triggers another lookup) IP address(es) directly
Works at apex? No Yes
Can coexist with other records at the name? No (must be alone) Yes (it is an A/AAAA)
Targets Any DNS name Specific AWS resources + same-zone records (see below)
Query cost Charged as a normal query Free when pointing at an AWS resource
Health/failover integration Manual Evaluate Target Health auto-tracks the target
TTL You set it Inherited from the target (you cannot set it)

Alias targets you can point at: CloudFront distributions, ELB load balancers (ALB/NLB/CLB), S3 website endpoints, API Gateway, VPC interface endpoints, Elastic Beanstalk environments, Global Accelerator, AppSync, and — very usefully — another record in the same hosted zone. That last one lets you alias www.example.com to example.com and maintain the IP in one place.

The Evaluate Target Health toggle on an Alias is the quiet superpower: set it to Yes and Route 53 stops returning that Alias if the underlying resource (e.g. all targets behind an ALB) is unhealthy — health checking you get for free, without creating a separate health check, as long as you are aliasing an ELB, CloudFront, or another Route 53 record that is itself health-checked.

When to use which: Alias for anything pointing at an AWS resource (always — it is free, faster, and apex-capable); CNAME only for pointing a subdomain at a non-AWS name (a SaaS endpoint, a partner’s host) or where you genuinely need standard-DNS behaviour.

Routing policies: all seven, with when-to-use

A routing policy is set per record and decides which value Route 53 returns when several records share the same name and type. This is where Route 53 stops being plain DNS and becomes a traffic director. There are seven.

1. Simple

One record, one answer (or, if you give multiple values, Route 53 returns them all in random order and the client picks). No health checking. This is ordinary DNS.

2. Weighted

Multiple records, same name/type, each with a weight (0–255). Route 53 returns each in proportion to its weight ÷ total. Weight 0 takes a record out of rotation (unless all are 0, in which case all are returned equally).

3. Latency-based

Multiple records, each tagged with an AWS Region. Route 53 returns the record whose Region gives the lowest network latency to the resolver, based on AWS’s continuously-measured latency map.

4. Failover

A primary and a secondary record. Route 53 returns the primary while its health check is healthy, and switches to the secondary when the primary fails. The classic active-passive pattern.

5. Geolocation

Returns a different record based on the geographic location of the user (resolver), matched by continent, country, or — for the US — state. You can set a default record for locations that match no rule.

6. Geoproximity

Routes based on the geographic distance between the user and your resources, with a bias you can dial (–99 to +99) to expand or shrink the geographic area a resource serves. Configured via Route 53 Traffic Flow (the visual policy editor).

7. Multivalue answer

Returns up to eight healthy records chosen at random from a larger set, each optionally health-checked. It is like a simple record with multiple values plus health checking, giving you crude DNS-level load distribution that automatically omits unhealthy endpoints.

Policy Returns based on Health-check aware? Signature use case
Simple Single config No One resource, plain DNS
Weighted Assigned weights Yes (per record) Canary / blue-green, A/B
Latency Lowest measured latency Yes (per record) Active-active multi-Region for speed
Failover Primary health Yes (required) Active-passive DR
Geolocation User’s country/continent/state Yes (per record) Localisation, compliance, geo-block
Geoproximity Distance + bias Yes (per record) Distance routing with tunable boundaries
Multivalue Up to 8 random healthy values Yes (per record) Health-aware spreading without an LB

You can nest policies with Traffic Flow — e.g. latency-based at the top to pick a Region, then weighted within each Region for a canary, then failover under each weight. That composition is how large multi-Region systems are actually expressed.

Worked example: reading a routing decision end to end

Policies are easier to trust once you trace a real answer. Three short scenarios, each computed by hand.

Weighted — the maths behind a canary. You run two stacks behind the name api.example.com (type A): blue with weight 90 and green with weight 10, both TTL 30. For each incoming query Route 53 rolls a weighted die: it returns green with probability 10 ÷ (90 + 10) = 10% of queries, blue for the other 90%. The subtle part is queries, not users: a resolver that gets green caches it for the 30-second TTL, so every user behind that resolver sticks to green for half a minute. Over a large, busy population the proportions converge on 10/90; over a handful of users on one corporate resolver they can be lumpy. To pull green out of rotation without deleting its record, set its weight to 0 — Route 53 stops returning it (unless every record at that name is 0, in which case they are all returned equally). To ramp the canary, walk the weights 10→25→50→100 while the low TTL keeps the split tracking reality.

Nesting latency over failover. Now make it multi-Region and self-healing with a Traffic Flow tree: a latency rule at the top chooses between eu-west-1 and us-east-1, and under each Region a failover pair points primary at that Region’s ALB (Alias, Evaluate Target Health = Yes) and secondary at an S3 “maintenance” page. Trace a query from a London resolver: latency map → eu-west-1 → primary healthy → return the ALB’s IPs. Now every target behind the eu-west-1 ALB fails health checks: the primary leaf goes unhealthy, failover returns the secondary S3 page — while a New York user is still served the healthy us-east-1 primary. The order matters: latency is evaluated first to pick the leaf, then failover is resolved inside the chosen leaf. Nesting is how large systems compose “fast and self-healing” from two simple rules.

Geolocation — most specific wins. You publish three geolocation records for www.example.com: CountryCode=FR → fr-stack, ContinentCode=EU → eu-stack, and a Default → global. A user in France matches all three conceptually, but Route 53 picks the most specific match — state beats country beats continent beats Default — so France gets fr-stack. A user in Germany matches EU (not FR) → eu-stack. A user in Brazil matches none of the named rules → Default → global. Delete that Default record and the Brazilian user gets no answer at all (NODATA), which is the single most common geolocation outage.

Health checks: every kind

A health check is a separate Route 53 object that monitors a target and reports healthy/unhealthy; routing policies consult it to decide whether to return a record. Route 53 health checkers run from multiple AWS locations worldwide and a target is considered up if more than 18% of checkers see it as healthy (this is why you must allow the Route 53 health-checker IP ranges through firewalls). There are three types.

Endpoint health checks

Monitor an IP or domain name on a chosen protocol. The settings:

Setting What it does Choices / default When to change / gotcha
Protocol How to probe HTTP, HTTPS, TCP HTTP(S) lets you check a path and status; TCP only checks the port opens. HTTPS does not validate the certificate by default.
Endpoint What to probe IP address or domain name + port If you use a domain name, Route 53 resolves it each check. Use an IP to pin it.
Path (HTTP/S) Which URL to request e.g. /health; default / Point at a deep health endpoint that checks dependencies, not a static page that is “up” while the app is broken.
Request interval Probe frequency Standard 30 s or Fast 10 s Fast detects failure sooner but costs more and is noisier.
Failure threshold Consecutive fails before “unhealthy” 1–10, default 3 Lower = faster failover, more false positives on a blip. 3×30s ≈ 90s to flip.
String matching Require a string in the first 5,120 bytes of the response body Off / on with search string Catches “200 OK but wrong content” — e.g. require "OK" in the body.
Latency graphs Record response time in CloudWatch Off (default) / on Turn on to alarm on slow-but-up endpoints.
Invert health status Treat healthy as unhealthy and vice-versa Off (default) Niche — e.g. fail over to a site only when a maintenance flag returns 200.
Health checker regions Which checker locations probe Default set / custom Reduce to fewer regions to cut noise, but keep enough for the 18% rule.
SNI (HTTPS) Send the hostname in the TLS handshake On by default Required for endpoints that serve multiple certs on one IP.

Calculated health checks

A health check whose status is derived from other health checks using a Boolean rule — “healthy if at least N of these child checks are healthy”. It probes nothing itself.

CloudWatch-alarm health checks

A health check that mirrors the state of a CloudWatch alarm. The check is unhealthy when the alarm is in ALARM. This lets you health-check anything CloudWatch can measure — DynamoDB throttles, SQS queue depth, ELB 5xx rate, a custom metric — not just an HTTP endpoint.

Health checks integrate with routing in two ways: associate a health check with a record (failover, weighted, latency, geolocation, multivalue all honour it and stop returning unhealthy records), or use Evaluate Target Health on an Alias to inherit the target’s health automatically. A frequent design is a failover record pair where the primary is an Alias to an ALB with Evaluate Target Health = Yes — no manual health check object needed.

Worked example: how long a failover really takes, and the 18% rule

Two numbers turn “Route 53 fails over” into a real recovery-time figure you can defend in a design review.

Detection time = interval × threshold. An endpoint health check with the Standard 30 s interval and the default failure threshold of 3 needs three consecutive failed probes before it flips to unhealthy: roughly 90 seconds. Only then does Route 53 stop returning the record — but a client that already resolved the name still holds the old answer until its cached TTL expires. So the user-visible failover is approximately detection + record TTL: with a 60 s TTL that is ~90 s + up to 60 s ≈ 150 s worst case. Switch to the Fast 10 s interval with a threshold of 2 and detection drops to ~20 s (plus TTL) — faster recovery, but more probes, more cost, and more false positives on a transient blip. This is why failover-eligible records live on a 60 s TTL: the TTL, not the health check, is often the bigger half of your recovery budget.

The 18% rule and why firewalls matter. Route 53 runs health checkers from many AWS Regions at once, and an endpoint is considered healthy when more than 18% of those checkers report success. The practical consequence is a footgun: if your security group, NACL, or on-prem firewall blocks the Route 53 health-checker source ranges, every checker fails, you fall below 18%, and Route 53 marks a perfectly healthy endpoint down — draining live traffic. Allowlist the published ranges (service ROUTE53_HEALTHCHECKS in AWS’s ip-ranges.json), or point the check at a public path that is already open.

Composing a calculated check. Suppose “the service is up” means app AND database are healthy, and at least one of three cache nodes is healthy. You cannot express that in one flat rule, so you nest: build endpoint checks for app, db, cache1/2/3; a calculated check caches-ok that is healthy when ≥ 1 of {cache1, cache2, cache3}; then an outer calculated check healthy when 3 of {app, db, caches-ok} are healthy (N = M is a logical AND, N = 1 is OR). For CloudWatch-alarm checks, also decide the insufficient-data behaviour (healthy / unhealthy / last known state) — a metric that simply goes quiet must not silently trigger a Region-wide failover.

TTL: the propagation lever

TTL (time to live), in seconds, tells every resolver how long it may cache an answer before re-asking Route 53. It is the single biggest control over how fast a DNS change reaches users — and a constant trade-off.

Record purpose Sensible TTL Reasoning
Stable apex/www pointing at CloudFront (Alias) n/a — Alias TTL is managed Route 53 handles it; you can’t set it.
Records you may fail over 60 s Fast failover; the small extra query cost is worth the recovery time.
Stable MX, TXT (SPF/DKIM), NS 3,600–86,400 s Rarely change; cache hard.
A record you’re about to migrate lower it to 60 s a day before So the cut-over propagates quickly, then raise it again.

Two things people miss: Alias records to AWS resources have a TTL managed by Route 53 (you cannot set it), and negative answers (NXDOMAIN) are cached according to the minimum TTL in the SOA record, so a typo that returns NXDOMAIN can stick in caches even after you fix it.

Amazon Route 53: records, routing policies, health checks

The diagram traces a single query from a client through the recursive resolver to a Route 53 hosted zone, then shows the same name resolving differently under each routing policy and how health checks gate which records are returned.

Going deeper

The core lesson gives you a resilient front door. This section is for the reader who has to operate it — the eighth routing policy the record screen never shows you, the truth about “the user’s location,” hybrid resolution, DNS signing, and the recovery-orchestration layer that sits above plain failover.

The eighth policy: IP-based routing

The record-creation screen offers seven policies, but Route 53 has an eighth — IP-based routing — configured through CIDR collections rather than the policy dropdown. You upload a set of named CIDR locations (each a bunch of IP ranges, e.g. mobile-isp-A = 198.51.100.0/24, 203.0.113.0/24), then create records that each map to a location. When a query arrives, Route 53 looks at the source IP, finds the location whose CIDR block contains it, and returns that location’s record.

Why it exists: geolocation and latency both infer where to send a client, and both can be wrong for a network you actually know something about. IP-based routing lets you assert ground truth — “traffic from this ISP’s ranges goes to this peering point,” “this partner’s corporate CIDR gets the internal build” — which is common in telco, CDN, and large-enterprise designs. A sketch:

# 1) create a CIDR collection and add a named location
COLL=$(aws route53 create-cidr-collection --name known-networks \
  --caller-reference "cidr-$(date +%s)" --query 'Collection.Id' --output text)
aws route53 change-cidr-collection --id "$COLL" \
  --changes '[{"LocationName":"isp-A","Action":"PUT","CidrList":["198.51.100.0/24","203.0.113.0/24"]}]'
# 2) a record then references CidrRoutingConfig: {CollectionId, LocationName}
#    (with a second record using LocationName "*" as the catch-all default)

The same “always define a default” discipline from geolocation applies: give one record the location name * so an IP that matches no block still gets an answer.

“The user’s location” is usually the resolver’s — until EDNS Client Subnet

A caveat runs through latency, geolocation, geoproximity, and IP-based routing: Route 53 sees the recursive resolver’s IP, not the end user’s. A user on a VPN or a centralised corporate resolver can be geolocated hundreds of kilometres away. The mitigation is EDNS Client Subnet (ECS), an extension where a resolver forwards a truncated slice of the client’s subnet (e.g. 203.0.113.0/24, not the full address) alongside the query. Public resolvers such as Google and OpenDNS support ECS; Route 53 (and CloudFront) use it when present to target the real user’s network. You cannot force it — many resolvers strip ECS for privacy — so treat resolver-based routing as usually right, occasionally coarse, and never as a compliance control on its own. For hard data-residency, gate with geolocation and enforce residency in the application/data tier, not in DNS alone.

Route 53 Resolver: the in-VPC recursive service and hybrid DNS

Everything so far is Route 53 the authoritative service. Inside a VPC there is a second, separate thing: the Route 53 Resolver — the recursive resolver every instance already uses at the VPC’s base-of-range .2 address (also reachable at 169.254.169.253). It answers amazonaws.com, resolves your private hosted zones, and recurses to the public internet. Left alone it needs no configuration. Hybrid networks need two add-ons, each a set of ENIs with IPs in your subnets:

This is a large topic — outbound filtering with Resolver DNS Firewall, rule precedence, and full hybrid patterns get their own treatment in the Route 53 Resolver lesson. The one thing to fix now: hosted zones answer for names you own; the Resolver decides how a VPC resolves everything.

Two query-logging features people conflate

“Turn on Route 53 query logging” means two different things depending on which side you are on, and mixing them up loses you the logs you actually needed:

DNSSEC signing: keys, the chain of trust, and the two-phase dance

Plain DNS answers are unauthenticated — a resolver cannot tell a genuine reply from a spoofed one, which is what cache-poisoning attacks exploit. DNSSEC fixes this by cryptographically signing a zone so validating resolvers can verify each answer. Route 53’s model:

The operational trap is ordering, and getting it wrong takes the entire domain down for validating resolvers (they return SERVFAIL, not a stale answer). Enable: turn on signing first, confirm the zone is signed, then add the DS at the parent. Disable or migrate registrars: remove the DS at the parent first, wait out the DS record’s TTL so every cache forgets it, and only then disable signing or move. KSK rotation follows the same add-new-before-removing-old discipline, and Route 53 supports it. Treat DNSSEC as a two-person-rule change with monitoring on signature expiry.

Route 53 Application Recovery Controller (ARC) and zonal shift

Failover routing is reactive and per-record. For large multi-Region systems that must fail over deliberately and safely, AWS offers Route 53 Application Recovery Controller:

Zonal shift is ARC’s cheaper, narrower cousin and worth knowing even if you never buy a cluster. When one Availability Zone is impaired, you start a zonal shift on a supported load balancer and Route 53 stops handing out that AZ’s IP addresses, draining it in seconds; the shift auto-expires so you cannot forget to undo it. Zonal autoshift lets AWS perform the shift for you the moment it detects an AZ problem (you opt in, and AWS runs periodic practice shifts to prove your app tolerates losing an AZ). It is the fastest blast-radius control you have for a single-AZ event.

Traffic Flow, traffic policies, and geoproximity bias

The nested trees from the worked example are built in Traffic Flow, Route 53’s visual editor, which produces a versioned traffic policy — a document describing the whole routing tree (weighted under latency under failover, etc.). You do not attach a policy directly to a name; you create a policy record, which applies one version of the policy to a DNS name in a hosted zone. Versioning is the quiet benefit: roll a change forward by creating a new version, and roll back by re-pointing the policy record at the previous one. Two facts to budget for: geoproximity (and its bias, −99 to +99, that grows or shrinks a resource’s service area on the map) is only configurable through Traffic Flow or the API — never the simple record screen — and policy records cost around $50/month each, so express high-fan-out designs as one policy rather than dozens of hand-built records. A bias example: with resources in us-east-1 and us-west-2, dialling us-east-1 to +40 pushes the dividing line westward so more of the central US resolves east — a clean way to pre-drain a Region before maintenance without touching the app.

The registrar side: a domain is not a hosted zone

The two most-confused Route 53 objects are Registered domains and hosted zones — different resources with different lifecycles. Registering example.com through Route 53 auto-creates a public hosted zone and wires the domain’s NS records to it, which is why it feels like one thing; it is two. Consequences and features worth knowing:

Hands-on lab

We will create a hosted zone, add records, build a failover pair backed by a health check, and clean up. This uses Route 53 features that incur small charges (see the cost note); there is no perpetual free tier for hosted zones, but the cost of doing this for an hour is a few cents. You do not need to own a domain — we will create a zone and inspect it; you would only delegate a real domain at the registrar step.

1. Set a zone name and create a public hosted zone.

ZONE=kloudvin-lab-$RANDOM.example
aws route53 create-hosted-zone \
  --name "$ZONE" \
  --caller-reference "lab-$(date +%s)" \
  --hosted-zone-config Comment="Route53 deep-dive lab"

Expected output includes a HostedZone.Id like /hostedzone/Z0123456789ABCDEFG and a DelegationSet.NameServers list of four ns-*.awsdns-* servers. Capture the ID:

ZID=$(aws route53 list-hosted-zones-by-name --dns-name "$ZONE" \
  --query 'HostedZones[0].Id' --output text | sed 's#/hostedzone/##')
echo "$ZID"

2. View the auto-created NS and SOA records.

aws route53 list-resource-record-sets --hosted-zone-id "$ZID" \
  --query "ResourceRecordSets[?Type=='NS' || Type=='SOA'].[Name,Type]" --output table

You should see the apex NS (four name servers) and the SOA — both created for you.

3. Add a simple A record.

cat > /tmp/r53-simple.json <<JSON
{ "Changes": [ {
  "Action": "UPSERT",
  "ResourceRecordSet": {
    "Name": "www.$ZONE",
    "Type": "A",
    "TTL": 60,
    "ResourceRecords": [ { "Value": "203.0.113.10" } ]
  } } ] }
JSON
aws route53 change-resource-record-sets --hosted-zone-id "$ZID" \
  --change-batch file:///tmp/r53-simple.json

The response shows a ChangeInfo.Status of PENDING. Route 53 changes are atomic and usually INSYNC within seconds.

4. Create a health check (endpoint, HTTPS to a known-good host) and a failover pair.

HCID=$(aws route53 create-health-check \
  --caller-reference "hc-$(date +%s)" \
  --health-check-config 'Type=HTTPS,FullyQualifiedDomainName=aws.amazon.com,Port=443,RequestInterval=30,FailureThreshold=3,ResourcePath=/' \
  --query 'HealthCheck.Id' --output text)
echo "Health check: $HCID"

cat > /tmp/r53-failover.json <<JSON
{ "Changes": [
  { "Action": "UPSERT", "ResourceRecordSet": {
      "Name": "app.$ZONE", "Type": "A", "TTL": 60,
      "SetIdentifier": "primary",
      "Failover": "PRIMARY",
      "HealthCheckId": "$HCID",
      "ResourceRecords": [ { "Value": "203.0.113.20" } ] } },
  { "Action": "UPSERT", "ResourceRecordSet": {
      "Name": "app.$ZONE", "Type": "A", "TTL": 60,
      "SetIdentifier": "secondary",
      "Failover": "SECONDARY",
      "ResourceRecords": [ { "Value": "198.51.100.30" } ] } }
] }
JSON
aws route53 change-resource-record-sets --hosted-zone-id "$ZID" \
  --change-batch file:///tmp/r53-failover.json

5. Validate. Confirm both failover records exist and check the health-check status:

aws route53 list-resource-record-sets --hosted-zone-id "$ZID" \
  --query "ResourceRecordSets[?Name=='app.$ZONE.'].[SetIdentifier,Failover,HealthCheckId]" \
  --output table

aws route53 get-health-check-status --health-check-id "$HCID" \
  --query 'HealthCheckObservations[].StatusReport.Status' --output table

You should see a primary/PRIMARY record bound to your health-check ID, a secondary/SECONDARY record, and several checker locations reporting Success: HTTP Status Code 200, OK. Because the records use placeholder documentation IPs, do not expect a real dig against them to reach a server — the point is the Route 53 configuration and the health-check signal.

6. Cleanup. Delete the records, the health check, and the zone (a zone with non-default records cannot be deleted):

# delete the failover records (Action must be DELETE with exact current values)
sed 's/"UPSERT"/"DELETE"/g' /tmp/r53-failover.json > /tmp/r53-failover-del.json
aws route53 change-resource-record-sets --hosted-zone-id "$ZID" \
  --change-batch file:///tmp/r53-failover-del.json

sed 's/"UPSERT"/"DELETE"/g' /tmp/r53-simple.json > /tmp/r53-simple-del.json
aws route53 change-resource-record-sets --hosted-zone-id "$ZID" \
  --change-batch file:///tmp/r53-simple-del.json

aws route53 delete-health-check --health-check-id "$HCID"
aws route53 delete-hosted-zone --id "$ZID"

Cost note. A hosted zone costs USD 0.50 per month (pro-rated only for the first 12 hours, then charged per month — so create and delete on the same day to keep it to the half-dollar). Standard queries are about USD 0.40 per million; Alias queries to AWS resources are free. Health checks of AWS endpoints are free; checks of non-AWS endpoints cost about USD 0.75 per check per month, with optional features (HTTPS, string matching, fast interval) adding small increments. This lab, deleted promptly, costs well under a dollar. Always run the cleanup — an orphaned hosted zone quietly bills USD 0.50 every month.

Common mistakes & troubleshooting

Symptom Likely cause Fix
Domain doesn’t resolve at all Registrar still points at old/auto name servers, not the zone’s four Copy the zone’s exact NS values into the registrar’s name-server settings; allow propagation.
“CNAME at apex not allowed” error Tried to CNAME example.com to an AWS resource Use an Alias A/AAAA record at the apex instead.
Failover never switches Primary record has no health check (or Alias Evaluate Target Health is off) Attach a health check to the primary, or enable Evaluate Target Health on the Alias.
Change made but users see old value for ages TTL was high (e.g. 86,400) Lower TTL before a planned change; for emergencies you can only wait out the cached TTL.
Geolocation users get no answer No default record for unmatched locations Add a geolocation record with location Default.
Private zone returns NXDOMAIN for a name that resolves publicly Private zone is authoritative for that name space and lacks the record Add the record to the private zone, or scope the private zone to a narrower name.
Health check flaps / always unhealthy Firewall blocks Route 53 health-checker IPs, or path returns non-2xx/3xx Allow the route53-healthchecks IP ranges; point the check at a real 200-returning health path.
Records in VPC don’t resolve from instances VPC enableDnsSupport/enableDnsHostnames off, or zone not associated Enable both VPC attributes and associate the private zone with the VPC.

Common beginner mistakes

These are misconceptions rather than symptoms — the wrong mental model that produces the outage in the first place. (The symptom→cause→fix table above is for when you are already staring at the failure.)

Best practices

Security notes

DNS is a security surface, not just plumbing. Enable DNSSEC signing on public zones to let resolvers cryptographically verify that answers are authentic and unmodified — this defends against DNS spoofing and cache poisoning (Route 53 supports DNSSEC signing with KMS-backed keys; you also add a DS record at the parent). Add CAA records to constrain certificate issuance to authorities you trust. Guard against dangling DNS / subdomain takeover: if a record points at a de-provisioned resource (a deleted S3 bucket or released Elastic IP), an attacker can claim that resource and serve content under your name — audit and remove records whose targets no longer exist. Turn on query logging to spot anomalous lookups (data-exfiltration tunnels, malware C2 patterns). Apply least-privilege IAM: scope route53:ChangeResourceRecordSets to specific hosted-zone ARNs so a compromised credential cannot rewrite every zone you own. Finally, for protecting outbound DNS from your VPCs (filtering what your workloads are allowed to resolve), use Route 53 Resolver DNS Firewall — covered in the resolver lesson linked below.

Interview & exam questions

  1. What is the difference between an Alias record and a CNAME, and when must you use an Alias? A CNAME is standard DNS that returns another name (forcing a second lookup), is charged as a query, must be alone at its name, and cannot sit at the zone apex. An Alias is a Route 53 extension on an A/AAAA record that returns the target AWS resource’s IPs directly, is free for AWS targets, can coexist as a normal record, and works at the apex. You must use an Alias to point the apex (example.com) at CloudFront, an ELB, S3 website, etc.

  2. You need to roll a new version of a service to 10% of users, then ramp up. Which routing policy? Weighted routing — give the new stack weight 10 and the old weight 90, then shift the weights. Keep TTL low so the proportions track reality and users re-resolve quickly.

  3. A multi-Region app should serve every user from the fastest Region. Which policy, and what’s its limitation? Latency-based routing. Limitation: it optimises for measured network latency from the resolver, not the user, and ignores geography/compliance — a user near a border can be sent across it. For data-residency, use geolocation.

  4. Failover routing isn’t switching to the secondary even though the primary is down. Why? The primary record almost certainly has no associated health check (and, if it’s an Alias, Evaluate Target Health is off). Without a health signal Route 53 keeps returning the primary. Attach a health check or enable Evaluate Target Health.

  5. What are the three types of health check? Endpoint (probe an IP/domain over HTTP/HTTPS/TCP), calculated (Boolean combination of other health checks — “N of M healthy”), and CloudWatch alarm (mirror the state of any CloudWatch alarm, so you can fail over on metrics like error rate or queue depth).

  6. Explain split-horizon DNS in Route 53 and one gotcha. Run a private hosted zone and a public hosted zone for the same domain; queries from associated VPCs hit the private zone, the internet hits the public one. Gotcha: the private zone is authoritative for its whole name space, so a name it lacks returns NXDOMAIN to the VPC rather than falling through to public resolution.

  7. What does TTL control, and what’s a safe value for a record you might need to fail over? TTL is how long resolvers cache an answer before re-querying. For failover-eligible records use a low TTL (~60 s) so a failover propagates quickly; the extra query cost is negligible versus recovery time.

  8. Why can’t you put a CNAME at example.com? The apex must carry the zone’s NS and SOA records, and DNS forbids a CNAME from coexisting with any other record at the same name. Route 53’s Alias record solves this because it is technically an A/AAAA record.

  9. Geolocation vs geoproximity — what’s the difference? Geolocation routes by the user’s named location (continent/country/US state) with a default fallback. Geoproximity routes by distance between user and resource and lets you apply a bias to expand or shrink each resource’s service area; it requires Traffic Flow.

  10. What is multivalue answer routing and when would you use it over a load balancer? It returns up to eight random, optionally health-checked values, omitting unhealthy ones — health-aware DNS spreading. Use it for a small set of independent endpoints when you want availability without an LB; it is not a true load balancer (no connection awareness, no even distribution).

  11. How does Route 53 decide an endpoint is healthy, and why does it matter for firewalls? Health checkers in multiple global locations probe the endpoint; it’s healthy if more than 18% report success. You must therefore allow the Route 53 health-checker IP ranges through security groups/NACLs/firewalls, or checks fail and traffic drains erroneously.

  12. What is subdomain takeover and how do you prevent it? A “dangling” DNS record points at a de-provisioned resource (deleted bucket, released EIP); an attacker re-creates/claims that resource and serves content under your name. Prevent it by removing records whose targets no longer exist and auditing zones regularly.

Quick check

  1. Which record type points a name at an IPv6 address?
  2. True/false: you can place a CNAME at the zone apex if it is the only record there.
  3. Which routing policy is the right choice for active-passive disaster recovery?
  4. What happens to a geolocation query that matches no rule and has no default record?
  5. Are Alias queries to an AWS resource charged?

Answers

  1. AAAA.
  2. False — a CNAME is never allowed at the apex (the apex must hold NS/SOA); use an Alias.
  3. Failover routing (primary + secondary, with a health check on the primary).
  4. It returns no answer (NODATA) — always configure a Default record for geolocation.
  5. No — Alias queries that resolve to AWS resources are free; standard queries are charged per million.

Exercise

Design the DNS for a two-Region web application (eu-west-1 and us-east-1) fronted by an ALB in each Region, that must (a) serve every user from the lower-latency Region, (b) fail a Region out automatically when its ALB has no healthy targets, and © serve EU users only from eu-west-1 for data-residency. Sketch the records: which routing policies, how they nest, what each record’s Alias target and health configuration are, and what TTLs you’d set. Then write the aws route53 change-resource-record-sets change-batch JSON for the apex records. (Hint: geolocation at the top to honour residency, latency for the rest, Alias-to-ALB with Evaluate Target Health for the per-Region failover, 60 s TTLs.)

Practice challenges

Six exercises, easy to hard. Try each before opening the solution. Commands use placeholder IDs and documentation IP ranges; nothing here needs a real domain, though a few incur the small charges noted in the lab.

1. (Beginner) Prove that a new zone comes with NS and SOA. Create a public hosted zone for shop.example and list only its auto-created NS and SOA records.

<details> <summary><strong>Solution</strong></summary>

ZID=$(aws route53 create-hosted-zone --name shop.example \
  --caller-reference "ex1-$(date +%s)" \
  --query 'HostedZone.Id' --output text | sed 's#/hostedzone/##')
aws route53 list-resource-record-sets --hosted-zone-id "$ZID" \
  --query "ResourceRecordSets[?Type=='NS' || Type=='SOA'].[Name,Type]" --output table

Why: every zone is born with four authoritative name servers (NS) and one SOA; those NS values are what you delegate to at the registrar, and you never delete them. </details>

2. (Beginner) Put a record at the apex where CNAME is illegal. Point the apex shop.example at a CloudFront distribution d111111abcdef8.cloudfront.net.

<details> <summary><strong>Solution</strong></summary>

{ "Changes": [ {
  "Action": "UPSERT",
  "ResourceRecordSet": {
    "Name": "shop.example", "Type": "A",
    "AliasTarget": {
      "HostedZoneId": "Z2FDTNDATAQYW2",
      "DNSName": "d111111abcdef8.cloudfront.net",
      "EvaluateTargetHealth": false
    } } } ] }

Z2FDTNDATAQYW2 is CloudFront’s fixed alias hosted-zone ID (the same for every distribution). Why: the apex must hold NS/SOA, so DNS forbids a CNAME there; an Alias A returns the resource’s IPs directly, works at the apex, and is free. </details>

3. (Intermediate) Ship a 90/10 canary, then pause it. Create weighted records for api.shop.example sending 10% to green (203.0.113.50) and 90% to blue (203.0.113.10), TTL 30. Then take green out of rotation without deleting it.

<details> <summary><strong>Solution</strong></summary>

{ "Changes": [
  { "Action": "UPSERT", "ResourceRecordSet": {
      "Name": "api.shop.example", "Type": "A", "TTL": 30,
      "SetIdentifier": "blue",  "Weight": 90,
      "ResourceRecords": [ { "Value": "203.0.113.10" } ] } },
  { "Action": "UPSERT", "ResourceRecordSet": {
      "Name": "api.shop.example", "Type": "A", "TTL": 30,
      "SetIdentifier": "green", "Weight": 10,
      "ResourceRecords": [ { "Value": "203.0.113.50" } ] } }
] }

To pause the canary, UPSERT the green record with "Weight": 0. Why: the returned share is weight ÷ sum, so 10 ÷ 100 = 10% of queries; weight 0 removes a record from rotation while keeping it defined for an instant resume. </details>

4. (Intermediate) Localise with no NODATA holes. Configure geolocation so users in Europe get eu (198.51.100.20) and everyone else gets global (198.51.100.99).

<details> <summary><strong>Solution</strong></summary>

{ "Changes": [
  { "Action": "UPSERT", "ResourceRecordSet": {
      "Name": "www.shop.example", "Type": "A", "TTL": 60,
      "SetIdentifier": "eu",     "GeoLocation": { "ContinentCode": "EU" },
      "ResourceRecords": [ { "Value": "198.51.100.20" } ] } },
  { "Action": "UPSERT", "ResourceRecordSet": {
      "Name": "www.shop.example", "Type": "A", "TTL": 60,
      "SetIdentifier": "global", "GeoLocation": { "CountryCode": "*" },
      "ResourceRecords": [ { "Value": "198.51.100.99" } ] } }
] }

Why: CountryCode: "*" is the Default record; without it, any location that matches no rule (most of the planet, here) gets no answer at all. </details>

5. (Advanced) Fastest Region, self-healing, no separate health-check object. Sketch the apex records for a two-Region app that serves every user from the lowest-latency Region and drains a Region automatically when its ALB has zero healthy targets.

<details> <summary><strong>Solution</strong></summary>

Two latency records for the apex, one per Region, each an Alias to that Region’s ALB with EvaluateTargetHealth: true — no standalone health check needed:

{ "Changes": [
  { "Action": "UPSERT", "ResourceRecordSet": {
      "Name": "app.shop.example", "Type": "A",
      "SetIdentifier": "eu", "Region": "eu-west-1",
      "AliasTarget": { "HostedZoneId": "Z32O12XQLNTSW2",
        "DNSName": "dualstack.eu-alb-123.eu-west-1.elb.amazonaws.com",
        "EvaluateTargetHealth": true } } },
  { "Action": "UPSERT", "ResourceRecordSet": {
      "Name": "app.shop.example", "Type": "A",
      "SetIdentifier": "us", "Region": "us-east-1",
      "AliasTarget": { "HostedZoneId": "Z35SXDOTRQ7X7K",
        "DNSName": "dualstack.us-alb-456.us-east-1.elb.amazonaws.com",
        "EvaluateTargetHealth": true } } }
] }

The HostedZoneId in each AliasTarget is the ELB zone ID for that Region (per-Region constants, not your zone). Why: latency chooses the Region; EvaluateTargetHealth makes Route 53 stop returning a Region whose ALB has no healthy targets — so the latency policy silently omits the dead Region. Add a maintenance page by wrapping each leaf in a failover pair via Traffic Flow. </details>

6. (Advanced) Sign a zone, and describe how to turn it off safely. Outline the steps to enable DNSSEC on a public zone ZID, then the correct order to disable it.

<details> <summary><strong>Solution</strong></summary>

# 1) asymmetric KMS key in us-east-1, ECC_NIST_P256, with a policy granting
#    dnssec-route53.amazonaws.com kms:DescribeKey/GetPublicKey/Sign
KID=arn:aws:kms:us-east-1:123456789012:key/EXAMPLE-KEY-ID
# 2) create the Key Signing Key and enable signing
aws route53 create-key-signing-key --caller-reference "ksk-$(date +%s)" \
  --hosted-zone-id "$ZID" --key-management-service-arn "$KID" \
  --name shopKSK --status ACTIVE
aws route53 enable-hosted-zone-dnssec --hosted-zone-id "$ZID"
# 3) copy the DS record Route 53 shows you into the PARENT (registrar/TLD)

Disable order (the part that bites): remove the DS record at the parent first, wait out its TTL so caches forget it, and only then run disable-hosted-zone-dnssec. Why: if signing stops while a DS still points at your zone, every validating resolver returns SERVFAIL for the entire domain — a self-inflicted outage that stale caches keep serving. </details>

Certification mapping

Glossary

Next steps

AWSRoute 53DNSRouting PoliciesHealth ChecksNetworking
Need this built for real?

Vinod is a Senior Cloud Architect (22+ yrs) — available for Azure / AWS / GCP architecture, landing zones, and migrations.

Work with me

Comments