In a nutshell
Picture the internet as a vast city where every server is a building with a numeric street address — its IP address. Nobody memorises street addresses; we remember names (“the coffee shop on Main Street”). DNS (the Domain Name System) is the city-wide directory that turns a name you can remember into the address a computer actually needs, and Amazon Route 53 is AWS’s version of that directory — with a smart traffic controller bolted on top.
Route 53 wears three hats, and beginners often blur them. It is a registrar (you can buy a name like example.com from it), it is the authoritative directory for that name (the single source of truth that says “example.com lives at this address”), and it is a traffic director that can hand back different addresses to different people — the nearest datacentre, a healthy server instead of a broken one, 5% of users to a new version for testing. That third hat is what lifts Route 53 above a plain phone book.
Why should a beginner care? Because Route 53 is the first thing every request touches, before any server you pay for. A small mistake here — the wrong record type, or a cache told to last a day when you need to fail over in a minute — becomes an outage that the rest of your perfectly healthy architecture cannot fix. Learn DNS and you can build a front door that is fast, resilient, and fails over in seconds.
The lesson goes deep, but you can read it in layers: the plain-language core first, then the Going deeper and Practice challenges sections when you want to push further.
Level: Intermediate · Time: ~52 min · You need: basic AWS Console/CLI comfort and a rough idea of what a load balancer does; every DNS concept itself is built up from scratch here.
Every connection on the internet begins with a question: what is the IP address for this name? The Domain Name System (DNS) answers it, and on AWS that answer is served by Amazon Route 53 — a highly available, globally distributed authoritative DNS service named after the port DNS runs on, port 53. Route 53 does three distinct jobs that people often conflate: it registers domain names (you can buy example.com through it), it hosts the authoritative records for a domain (it is the source of truth resolvers ask), and it routes traffic intelligently using policies that go far beyond plain DNS — weighting answers for canary releases, sending users to the lowest-latency Region, failing over to a backup when a health check goes red, and answering differently based on the user’s geography.
Route 53 sits at the very front of almost every architecture. It is the first thing a client touches, before the load balancer, before CloudFront, before any compute you pay for. Get it right and you have a resilient, fast-resolving front door with sub-second failover. Get it wrong — a stray CNAME at the apex, a TTL of 86,400 seconds on a record you need to fail over, a health check pointed at the wrong port — and you have an outage that DNS caches will keep serving for hours. This lesson is the exhaustive version: every record type, the Alias-versus-CNAME distinction that trips up nearly everyone, all seven routing policies with explicit when-to-use guidance, the three kinds of health check, and how TTL governs how fast the world sees your changes.
Learning objectives
By the end of this lesson you will be able to:
- Distinguish Route 53’s three roles — domain registration, authoritative hosting, and traffic routing — and explain how a recursive resolver finds your records.
- Create and manage public and private hosted zones, and explain how split-horizon DNS works.
- Choose the correct record type for any scenario (A, AAAA, CNAME, MX, TXT, NS, SOA, SRV, CAA, PTR) and use Alias records correctly — including at the zone apex where CNAME is forbidden.
- Select the right routing policy (simple, weighted, latency, failover, geolocation, geoproximity, multivalue answer) for a given requirement, and combine them.
- Configure health checks (endpoint, calculated, and CloudWatch-alarm) and wire them into failover so traffic drains from unhealthy targets automatically.
- Reason about TTL as the lever that trades query cost and resolution speed against how fast a change propagates.
Prerequisites & where this fits
You should be comfortable with the AWS Console and the AWS CLI (see AWS Hands-On First Steps: Console, CLI, CloudShell, SDKs & Access Keys), and you should have met load balancers in AWS Elastic Load Balancing, In Depth: ALB, NLB, GWLB & Target Groups — Route 53 most often points at an ALB or CloudFront distribution. A working mental model of how a domain delegates to name servers helps but is not assumed; we build it here. This lesson belongs in the Networking module of the AWS Zero-to-Hero course, immediately after load balancing, because the natural progression is: distribute traffic within a Region (ELB), then distribute and fail over traffic across Regions and endpoints (Route 53). It also feeds directly into the edge and DNS-security lessons cross-linked at the end.
Core concepts
Before the settings, the mental model. DNS is a hierarchical, distributed database that maps human-readable domain names to machine-usable data, most commonly IP addresses. The hierarchy reads right-to-left: the root (.), then the top-level domain (TLD — .com, .org, .co.uk), then your registered domain (example.com), then any subdomains (api.example.com).
When a user’s machine needs to resolve api.example.com, it does not ask Route 53 directly. It asks a recursive resolver (typically run by the user’s ISP, or a public one like 8.8.8.8). That resolver walks the hierarchy: it asks a root server “who handles .com?”, asks the .com TLD servers “who is authoritative for example.com?”, and they answer with the name servers for your domain — which, if you host the zone in Route 53, are four Route 53 name servers. The resolver then asks one of those Route 53 name servers for api.example.com, gets the record, caches it for the TTL, and returns it to the user. Route 53 is the authoritative server at the end of that chain — it holds the truth. Everything else is asking and caching.
A few load-bearing terms:
- Hosted zone — a container for all the DNS records of one domain (and its subdomains). It is the Route 53 representation of a zone. Creating one gives you a set of four authoritative name servers.
- Record set (resource record set) — an individual entry in a zone: a name, a type, a TTL, and a value (or, for Alias records, a target AWS resource). Route 53 calls them “records” in the console.
- Authoritative vs recursive — Route 53 is authoritative (it answers for zones you own). Resolvers are recursive (they chase answers on a client’s behalf and cache them). Route 53 is not a recursive resolver for arbitrary internet names — the Route 53 Resolver (covered in a separate lesson) is the in-VPC recursive service.
- TTL (time to live) — how many seconds a resolver may cache an answer before asking again. It governs propagation speed.
- Delegation — pointing a parent zone (the registrar’s
.comentry, or a parent hosted zone) at the name servers of a child zone viaNSrecords. This is how the hierarchy connects.
With that, the settings.
Hosted zones: every setting
A hosted zone is where you start. There are two kinds, and the difference is who can see the records.
Public hosted zones
A public hosted zone is authoritative on the public internet. Any resolver in the world can query it. You create one for a domain you intend to serve publicly (example.com). On creation Route 53 assigns four name servers (an NS record) and a SOA record, both created automatically — never delete them. To make the zone live, you must delegate to those four name servers from the parent: if you registered the domain through Route 53, you point the registered domain at the zone’s name servers (Route 53 can do this automatically); if you registered elsewhere, you copy the four name servers into your registrar’s control panel.
| Setting | What it does | Choices / default | When to change / gotcha |
|---|---|---|---|
| Domain name | The apex/root the zone is authoritative for | Any DNS name; immutable after creation | Must match the registered domain. Typo means nothing resolves — delete and recreate. |
| Type | Public vs private | Public (default) | Public = internet-visible. Cannot convert a zone’s type after creation. |
| Comment | Free-text description | Optional | Use it to record owner/ticket; editable later. |
| Name servers (NS) | The four authoritative servers Route 53 assigns | Auto-generated, 4 per zone | Copy exactly these into your registrar. Two zones for the same name get different name servers — only the delegated set is live. |
| SOA | Start-of-authority metadata (primary NS, contact, serial, timers) | Auto-generated | Rarely edited; the minimum TTL field affects negative-caching. Do not delete. |
A subtle and very common mistake: deleting a public hosted zone and recreating it gives you a brand-new set of four name servers. The old delegation at the registrar now points at name servers that no longer host the zone, and the domain goes dark until you update the registrar. Treat hosted-zone name servers as something to wire once and leave alone.
Private hosted zones
A private hosted zone is authoritative only inside one or more VPCs you associate with it. Queries from those VPCs (via the Route 53 Resolver at the VPC .2 address) get the private answers; the rest of the internet cannot see the zone at all. This is how you give internal services friendly names — db.internal.example.com resolving to a private IP — without exposing topology publicly.
| Setting | What it does | Choices / default | When to change / gotcha |
|---|---|---|---|
| Domain name | The zone the records cover | Any name (often a real domain or an internal-only one) | You can run a private zone for the same name as a public zone — this is split-horizon DNS (below). |
| VPCs to associate | Which VPCs see these records | One or more, any Region/account | The VPC must have enableDnsHostnames and enableDnsSupport both on, or resolution fails silently. Cross-account association needs a CLI authorisation step. |
| Region | Where the zone is created | Any; association can span Regions | Private zones are global objects but associations are per-VPC. |
Split-horizon (split-view) DNS is the headline use case: associate a private zone for example.com with your VPCs and keep a public zone for example.com on the internet. Inside the VPC, app.example.com resolves to a private ALB; outside, the same name resolves to a public CloudFront distribution. Route 53 evaluates the most specific matching private zone first for queries from an associated VPC, falling back to public resolution only if no private zone covers the name. The classic gotcha: if a private zone for example.com exists but lacks a record for legacy.example.com, queries from the VPC get NXDOMAIN rather than falling through to the public zone — the private zone is authoritative for the whole name space it covers.
Record types: every type you will meet
A record (resource record set) has a name, a type, a TTL (except Alias), and a value. The type tells resolvers what kind of data to expect. Route 53 supports the full standard set; these are the ones you will actually configure.
| Type | What it holds | Typical value | Notes & gotchas |
|---|---|---|---|
| A | IPv4 address | 203.0.113.10 |
The workhorse. Can be an Alias (see below) instead of a literal IP. |
| AAAA | IPv6 address | 2001:db8::1 |
The IPv6 equivalent of A; also Alias-capable. Add it whenever you serve IPv6. |
| CNAME | Canonical name (an alias to another name) | lb-123.eu-west-1.elb.amazonaws.com |
Returns a name, not an IP; the resolver must look that up too. Forbidden at the zone apex and must be the only record at its name. |
| NS | Name servers for a zone or delegated subdomain | four ns-xxx.awsdns-xx.* |
Created automatically for the zone apex. Add your own NS records to delegate a subdomain to a different zone/provider. |
| SOA | Start of authority — zone metadata and timers | ns-... hostmaster... serial refresh retry expire minTTL |
One per zone, auto-created. The last field sets negative caching (how long NXDOMAIN is cached). |
| MX | Mail exchanger — where email for the domain goes | 10 mail.example.com |
The number is priority (lower = preferred). Required for receiving email. |
| TXT | Arbitrary text | "v=spf1 include:_spf.google.com ~all" |
Used for SPF, DKIM, DMARC, and domain-ownership verification. Quote each string; 255-char chunks. |
| SRV | Service location (host + port + priority + weight) | 10 60 5060 sip.example.com |
For protocols that advertise host and port (SIP, LDAP, some game and chat services). |
| CAA | Which Certificate Authorities may issue certs for the domain | 0 issue "amazon.com" |
A security control — stops a rogue CA issuing a cert for your domain. Add amazon.com so ACM can issue. |
| PTR | Reverse DNS (IP → name) | host.example.com |
Lives in special in-addr.arpa / ip6.arpa zones; used for reverse lookups and mail-server reputation. |
| NAPTR / DS / SPF / others | Telephony rewriting, DNSSEC delegation signer, legacy SPF | varies | Less common; Route 53 supports them. SPF-the-type is deprecated — use TXT for SPF. |
Two practical rules that catch people: a CNAME must be alone at its name (you cannot have a CNAME and an A for www.example.com), and you cannot put a CNAME at the apex (example.com itself), because the apex must also carry the NS and SOA records and the DNS spec forbids a CNAME coexisting with other records. The fix for the apex is the Alias record.
Alias vs CNAME: the distinction that trips everyone
This is the single most-asked Route 53 interview question, so understand it cold.
A CNAME is standard DNS: it says “this name is really that name; go look that up.” It works for any target, AWS or not, but it returns a name, forcing the resolver to do a second lookup, and it is forbidden at the apex and must be alone at its record name.
An Alias record is a Route 53 extension (not standard DNS). It points an A or AAAA record directly at a supported AWS resource — and at resolution time Route 53 substitutes that resource’s current IP address(es) into the answer. To the resolver it looks like a normal A/AAAA answer (it gets IPs, not a name), so there is no second lookup and no charge for Alias queries to AWS resources. Crucially, an Alias works at the zone apex, which is why example.com → CloudFront is always an Alias, never a CNAME.
| Dimension | CNAME | Alias |
|---|---|---|
| Standard? | Yes (RFC) | No — Route 53 only |
| Returns | A name (triggers another lookup) | IP address(es) directly |
| Works at apex? | No | Yes |
| Can coexist with other records at the name? | No (must be alone) | Yes (it is an A/AAAA) |
| Targets | Any DNS name | Specific AWS resources + same-zone records (see below) |
| Query cost | Charged as a normal query | Free when pointing at an AWS resource |
| Health/failover integration | Manual | Evaluate Target Health auto-tracks the target |
| TTL | You set it | Inherited from the target (you cannot set it) |
Alias targets you can point at: CloudFront distributions, ELB load balancers (ALB/NLB/CLB), S3 website endpoints, API Gateway, VPC interface endpoints, Elastic Beanstalk environments, Global Accelerator, AppSync, and — very usefully — another record in the same hosted zone. That last one lets you alias www.example.com to example.com and maintain the IP in one place.
The Evaluate Target Health toggle on an Alias is the quiet superpower: set it to Yes and Route 53 stops returning that Alias if the underlying resource (e.g. all targets behind an ALB) is unhealthy — health checking you get for free, without creating a separate health check, as long as you are aliasing an ELB, CloudFront, or another Route 53 record that is itself health-checked.
When to use which: Alias for anything pointing at an AWS resource (always — it is free, faster, and apex-capable); CNAME only for pointing a subdomain at a non-AWS name (a SaaS endpoint, a partner’s host) or where you genuinely need standard-DNS behaviour.
Routing policies: all seven, with when-to-use
A routing policy is set per record and decides which value Route 53 returns when several records share the same name and type. This is where Route 53 stops being plain DNS and becomes a traffic director. There are seven.
1. Simple
One record, one answer (or, if you give multiple values, Route 53 returns them all in random order and the client picks). No health checking. This is ordinary DNS.
- When to use: a single resource, or when you do not need conditional logic. The default.
- Gotcha: with multiple values in a simple record you get crude client-side load spreading but no failover — a dead IP stays in the answer set.
2. Weighted
Multiple records, same name/type, each with a weight (0–255). Route 53 returns each in proportion to its weight ÷ total. Weight 0 takes a record out of rotation (unless all are 0, in which case all are returned equally).
- When to use: canary / blue-green releases (send 5% to the new stack, 95% to the old, then shift), A/B testing, gradually migrating between Regions or providers.
- Gotcha: proportions are over queries, and resolver caching means an individual user sticks to one answer for the TTL — keep TTLs low while shifting weights, and remember weight is per-record so percentages are weight/sum, not absolute.
3. Latency-based
Multiple records, each tagged with an AWS Region. Route 53 returns the record whose Region gives the lowest network latency to the resolver, based on AWS’s continuously-measured latency map.
- When to use: active-active multi-Region deployments where you want each user served from the fastest Region (a global web app with stacks in
eu-west-1andus-east-1). - Gotcha: it optimises for latency, not geography or compliance — a user near a border can be routed across it. It measures from the resolver, not the user, so users behind a distant DNS resolver may be routed sub-optimally. For data-residency rules use geolocation instead.
4. Failover
A primary and a secondary record. Route 53 returns the primary while its health check is healthy, and switches to the secondary when the primary fails. The classic active-passive pattern.
- When to use: disaster recovery — a hot primary site with a standby (which can be an S3 static “we’ll be right back” page, or a second-Region stack).
- Gotcha: the primary must have an associated health check (or be an Alias with Evaluate Target Health) or failover never triggers. The secondary should be reachable independently of whatever took the primary down.
5. Geolocation
Returns a different record based on the geographic location of the user (resolver), matched by continent, country, or — for the US — state. You can set a default record for locations that match no rule.
- When to use: content localisation (serve a French site to users in France), licensing / compliance (keep EU users on EU infrastructure), geo-blocking (return a “not available in your region” answer).
- Gotcha: always set a default record — a user whose location matches no rule and has no default gets no answer (NODATA). Location is inferred from the resolver’s IP and can be wrong for VPN/corporate-resolver users.
6. Geoproximity
Routes based on the geographic distance between the user and your resources, with a bias you can dial (–99 to +99) to expand or shrink the geographic area a resource serves. Configured via Route 53 Traffic Flow (the visual policy editor).
- When to use: when you want geographic routing but with control over the boundaries — shift more traffic to a resource by increasing its bias, e.g. to drain a Region gradually or to balance load between two nearby Regions.
- Gotcha: distinct from geolocation — geolocation matches named regions (country/state); geoproximity matches distance and lets you warp the map with bias. It requires Traffic Flow.
7. Multivalue answer
Returns up to eight healthy records chosen at random from a larger set, each optionally health-checked. It is like a simple record with multiple values plus health checking, giving you crude DNS-level load distribution that automatically omits unhealthy endpoints.
- When to use: improving availability for a set of independent endpoints (several web servers by IP) where you want health-aware spreading without a load balancer.
- Gotcha: it is not a substitute for a real load balancer — there is no connection awareness, no even distribution guarantee, and clients cache one answer per TTL. Each value needs its own health check to benefit from the health-aware behaviour.
| Policy | Returns based on | Health-check aware? | Signature use case |
|---|---|---|---|
| Simple | Single config | No | One resource, plain DNS |
| Weighted | Assigned weights | Yes (per record) | Canary / blue-green, A/B |
| Latency | Lowest measured latency | Yes (per record) | Active-active multi-Region for speed |
| Failover | Primary health | Yes (required) | Active-passive DR |
| Geolocation | User’s country/continent/state | Yes (per record) | Localisation, compliance, geo-block |
| Geoproximity | Distance + bias | Yes (per record) | Distance routing with tunable boundaries |
| Multivalue | Up to 8 random healthy values | Yes (per record) | Health-aware spreading without an LB |
You can nest policies with Traffic Flow — e.g. latency-based at the top to pick a Region, then weighted within each Region for a canary, then failover under each weight. That composition is how large multi-Region systems are actually expressed.
Worked example: reading a routing decision end to end
Policies are easier to trust once you trace a real answer. Three short scenarios, each computed by hand.
Weighted — the maths behind a canary. You run two stacks behind the name api.example.com (type A): blue with weight 90 and green with weight 10, both TTL 30. For each incoming query Route 53 rolls a weighted die: it returns green with probability 10 ÷ (90 + 10) = 10% of queries, blue for the other 90%. The subtle part is queries, not users: a resolver that gets green caches it for the 30-second TTL, so every user behind that resolver sticks to green for half a minute. Over a large, busy population the proportions converge on 10/90; over a handful of users on one corporate resolver they can be lumpy. To pull green out of rotation without deleting its record, set its weight to 0 — Route 53 stops returning it (unless every record at that name is 0, in which case they are all returned equally). To ramp the canary, walk the weights 10→25→50→100 while the low TTL keeps the split tracking reality.
Nesting latency over failover. Now make it multi-Region and self-healing with a Traffic Flow tree: a latency rule at the top chooses between eu-west-1 and us-east-1, and under each Region a failover pair points primary at that Region’s ALB (Alias, Evaluate Target Health = Yes) and secondary at an S3 “maintenance” page. Trace a query from a London resolver: latency map → eu-west-1 → primary healthy → return the ALB’s IPs. Now every target behind the eu-west-1 ALB fails health checks: the primary leaf goes unhealthy, failover returns the secondary S3 page — while a New York user is still served the healthy us-east-1 primary. The order matters: latency is evaluated first to pick the leaf, then failover is resolved inside the chosen leaf. Nesting is how large systems compose “fast and self-healing” from two simple rules.
Geolocation — most specific wins. You publish three geolocation records for www.example.com: CountryCode=FR → fr-stack, ContinentCode=EU → eu-stack, and a Default → global. A user in France matches all three conceptually, but Route 53 picks the most specific match — state beats country beats continent beats Default — so France gets fr-stack. A user in Germany matches EU (not FR) → eu-stack. A user in Brazil matches none of the named rules → Default → global. Delete that Default record and the Brazilian user gets no answer at all (NODATA), which is the single most common geolocation outage.
Health checks: every kind
A health check is a separate Route 53 object that monitors a target and reports healthy/unhealthy; routing policies consult it to decide whether to return a record. Route 53 health checkers run from multiple AWS locations worldwide and a target is considered up if more than 18% of checkers see it as healthy (this is why you must allow the Route 53 health-checker IP ranges through firewalls). There are three types.
Endpoint health checks
Monitor an IP or domain name on a chosen protocol. The settings:
| Setting | What it does | Choices / default | When to change / gotcha |
|---|---|---|---|
| Protocol | How to probe | HTTP, HTTPS, TCP | HTTP(S) lets you check a path and status; TCP only checks the port opens. HTTPS does not validate the certificate by default. |
| Endpoint | What to probe | IP address or domain name + port | If you use a domain name, Route 53 resolves it each check. Use an IP to pin it. |
| Path (HTTP/S) | Which URL to request | e.g. /health; default / |
Point at a deep health endpoint that checks dependencies, not a static page that is “up” while the app is broken. |
| Request interval | Probe frequency | Standard 30 s or Fast 10 s | Fast detects failure sooner but costs more and is noisier. |
| Failure threshold | Consecutive fails before “unhealthy” | 1–10, default 3 | Lower = faster failover, more false positives on a blip. 3×30s ≈ 90s to flip. |
| String matching | Require a string in the first 5,120 bytes of the response body | Off / on with search string | Catches “200 OK but wrong content” — e.g. require "OK" in the body. |
| Latency graphs | Record response time in CloudWatch | Off (default) / on | Turn on to alarm on slow-but-up endpoints. |
| Invert health status | Treat healthy as unhealthy and vice-versa | Off (default) | Niche — e.g. fail over to a site only when a maintenance flag returns 200. |
| Health checker regions | Which checker locations probe | Default set / custom | Reduce to fewer regions to cut noise, but keep enough for the 18% rule. |
| SNI (HTTPS) | Send the hostname in the TLS handshake | On by default | Required for endpoints that serve multiple certs on one IP. |
Calculated health checks
A health check whose status is derived from other health checks using a Boolean rule — “healthy if at least N of these child checks are healthy”. It probes nothing itself.
- When to use: model a service that is up only if several components are up (app and database and cache), or, with a low threshold, an any-of OR.
- Gotcha: you set “report healthy when N of M are healthy”; choose N deliberately (N=M is AND, N=1 is OR).
CloudWatch-alarm health checks
A health check that mirrors the state of a CloudWatch alarm. The check is unhealthy when the alarm is in ALARM. This lets you health-check anything CloudWatch can measure — DynamoDB throttles, SQS queue depth, ELB 5xx rate, a custom metric — not just an HTTP endpoint.
- When to use: failover driven by an internal signal that is not a simple endpoint probe (e.g. “fail over the Region when its error-rate alarm fires”).
- Gotcha: choose what happens when the alarm has insufficient data (treat as healthy / unhealthy / last known) — getting this wrong causes spurious failovers when a metric goes quiet.
Health checks integrate with routing in two ways: associate a health check with a record (failover, weighted, latency, geolocation, multivalue all honour it and stop returning unhealthy records), or use Evaluate Target Health on an Alias to inherit the target’s health automatically. A frequent design is a failover record pair where the primary is an Alias to an ALB with Evaluate Target Health = Yes — no manual health check object needed.
Worked example: how long a failover really takes, and the 18% rule
Two numbers turn “Route 53 fails over” into a real recovery-time figure you can defend in a design review.
Detection time = interval × threshold. An endpoint health check with the Standard 30 s interval and the default failure threshold of 3 needs three consecutive failed probes before it flips to unhealthy: roughly 90 seconds. Only then does Route 53 stop returning the record — but a client that already resolved the name still holds the old answer until its cached TTL expires. So the user-visible failover is approximately detection + record TTL: with a 60 s TTL that is ~90 s + up to 60 s ≈ 150 s worst case. Switch to the Fast 10 s interval with a threshold of 2 and detection drops to ~20 s (plus TTL) — faster recovery, but more probes, more cost, and more false positives on a transient blip. This is why failover-eligible records live on a 60 s TTL: the TTL, not the health check, is often the bigger half of your recovery budget.
The 18% rule and why firewalls matter. Route 53 runs health checkers from many AWS Regions at once, and an endpoint is considered healthy when more than 18% of those checkers report success. The practical consequence is a footgun: if your security group, NACL, or on-prem firewall blocks the Route 53 health-checker source ranges, every checker fails, you fall below 18%, and Route 53 marks a perfectly healthy endpoint down — draining live traffic. Allowlist the published ranges (service ROUTE53_HEALTHCHECKS in AWS’s ip-ranges.json), or point the check at a public path that is already open.
Composing a calculated check. Suppose “the service is up” means app AND database are healthy, and at least one of three cache nodes is healthy. You cannot express that in one flat rule, so you nest: build endpoint checks for app, db, cache1/2/3; a calculated check caches-ok that is healthy when ≥ 1 of {cache1, cache2, cache3}; then an outer calculated check healthy when 3 of {app, db, caches-ok} are healthy (N = M is a logical AND, N = 1 is OR). For CloudWatch-alarm checks, also decide the insufficient-data behaviour (healthy / unhealthy / last known state) — a metric that simply goes quiet must not silently trigger a Region-wide failover.
TTL: the propagation lever
TTL (time to live), in seconds, tells every resolver how long it may cache an answer before re-asking Route 53. It is the single biggest control over how fast a DNS change reaches users — and a constant trade-off.
- High TTL (e.g. 86,400 = 1 day): fewer queries to Route 53 (lower cost), faster resolution for repeat users (it’s cached), but a change — including a failover — can take up to a full TTL to propagate. A day-long TTL on a record you need to fail over means up to a day of stale answers.
- Low TTL (e.g. 60 s): changes propagate fast and failover is quick, but more queries hit Route 53 (more cost) and resolvers cache less.
| Record purpose | Sensible TTL | Reasoning |
|---|---|---|
Stable apex/www pointing at CloudFront (Alias) |
n/a — Alias TTL is managed | Route 53 handles it; you can’t set it. |
| Records you may fail over | 60 s | Fast failover; the small extra query cost is worth the recovery time. |
Stable MX, TXT (SPF/DKIM), NS |
3,600–86,400 s | Rarely change; cache hard. |
| A record you’re about to migrate | lower it to 60 s a day before | So the cut-over propagates quickly, then raise it again. |
Two things people miss: Alias records to AWS resources have a TTL managed by Route 53 (you cannot set it), and negative answers (NXDOMAIN) are cached according to the minimum TTL in the SOA record, so a typo that returns NXDOMAIN can stick in caches even after you fix it.
The diagram traces a single query from a client through the recursive resolver to a Route 53 hosted zone, then shows the same name resolving differently under each routing policy and how health checks gate which records are returned.
Going deeper
The core lesson gives you a resilient front door. This section is for the reader who has to operate it — the eighth routing policy the record screen never shows you, the truth about “the user’s location,” hybrid resolution, DNS signing, and the recovery-orchestration layer that sits above plain failover.
The eighth policy: IP-based routing
The record-creation screen offers seven policies, but Route 53 has an eighth — IP-based routing — configured through CIDR collections rather than the policy dropdown. You upload a set of named CIDR locations (each a bunch of IP ranges, e.g. mobile-isp-A = 198.51.100.0/24, 203.0.113.0/24), then create records that each map to a location. When a query arrives, Route 53 looks at the source IP, finds the location whose CIDR block contains it, and returns that location’s record.
Why it exists: geolocation and latency both infer where to send a client, and both can be wrong for a network you actually know something about. IP-based routing lets you assert ground truth — “traffic from this ISP’s ranges goes to this peering point,” “this partner’s corporate CIDR gets the internal build” — which is common in telco, CDN, and large-enterprise designs. A sketch:
# 1) create a CIDR collection and add a named location
COLL=$(aws route53 create-cidr-collection --name known-networks \
--caller-reference "cidr-$(date +%s)" --query 'Collection.Id' --output text)
aws route53 change-cidr-collection --id "$COLL" \
--changes '[{"LocationName":"isp-A","Action":"PUT","CidrList":["198.51.100.0/24","203.0.113.0/24"]}]'
# 2) a record then references CidrRoutingConfig: {CollectionId, LocationName}
# (with a second record using LocationName "*" as the catch-all default)
The same “always define a default” discipline from geolocation applies: give one record the location name * so an IP that matches no block still gets an answer.
“The user’s location” is usually the resolver’s — until EDNS Client Subnet
A caveat runs through latency, geolocation, geoproximity, and IP-based routing: Route 53 sees the recursive resolver’s IP, not the end user’s. A user on a VPN or a centralised corporate resolver can be geolocated hundreds of kilometres away. The mitigation is EDNS Client Subnet (ECS), an extension where a resolver forwards a truncated slice of the client’s subnet (e.g. 203.0.113.0/24, not the full address) alongside the query. Public resolvers such as Google and OpenDNS support ECS; Route 53 (and CloudFront) use it when present to target the real user’s network. You cannot force it — many resolvers strip ECS for privacy — so treat resolver-based routing as usually right, occasionally coarse, and never as a compliance control on its own. For hard data-residency, gate with geolocation and enforce residency in the application/data tier, not in DNS alone.
Route 53 Resolver: the in-VPC recursive service and hybrid DNS
Everything so far is Route 53 the authoritative service. Inside a VPC there is a second, separate thing: the Route 53 Resolver — the recursive resolver every instance already uses at the VPC’s base-of-range .2 address (also reachable at 169.254.169.253). It answers amazonaws.com, resolves your private hosted zones, and recurses to the public internet. Left alone it needs no configuration. Hybrid networks need two add-ons, each a set of ENIs with IPs in your subnets:
- Inbound endpoint — lets DNS queries from outside AWS (your on-prem data centre over VPN/Direct Connect) resolve names in your private hosted zones. You point your on-prem forwarders at the inbound endpoint’s IPs.
- Outbound endpoint + forwarding rules — lets the Resolver forward queries for specific domains from your VPC to an external resolver. A FORWARD rule (“
corp.internal→ 10.0.0.53, 10.0.0.54”) sends matching queries on-prem; a SYSTEM rule carves an exception back to normal resolution for a sub-name. Rules attach to VPCs and can be shared across accounts with AWS RAM, so one networking account owns hybrid DNS for the whole org.
This is a large topic — outbound filtering with Resolver DNS Firewall, rule precedence, and full hybrid patterns get their own treatment in the Route 53 Resolver lesson. The one thing to fix now: hosted zones answer for names you own; the Resolver decides how a VPC resolves everything.
Two query-logging features people conflate
“Turn on Route 53 query logging” means two different things depending on which side you are on, and mixing them up loses you the logs you actually needed:
- Public hosted-zone query logging records the DNS queries Route 53 answers authoritatively for a public zone. It goes to CloudWatch Logs only, and the log group must live in
us-east-1regardless of anything else. Use it to see who is looking up your public names. - Resolver query logging records the DNS queries originating from resources inside a VPC — every lookup an instance makes, whether answered by a private zone, forwarded on-prem, or recursed to the internet. It can go to CloudWatch Logs, S3, or Kinesis Data Firehose. This is the security-relevant one: DNS is a favourite tunnel for data exfiltration and malware command-and-control, and Resolver logs are where you spot the anomalous, high-entropy, high-volume lookups. Ship them somewhere queryable — see the observability lesson for the pipeline.
DNSSEC signing: keys, the chain of trust, and the two-phase dance
Plain DNS answers are unauthenticated — a resolver cannot tell a genuine reply from a spoofed one, which is what cache-poisoning attacks exploit. DNSSEC fixes this by cryptographically signing a zone so validating resolvers can verify each answer. Route 53’s model:
- You provide a Key Signing Key (KSK) backed by an asymmetric KMS key, algorithm ECC_NIST_P256 (ECDSAP256SHA256), created in
us-east-1— hosted zones are global, but the DNSSEC KSK is us-east-1 regardless of where you think of the zone as living. The KMS key policy must let the service principaldnssec-route53.amazonaws.comuse it. Route 53 manages the Zone Signing Key (ZSK) internally. - Enabling signing makes Route 53 add RRSIG (signatures), DNSKEY (public keys), and NSEC (authenticated denial-of-existence) records.
- Trust is established by publishing a DS record — a hash of your KSK — in the parent zone (the TLD, set through your registrar). Route 53 hands you the exact DS contents to paste.
The operational trap is ordering, and getting it wrong takes the entire domain down for validating resolvers (they return SERVFAIL, not a stale answer). Enable: turn on signing first, confirm the zone is signed, then add the DS at the parent. Disable or migrate registrars: remove the DS at the parent first, wait out the DS record’s TTL so every cache forgets it, and only then disable signing or move. KSK rotation follows the same add-new-before-removing-old discipline, and Route 53 supports it. Treat DNSSEC as a two-person-rule change with monitoring on signature expiry.
Route 53 Application Recovery Controller (ARC) and zonal shift
Failover routing is reactive and per-record. For large multi-Region systems that must fail over deliberately and safely, AWS offers Route 53 Application Recovery Controller:
- Readiness checks continuously audit that your standby is actually ready — capacity, quotas, and configuration matched across cells — so failover does not tip traffic into a Region that cannot take it.
- Routing controls are simple on/off switches you flip during a failover. Each is backed by a routing-control health check that ordinary Route 53 records reference, so flipping the switch changes which records resolve. The control plane is a highly available cluster spread across five Regions with Regional endpoints, so you can still operate the switches during a Regional impairment — the whole point of a recovery tool.
- Safety rules (assertion and gating rules) stop unsafe states, e.g. “never turn both cells off at once.” ARC clusters are not cheap (on the order of ~$2.50/hour, ≈ $1,800/month per cluster), so they are for tier-1 systems where a botched failover is the bigger risk. Pair this with a broader DR strategy such as cross-Region recovery with AWS DRS.
Zonal shift is ARC’s cheaper, narrower cousin and worth knowing even if you never buy a cluster. When one Availability Zone is impaired, you start a zonal shift on a supported load balancer and Route 53 stops handing out that AZ’s IP addresses, draining it in seconds; the shift auto-expires so you cannot forget to undo it. Zonal autoshift lets AWS perform the shift for you the moment it detects an AZ problem (you opt in, and AWS runs periodic practice shifts to prove your app tolerates losing an AZ). It is the fastest blast-radius control you have for a single-AZ event.
Traffic Flow, traffic policies, and geoproximity bias
The nested trees from the worked example are built in Traffic Flow, Route 53’s visual editor, which produces a versioned traffic policy — a document describing the whole routing tree (weighted under latency under failover, etc.). You do not attach a policy directly to a name; you create a policy record, which applies one version of the policy to a DNS name in a hosted zone. Versioning is the quiet benefit: roll a change forward by creating a new version, and roll back by re-pointing the policy record at the previous one. Two facts to budget for: geoproximity (and its bias, −99 to +99, that grows or shrinks a resource’s service area on the map) is only configurable through Traffic Flow or the API — never the simple record screen — and policy records cost around $50/month each, so express high-fan-out designs as one policy rather than dozens of hand-built records. A bias example: with resources in us-east-1 and us-west-2, dialling us-east-1 to +40 pushes the dividing line westward so more of the central US resolves east — a clean way to pre-drain a Region before maintenance without touching the app.
The registrar side: a domain is not a hosted zone
The two most-confused Route 53 objects are Registered domains and hosted zones — different resources with different lifecycles. Registering example.com through Route 53 auto-creates a public hosted zone and wires the domain’s NS records to it, which is why it feels like one thing; it is two. Consequences and features worth knowing:
- Deleting the hosted zone does not release the domain, and letting the domain expire does not delete the zone. Bill and lifecycle are independent.
- Auto-renew is on by default — good (domains lapsing is a classic outage) but confirm it for anything important. Expired domains pass through grace and redemption periods (redemption fees are steep) before they drop.
- Transfer lock (
clientTransferProhibited) blocks unauthorised transfers; WHOIS privacy masks your contact details and is free where the TLD allows it. To transfer a domain in, you must unlock it at the current registrar, disable privacy, and get the authorisation (EPP) code. - Domain registration is a global capability, but the
route53domainsAPI and its billing operate out ofus-east-1— target that Region when scripting registrations. Feature support (privacy, lock, transfer) varies by TLD, so check the specific extension rather than assuming.
Hands-on lab
We will create a hosted zone, add records, build a failover pair backed by a health check, and clean up. This uses Route 53 features that incur small charges (see the cost note); there is no perpetual free tier for hosted zones, but the cost of doing this for an hour is a few cents. You do not need to own a domain — we will create a zone and inspect it; you would only delegate a real domain at the registrar step.
1. Set a zone name and create a public hosted zone.
ZONE=kloudvin-lab-$RANDOM.example
aws route53 create-hosted-zone \
--name "$ZONE" \
--caller-reference "lab-$(date +%s)" \
--hosted-zone-config Comment="Route53 deep-dive lab"
Expected output includes a HostedZone.Id like /hostedzone/Z0123456789ABCDEFG and a DelegationSet.NameServers list of four ns-*.awsdns-* servers. Capture the ID:
ZID=$(aws route53 list-hosted-zones-by-name --dns-name "$ZONE" \
--query 'HostedZones[0].Id' --output text | sed 's#/hostedzone/##')
echo "$ZID"
2. View the auto-created NS and SOA records.
aws route53 list-resource-record-sets --hosted-zone-id "$ZID" \
--query "ResourceRecordSets[?Type=='NS' || Type=='SOA'].[Name,Type]" --output table
You should see the apex NS (four name servers) and the SOA — both created for you.
3. Add a simple A record.
cat > /tmp/r53-simple.json <<JSON
{ "Changes": [ {
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "www.$ZONE",
"Type": "A",
"TTL": 60,
"ResourceRecords": [ { "Value": "203.0.113.10" } ]
} } ] }
JSON
aws route53 change-resource-record-sets --hosted-zone-id "$ZID" \
--change-batch file:///tmp/r53-simple.json
The response shows a ChangeInfo.Status of PENDING. Route 53 changes are atomic and usually INSYNC within seconds.
4. Create a health check (endpoint, HTTPS to a known-good host) and a failover pair.
HCID=$(aws route53 create-health-check \
--caller-reference "hc-$(date +%s)" \
--health-check-config 'Type=HTTPS,FullyQualifiedDomainName=aws.amazon.com,Port=443,RequestInterval=30,FailureThreshold=3,ResourcePath=/' \
--query 'HealthCheck.Id' --output text)
echo "Health check: $HCID"
cat > /tmp/r53-failover.json <<JSON
{ "Changes": [
{ "Action": "UPSERT", "ResourceRecordSet": {
"Name": "app.$ZONE", "Type": "A", "TTL": 60,
"SetIdentifier": "primary",
"Failover": "PRIMARY",
"HealthCheckId": "$HCID",
"ResourceRecords": [ { "Value": "203.0.113.20" } ] } },
{ "Action": "UPSERT", "ResourceRecordSet": {
"Name": "app.$ZONE", "Type": "A", "TTL": 60,
"SetIdentifier": "secondary",
"Failover": "SECONDARY",
"ResourceRecords": [ { "Value": "198.51.100.30" } ] } }
] }
JSON
aws route53 change-resource-record-sets --hosted-zone-id "$ZID" \
--change-batch file:///tmp/r53-failover.json
5. Validate. Confirm both failover records exist and check the health-check status:
aws route53 list-resource-record-sets --hosted-zone-id "$ZID" \
--query "ResourceRecordSets[?Name=='app.$ZONE.'].[SetIdentifier,Failover,HealthCheckId]" \
--output table
aws route53 get-health-check-status --health-check-id "$HCID" \
--query 'HealthCheckObservations[].StatusReport.Status' --output table
You should see a primary/PRIMARY record bound to your health-check ID, a secondary/SECONDARY record, and several checker locations reporting Success: HTTP Status Code 200, OK. Because the records use placeholder documentation IPs, do not expect a real dig against them to reach a server — the point is the Route 53 configuration and the health-check signal.
6. Cleanup. Delete the records, the health check, and the zone (a zone with non-default records cannot be deleted):
# delete the failover records (Action must be DELETE with exact current values)
sed 's/"UPSERT"/"DELETE"/g' /tmp/r53-failover.json > /tmp/r53-failover-del.json
aws route53 change-resource-record-sets --hosted-zone-id "$ZID" \
--change-batch file:///tmp/r53-failover-del.json
sed 's/"UPSERT"/"DELETE"/g' /tmp/r53-simple.json > /tmp/r53-simple-del.json
aws route53 change-resource-record-sets --hosted-zone-id "$ZID" \
--change-batch file:///tmp/r53-simple-del.json
aws route53 delete-health-check --health-check-id "$HCID"
aws route53 delete-hosted-zone --id "$ZID"
Cost note. A hosted zone costs USD 0.50 per month (pro-rated only for the first 12 hours, then charged per month — so create and delete on the same day to keep it to the half-dollar). Standard queries are about USD 0.40 per million; Alias queries to AWS resources are free. Health checks of AWS endpoints are free; checks of non-AWS endpoints cost about USD 0.75 per check per month, with optional features (HTTPS, string matching, fast interval) adding small increments. This lab, deleted promptly, costs well under a dollar. Always run the cleanup — an orphaned hosted zone quietly bills USD 0.50 every month.
Common mistakes & troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Domain doesn’t resolve at all | Registrar still points at old/auto name servers, not the zone’s four | Copy the zone’s exact NS values into the registrar’s name-server settings; allow propagation. |
| “CNAME at apex not allowed” error | Tried to CNAME example.com to an AWS resource |
Use an Alias A/AAAA record at the apex instead. |
| Failover never switches | Primary record has no health check (or Alias Evaluate Target Health is off) |
Attach a health check to the primary, or enable Evaluate Target Health on the Alias. |
| Change made but users see old value for ages | TTL was high (e.g. 86,400) | Lower TTL before a planned change; for emergencies you can only wait out the cached TTL. |
| Geolocation users get no answer | No default record for unmatched locations | Add a geolocation record with location Default. |
| Private zone returns NXDOMAIN for a name that resolves publicly | Private zone is authoritative for that name space and lacks the record | Add the record to the private zone, or scope the private zone to a narrower name. |
| Health check flaps / always unhealthy | Firewall blocks Route 53 health-checker IPs, or path returns non-2xx/3xx | Allow the route53-healthchecks IP ranges; point the check at a real 200-returning health path. |
| Records in VPC don’t resolve from instances | VPC enableDnsSupport/enableDnsHostnames off, or zone not associated |
Enable both VPC attributes and associate the private zone with the VPC. |
Common beginner mistakes
These are misconceptions rather than symptoms — the wrong mental model that produces the outage in the first place. (The symptom→cause→fix table above is for when you are already staring at the failure.)
-
“DNS changes are instant.” They are not. You update the authoritative source in seconds, but the world only catches up as each resolver’s cached copy expires — governed by the record’s TTL, not by how fast you clicked Save. The right model: publishing a DNS change is starting a countdown, not finishing a switch. Lower the TTL a day before a planned change, then raise it again after.
-
“An Alias is just AWS’s word for a CNAME.” No. An Alias returns the target’s IP addresses directly, works at the apex, is free for AWS targets, and can carry Evaluate Target Health. A CNAME returns another name (a second lookup), is charged, must be alone at its record, and is banned at the apex. Reaching for a CNAME where an Alias belongs is the most common Route 53 design error.
-
“I’ll just CNAME
example.comto my load balancer.” The apex must also holdNSandSOA, and DNS forbids a CNAME coexisting with other records, so this is rejected. Use an Alias A/AAAA at the apex instead. -
“Weighted 10/90 means exactly 10% of my users hit green right now.” It means 10% of queries over time. Resolver caching pins each user to one answer for the TTL, and small or lopsided user populations come out lumpy. Keep TTLs low while shifting weights and judge the split over minutes, not seconds.
-
“Latency routing sends users to the geographically closest Region.” It sends them to the lowest measured latency from their resolver — which usually correlates with distance but is not the same thing, and is measured from the resolver, not the person. It ignores borders and compliance entirely; a user right by a national border can be routed across it. For residency, use geolocation and enforce it in the app tier too.
-
“I made a primary and a secondary, so failover will work.” Failover only triggers when the primary has a health check (or is an Alias with Evaluate Target Health = Yes). With no health signal Route 53 keeps returning the primary forever. Wire the health check, and make sure the secondary can survive whatever took the primary down.
-
“My private zone will fall through to public DNS for names it doesn’t have.” It will not. A private hosted zone is authoritative for its entire namespace, so a name it lacks returns NXDOMAIN to the VPC rather than resolving publicly. Either add the record to the private zone or scope the private zone to a narrower name.
-
“Deleting and recreating a hosted zone is harmless — same domain, same records.” Recreating a zone assigns a brand-new set of four name servers. The delegation at your registrar now points at servers that no longer host the zone, and the domain goes dark until you update the registrar. Treat a live zone’s name servers as wire-once-and-leave.
-
“A green health check means my application works.” It means the one thing you pointed the check at responded. A check against
/can stay green while the app is broken behind it. Point checks at a deep health path that exercises real dependencies, and add string matching to catch “HTTP 200 but wrong content.”
Best practices
- Always use Alias for AWS targets — apex-capable, free, faster, and health-aware via Evaluate Target Health.
- Keep failover-eligible records on low TTL (60 s) so recovery is fast; keep stable records (MX/TXT/NS) high to cut cost.
- Point health checks at a deep health endpoint that verifies real dependencies, with string matching to catch “200 but broken”.
- Always set a default record for geolocation/geoproximity policies.
- Manage zones as code (Terraform/CloudFormation) so records are reviewed and reproducible — manual record edits are a common outage source.
- Enable query logging (to CloudWatch Logs) on important public zones for security and debugging visibility.
- Use one delegated set of name servers — never recreate a live zone casually; if you must, update the registrar immediately.
- Add a
CAArecord listingamazon.comso only intended CAs (including ACM) can issue certificates for your domain.
Security notes
DNS is a security surface, not just plumbing. Enable DNSSEC signing on public zones to let resolvers cryptographically verify that answers are authentic and unmodified — this defends against DNS spoofing and cache poisoning (Route 53 supports DNSSEC signing with KMS-backed keys; you also add a DS record at the parent). Add CAA records to constrain certificate issuance to authorities you trust. Guard against dangling DNS / subdomain takeover: if a record points at a de-provisioned resource (a deleted S3 bucket or released Elastic IP), an attacker can claim that resource and serve content under your name — audit and remove records whose targets no longer exist. Turn on query logging to spot anomalous lookups (data-exfiltration tunnels, malware C2 patterns). Apply least-privilege IAM: scope route53:ChangeResourceRecordSets to specific hosted-zone ARNs so a compromised credential cannot rewrite every zone you own. Finally, for protecting outbound DNS from your VPCs (filtering what your workloads are allowed to resolve), use Route 53 Resolver DNS Firewall — covered in the resolver lesson linked below.
Interview & exam questions
-
What is the difference between an Alias record and a CNAME, and when must you use an Alias? A CNAME is standard DNS that returns another name (forcing a second lookup), is charged as a query, must be alone at its name, and cannot sit at the zone apex. An Alias is a Route 53 extension on an A/AAAA record that returns the target AWS resource’s IPs directly, is free for AWS targets, can coexist as a normal record, and works at the apex. You must use an Alias to point the apex (
example.com) at CloudFront, an ELB, S3 website, etc. -
You need to roll a new version of a service to 10% of users, then ramp up. Which routing policy? Weighted routing — give the new stack weight 10 and the old weight 90, then shift the weights. Keep TTL low so the proportions track reality and users re-resolve quickly.
-
A multi-Region app should serve every user from the fastest Region. Which policy, and what’s its limitation? Latency-based routing. Limitation: it optimises for measured network latency from the resolver, not the user, and ignores geography/compliance — a user near a border can be sent across it. For data-residency, use geolocation.
-
Failover routing isn’t switching to the secondary even though the primary is down. Why? The primary record almost certainly has no associated health check (and, if it’s an Alias, Evaluate Target Health is off). Without a health signal Route 53 keeps returning the primary. Attach a health check or enable Evaluate Target Health.
-
What are the three types of health check? Endpoint (probe an IP/domain over HTTP/HTTPS/TCP), calculated (Boolean combination of other health checks — “N of M healthy”), and CloudWatch alarm (mirror the state of any CloudWatch alarm, so you can fail over on metrics like error rate or queue depth).
-
Explain split-horizon DNS in Route 53 and one gotcha. Run a private hosted zone and a public hosted zone for the same domain; queries from associated VPCs hit the private zone, the internet hits the public one. Gotcha: the private zone is authoritative for its whole name space, so a name it lacks returns NXDOMAIN to the VPC rather than falling through to public resolution.
-
What does TTL control, and what’s a safe value for a record you might need to fail over? TTL is how long resolvers cache an answer before re-querying. For failover-eligible records use a low TTL (~60 s) so a failover propagates quickly; the extra query cost is negligible versus recovery time.
-
Why can’t you put a CNAME at
example.com? The apex must carry the zone’sNSandSOArecords, and DNS forbids a CNAME from coexisting with any other record at the same name. Route 53’s Alias record solves this because it is technically an A/AAAA record. -
Geolocation vs geoproximity — what’s the difference? Geolocation routes by the user’s named location (continent/country/US state) with a default fallback. Geoproximity routes by distance between user and resource and lets you apply a bias to expand or shrink each resource’s service area; it requires Traffic Flow.
-
What is multivalue answer routing and when would you use it over a load balancer? It returns up to eight random, optionally health-checked values, omitting unhealthy ones — health-aware DNS spreading. Use it for a small set of independent endpoints when you want availability without an LB; it is not a true load balancer (no connection awareness, no even distribution).
-
How does Route 53 decide an endpoint is healthy, and why does it matter for firewalls? Health checkers in multiple global locations probe the endpoint; it’s healthy if more than 18% report success. You must therefore allow the Route 53 health-checker IP ranges through security groups/NACLs/firewalls, or checks fail and traffic drains erroneously.
-
What is subdomain takeover and how do you prevent it? A “dangling” DNS record points at a de-provisioned resource (deleted bucket, released EIP); an attacker re-creates/claims that resource and serves content under your name. Prevent it by removing records whose targets no longer exist and auditing zones regularly.
Quick check
- Which record type points a name at an IPv6 address?
- True/false: you can place a CNAME at the zone apex if it is the only record there.
- Which routing policy is the right choice for active-passive disaster recovery?
- What happens to a geolocation query that matches no rule and has no default record?
- Are Alias queries to an AWS resource charged?
Answers
- AAAA.
- False — a CNAME is never allowed at the apex (the apex must hold
NS/SOA); use an Alias. - Failover routing (primary + secondary, with a health check on the primary).
- It returns no answer (NODATA) — always configure a Default record for geolocation.
- No — Alias queries that resolve to AWS resources are free; standard queries are charged per million.
Exercise
Design the DNS for a two-Region web application (eu-west-1 and us-east-1) fronted by an ALB in each Region, that must (a) serve every user from the lower-latency Region, (b) fail a Region out automatically when its ALB has no healthy targets, and © serve EU users only from eu-west-1 for data-residency. Sketch the records: which routing policies, how they nest, what each record’s Alias target and health configuration are, and what TTLs you’d set. Then write the aws route53 change-resource-record-sets change-batch JSON for the apex records. (Hint: geolocation at the top to honour residency, latency for the rest, Alias-to-ALB with Evaluate Target Health for the per-Region failover, 60 s TTLs.)
Practice challenges
Six exercises, easy to hard. Try each before opening the solution. Commands use placeholder IDs and documentation IP ranges; nothing here needs a real domain, though a few incur the small charges noted in the lab.
1. (Beginner) Prove that a new zone comes with NS and SOA. Create a public hosted zone for shop.example and list only its auto-created NS and SOA records.
<details> <summary><strong>Solution</strong></summary>
ZID=$(aws route53 create-hosted-zone --name shop.example \
--caller-reference "ex1-$(date +%s)" \
--query 'HostedZone.Id' --output text | sed 's#/hostedzone/##')
aws route53 list-resource-record-sets --hosted-zone-id "$ZID" \
--query "ResourceRecordSets[?Type=='NS' || Type=='SOA'].[Name,Type]" --output table
Why: every zone is born with four authoritative name servers (NS) and one SOA; those NS values are what you delegate to at the registrar, and you never delete them.
</details>
2. (Beginner) Put a record at the apex where CNAME is illegal. Point the apex shop.example at a CloudFront distribution d111111abcdef8.cloudfront.net.
<details> <summary><strong>Solution</strong></summary>
{ "Changes": [ {
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "shop.example", "Type": "A",
"AliasTarget": {
"HostedZoneId": "Z2FDTNDATAQYW2",
"DNSName": "d111111abcdef8.cloudfront.net",
"EvaluateTargetHealth": false
} } } ] }
Z2FDTNDATAQYW2 is CloudFront’s fixed alias hosted-zone ID (the same for every distribution). Why: the apex must hold NS/SOA, so DNS forbids a CNAME there; an Alias A returns the resource’s IPs directly, works at the apex, and is free.
</details>
3. (Intermediate) Ship a 90/10 canary, then pause it. Create weighted records for api.shop.example sending 10% to green (203.0.113.50) and 90% to blue (203.0.113.10), TTL 30. Then take green out of rotation without deleting it.
<details> <summary><strong>Solution</strong></summary>
{ "Changes": [
{ "Action": "UPSERT", "ResourceRecordSet": {
"Name": "api.shop.example", "Type": "A", "TTL": 30,
"SetIdentifier": "blue", "Weight": 90,
"ResourceRecords": [ { "Value": "203.0.113.10" } ] } },
{ "Action": "UPSERT", "ResourceRecordSet": {
"Name": "api.shop.example", "Type": "A", "TTL": 30,
"SetIdentifier": "green", "Weight": 10,
"ResourceRecords": [ { "Value": "203.0.113.50" } ] } }
] }
To pause the canary, UPSERT the green record with "Weight": 0. Why: the returned share is weight ÷ sum, so 10 ÷ 100 = 10% of queries; weight 0 removes a record from rotation while keeping it defined for an instant resume.
</details>
4. (Intermediate) Localise with no NODATA holes. Configure geolocation so users in Europe get eu (198.51.100.20) and everyone else gets global (198.51.100.99).
<details> <summary><strong>Solution</strong></summary>
{ "Changes": [
{ "Action": "UPSERT", "ResourceRecordSet": {
"Name": "www.shop.example", "Type": "A", "TTL": 60,
"SetIdentifier": "eu", "GeoLocation": { "ContinentCode": "EU" },
"ResourceRecords": [ { "Value": "198.51.100.20" } ] } },
{ "Action": "UPSERT", "ResourceRecordSet": {
"Name": "www.shop.example", "Type": "A", "TTL": 60,
"SetIdentifier": "global", "GeoLocation": { "CountryCode": "*" },
"ResourceRecords": [ { "Value": "198.51.100.99" } ] } }
] }
Why: CountryCode: "*" is the Default record; without it, any location that matches no rule (most of the planet, here) gets no answer at all.
</details>
5. (Advanced) Fastest Region, self-healing, no separate health-check object. Sketch the apex records for a two-Region app that serves every user from the lowest-latency Region and drains a Region automatically when its ALB has zero healthy targets.
<details> <summary><strong>Solution</strong></summary>
Two latency records for the apex, one per Region, each an Alias to that Region’s ALB with EvaluateTargetHealth: true — no standalone health check needed:
{ "Changes": [
{ "Action": "UPSERT", "ResourceRecordSet": {
"Name": "app.shop.example", "Type": "A",
"SetIdentifier": "eu", "Region": "eu-west-1",
"AliasTarget": { "HostedZoneId": "Z32O12XQLNTSW2",
"DNSName": "dualstack.eu-alb-123.eu-west-1.elb.amazonaws.com",
"EvaluateTargetHealth": true } } },
{ "Action": "UPSERT", "ResourceRecordSet": {
"Name": "app.shop.example", "Type": "A",
"SetIdentifier": "us", "Region": "us-east-1",
"AliasTarget": { "HostedZoneId": "Z35SXDOTRQ7X7K",
"DNSName": "dualstack.us-alb-456.us-east-1.elb.amazonaws.com",
"EvaluateTargetHealth": true } } }
] }
The HostedZoneId in each AliasTarget is the ELB zone ID for that Region (per-Region constants, not your zone). Why: latency chooses the Region; EvaluateTargetHealth makes Route 53 stop returning a Region whose ALB has no healthy targets — so the latency policy silently omits the dead Region. Add a maintenance page by wrapping each leaf in a failover pair via Traffic Flow.
</details>
6. (Advanced) Sign a zone, and describe how to turn it off safely. Outline the steps to enable DNSSEC on a public zone ZID, then the correct order to disable it.
<details> <summary><strong>Solution</strong></summary>
# 1) asymmetric KMS key in us-east-1, ECC_NIST_P256, with a policy granting
# dnssec-route53.amazonaws.com kms:DescribeKey/GetPublicKey/Sign
KID=arn:aws:kms:us-east-1:123456789012:key/EXAMPLE-KEY-ID
# 2) create the Key Signing Key and enable signing
aws route53 create-key-signing-key --caller-reference "ksk-$(date +%s)" \
--hosted-zone-id "$ZID" --key-management-service-arn "$KID" \
--name shopKSK --status ACTIVE
aws route53 enable-hosted-zone-dnssec --hosted-zone-id "$ZID"
# 3) copy the DS record Route 53 shows you into the PARENT (registrar/TLD)
Disable order (the part that bites): remove the DS record at the parent first, wait out its TTL so caches forget it, and only then run disable-hosted-zone-dnssec. Why: if signing stops while a DS still points at your zone, every validating resolver returns SERVFAIL for the entire domain — a self-inflicted outage that stale caches keep serving.
</details>
Certification mapping
- AWS Certified Solutions Architect – Associate (SAA-C03): routing policies and when to use each, Alias vs CNAME, failover and health checks, latency-based routing for multi-Region, private vs public hosted zones — heavily tested.
- AWS Certified Advanced Networking – Specialty (ANS-C01): deep DNS — split-horizon, DNSSEC, geoproximity/Traffic Flow, calculated and CloudWatch health checks, hybrid resolution (with the Resolver lesson), and health-checker IP allow-listing.
- Touches SOA-C02 (operating failover and health checks) and DVA-C02 (pointing app endpoints at AWS resources).
Glossary
- Hosted zone — Route 53 container for one domain’s records; gives you four authoritative name servers.
- Authoritative server — the server that holds the real records for a zone (Route 53). Distinct from a recursive resolver, which chases and caches answers.
- Record set — one DNS entry: name, type, TTL, value (or Alias target).
- Alias — a Route 53-only A/AAAA record that returns an AWS resource’s IPs directly; apex-capable, free, health-aware.
- CNAME — standard record aliasing one name to another; not allowed at the apex.
- Routing policy — the rule deciding which record value Route 53 returns (simple, weighted, latency, failover, geolocation, geoproximity, multivalue).
- Health check — a Route 53 object monitoring a target (endpoint, calculated, or CloudWatch-alarm) that gates whether a record is returned.
- Evaluate Target Health — an Alias setting that inherits the health of the AWS target automatically.
- TTL — seconds a resolver may cache an answer before re-querying.
- Split-horizon DNS — same domain served differently to VPCs (private zone) and the internet (public zone).
- Delegation — pointing a parent at a child zone’s name servers via
NSrecords. - DNSSEC — cryptographic signing of DNS answers so resolvers can verify authenticity.
- IP-based routing — the eighth routing policy: return a record based on the client/resolver source IP matching a CIDR location you define, rather than an inferred geography.
- CIDR collection / CIDR location — a named set of IP ranges (collection) grouped into buckets (locations) that IP-based routing records point at; use location
*as the catch-all. - EDNS Client Subnet (ECS) — a DNS extension where a resolver forwards a truncated slice of the client’s subnet, letting Route 53 target the real user’s network instead of the resolver’s; supported by some public resolvers, not all.
- Traffic Flow — Route 53’s visual editor for building nested routing trees; it produces versioned traffic policies.
- Traffic policy / policy record — a versioned routing-tree document (traffic policy) applied to a DNS name by creating a policy record; versions let you roll changes forward and back. Policy records carry a monthly charge.
- Geoproximity bias — a value from −99 to +99 that expands or shrinks a resource’s geographic service area under geoproximity routing; configurable only via Traffic Flow or the API.
- Route 53 Resolver — the recursive DNS service inside a VPC (at the
.2/169.254.169.253address); distinct from authoritative hosted zones. - Inbound / outbound Resolver endpoint — ENIs that let on-prem query AWS private zones (inbound) or let a VPC forward chosen domains to external resolvers (outbound) for hybrid DNS.
- Forwarding rule (FORWARD / SYSTEM) — a Resolver rule that forwards queries for a domain to target IPs (FORWARD) or restores default resolution for a sub-name (SYSTEM); shareable across accounts via AWS RAM.
- Route 53 ARC (Application Recovery Controller) — recovery orchestration: readiness checks, routing controls (on/off switches backed by health checks), and safety rules, run from a five-Region cluster.
- Zonal shift / zonal autoshift — temporarily drain traffic from one impaired Availability Zone on a supported load balancer (manual shift), or let AWS do it automatically on detected AZ impairment (autoshift).
- KSK / ZSK — DNSSEC keys: the Key Signing Key (you supply it via an asymmetric KMS key) signs the keyset; the Zone Signing Key (Route 53-managed) signs the records.
- DS record — a hash of your KSK published in the parent zone to establish the DNSSEC chain of trust; remove it before disabling signing.
- Negative caching — how long a “no such name” answer is cached, set by the lesser of the SOA record’s TTL and its minimum field; a typo returning NXDOMAIN can linger in caches after the fix.
- NXDOMAIN vs NODATA — NXDOMAIN means the name does not exist; NODATA means the name exists but has no record of the requested type (what an unmatched geolocation query with no default returns).
- Registered domain vs hosted zone — two separate Route 53 objects with independent lifecycles: the domain registration (registrar side) and the zone that holds its records; deleting one does not affect the other.
- Delegation set — the group of four name servers Route 53 assigns a zone; recreating a zone yields a new set, breaking any existing registrar delegation.
Next steps
- Build the resilient global front door these records feed into: Global Edge Architecture with CloudFront and Route 53: Failover Routing, Origin Shielding, and WAF Protection (
cloudfront-route53-global-edge-failover-waf-origin-protection). - Go deeper on resolving and filtering DNS inside your VPCs and across hybrid networks: Route 53 Resolver: DNS Firewall, Endpoints, Rules & Hybrid Resolution (
route53-resolver-dns-firewall-endpoints-rules-hybrid-resolution). - Continue the Networking module with observability: AWS Observability, In Depth: CloudWatch, CloudTrail, Config & EventBridge (
aws-cloudwatch-cloudtrail-observability-deep-dive).