In a nutshell
A VPC is your own private network inside AWS — think of it as leasing an empty office building in one AWS Region. AWS hands you the building and a block of internal extensions (your IP address range); you decide how to divide the floors into rooms (subnets), which doors open onto the street (an Internet Gateway), which rooms may call out but can never be called (private subnets behind a NAT Gateway), and which corridors connect one room to another (route tables). Nothing in AWS that uses the network — a server, a database, a load balancer — ever floats free; it always sits in a room on one of your floors.
The single idea that unlocks everything else: a subnet is “public” only because a route sends its internet-bound traffic to an Internet Gateway. There is no “make this public” switch — being public is a consequence of routing, exactly the way a room reaches the street only if a corridor actually leads to an exit. Internalise that one sentence and the rest of this lesson — reserved IPs, NAT, endpoints, peering — becomes a set of predictable follow-ons instead of a pile of disconnected features.
Why a beginner should care: almost every “my instance can’t reach the internet”, “my database is unexpectedly exposed”, or “I can’t connect these two networks” problem in AWS is really a VPC-design problem. Get the address plan and the routes right once and everything you build on top just works; get them wrong and you re-IP networks for years. This lesson is long because the VPC is the foundation — but every term is defined the moment it appears, so you can read straight through.
Level: Intermediate (beginner-accessible) · Time: ~50 min · You need: an AWS account and a rough idea of what an IP address is — CIDR, routing, and NAT are all built up from first principles here. After it you can: plan a non-overlapping CIDR, place public/private/isolated subnets across Availability Zones, wire an IGW plus a per-AZ NAT, add free S3 gateway endpoints, and read a route table like a map.
Every workload you run in AWS that touches the network — an EC2 instance, an RDS database, a Lambda function reaching a private API, a load balancer — lives inside a Virtual Private Cloud (VPC): your own logically isolated, software-defined slice of the AWS network where you choose the IP address range, carve it into subnets across Availability Zones, and decide exactly what can reach the internet, what stays private, and how packets are routed. Get the design right and everything downstream just works. Get it wrong and you feel it for years: you run out of addresses mid-migration, you cannot peer two VPCs because their ranges overlap, traffic that should stay private egresses through a NAT gateway you are paying for by the gigabyte, or a “private” instance silently has a route to the internet.
This is the exhaustive lesson. We go component by component — the VPC CIDR and how to add IPv6 and secondary ranges, every field on a subnet and the five IP addresses AWS reserves in each one, route tables and the immovable local route, the Internet Gateway, the long-running argument of NAT Gateway versus a self-managed NAT instance, DHCP option sets, the two DNS attributes that break name resolution when they are off, the difference between gateway and interface (PrivateLink) endpoints, where peering ends and Transit Gateway begins, and VPC Flow Logs — until you can whiteboard a production VPC from memory and answer the follow-up questions a Solutions Architect interview or the SAA-C03 and ANS-C01 exams will throw at you. It is beginner-accessible — every term is defined as it appears — but complete: read it once and you know the service end to end.
Learning objectives
By the end of this lesson you will be able to:
- Plan a VPC CIDR block with room to grow, add secondary IPv4 ranges, and enable IPv6, understanding what you can and cannot change after creation.
- Size and place subnets across Availability Zones, classify them as public or private by their routing, and account for the five reserved IP addresses AWS takes in every subnet.
- Build route tables correctly — the main vs custom distinction, the unremovable
localroute, the0.0.0.0/0default route, and route priority (longest-prefix match). - Attach an Internet Gateway and reason about what actually makes a subnet public.
- Choose between a NAT Gateway and a NAT instance for outbound-only internet access, and size both for cost and throughput.
- Configure DHCP option sets and the
enableDnsSupport/enableDnsHostnamesattributes, and explain what each one controls. - Decide between gateway endpoints (S3, DynamoDB) and interface endpoints / PrivateLink, and know when each saves money or is required.
- Connect VPCs with peering and know where Transit Gateway takes over, and turn on VPC Flow Logs for visibility.
Prerequisites & where this fits
You need an AWS account and the basics of regions, Availability Zones, and the CLI/console from the earlier Fundamentals lessons, plus a working idea of what an IP address and a subnet are. No deep networking background is assumed — CIDR, routing, and NAT are all explained from first principles. This is the opening Networking deep-dive of the AWS Zero-to-Hero course and the foundation that every later networking lesson builds on. The very next lesson, AWS Security Groups vs Network ACLs, In Depth, covers the filtering layer that sits on top of the routing layer you design here; this lesson deliberately stays on addressing, routing, and connectivity, and points you there for firewalls. When your address planning outgrows a spreadsheet, Amazon VPC IPAM: Hierarchical CIDR Planning, Allocation, and BYOIP at Scale automates it; when one VPC becomes dozens, Designing Multi-Account VPC Connectivity with Transit Gateway replaces the peering mesh.
Core concepts
A VPC (Virtual Private Cloud) is a regional resource — it spans every Availability Zone in one AWS Region but cannot cross Regions — that defines a private IPv4 address range (and optionally IPv6) which is yours alone. Inside it you build a network using a small set of primitives that fit together predictably. Anchor everything that follows on these mental models:
- The VPC is the building; subnets are the floors. You give the building an address space (e.g.
10.0.0.0/16) and partition it into subnets (10.0.1.0/24,10.0.2.0/24, …). Every network interface — and therefore every instance, database, or load balancer node — attaches to a subnet, never to the VPC directly. - A subnet lives in exactly one Availability Zone. This is the single most important fact for designing high availability: to survive an AZ failure you need at least two subnets in two different AZs, and you place a copy of your workload in each.
- Routing decides “public” vs “private”, not the subnet itself. A subnet is “public” only because its route table sends
0.0.0.0/0to an Internet Gateway. There is no checkbox called “public”; it is a property of the routes. - Everything inside is reachable by default; the edges are controlled. Every subnet in a VPC can reach every other subnet via the built-in
localroute. What crosses the edge — to the internet, to another VPC, to on-premises — is what you explicitly enable with gateways and routes. - Filtering is a separate layer. Routing gets a packet to a destination; security groups (stateful, on the network interface) and network ACLs (stateless, on the subnet) decide whether it is allowed. They are covered in the next lesson — keep them mentally separate from routing.
Key terms you will see throughout: CIDR (Classless Inter-Domain Routing — the /16, /24 notation that defines how many addresses a block holds and how the prefix is split between network and host), ENI (Elastic Network Interface — the virtual NIC that everything in a VPC actually attaches to), IGW (Internet Gateway — the VPC’s door to the public internet), NAT (Network Address Translation — letting many private addresses share one public address for outbound traffic), route table (the ordered set of rules that decides where a packet goes next), and endpoint (a private on-ramp to an AWS service that keeps traffic off the internet).
Default VPC vs a custom VPC
Every Region in a new account comes with a default VPC so that you can launch an instance immediately without designing a network first. Understanding what makes it “default” tells you what a custom VPC does not give you for free.
| Property | Default VPC | Custom VPC (one you create) |
|---|---|---|
| CIDR | 172.31.0.0/16, fixed |
You choose |
| Subnets | One default subnet per AZ, all public | None until you create them |
| Internet Gateway | Created and attached | You attach it yourself |
Route to 0.0.0.0/0 |
Present in the main route table → IGW | You add it |
| Public IP on launch | Auto-assign public IPv4 = on in default subnets | Off by default |
| DNS hostnames | Enabled | Disabled by default |
| Good for | Quick demos, getting started | Everything real — explicit control |
The convenience of the default VPC is also its danger: every default subnet is public and auto-assigns a public IP, so an instance launched there is internet-reachable the moment a permissive security group is attached. For anything beyond a throwaway test, build a custom VPC where nothing is public unless you deliberately route it that way. You can delete the default VPC, and recreate it later from the console if you ever need it back.
VPC CIDR: primary, secondary, and IPv6
When you create a VPC the one truly load-bearing decision is the primary IPv4 CIDR block. It defines the pool of private addresses every subnet will be carved from, and it cannot be changed or removed for the life of the VPC — you can only add secondary blocks.
| Setting | What it is | Choices / limits | Default | When to change / gotcha |
|---|---|---|---|---|
| Primary IPv4 CIDR | The main private address range | /16 (65,536 addresses) down to /28 (16 addresses); use RFC 1918 private ranges |
None — required | Permanent. Pick a /16 for production so subnets have room. Cannot overlap with any network you will peer or connect to on-prem. |
| Secondary IPv4 CIDRs | Extra ranges added later when you run out | Up to 5 by default (raise to ~50 via quota); must not overlap existing blocks or reserved AWS ranges | None | Add when subnets fill up. Cannot fall inside an existing block; choose from the same private range family to keep routing sane. |
| IPv6 CIDR | An optional /56 block |
Amazon-provided (you get a /56, subnets are /64) or your own (BYOIP) |
Off | Enable for IPv6 workloads or to use egress-only internet gateways. IPv6 addresses are public and globally routable — there is no “private IPv6” in the RFC 1918 sense. |
Two rules save most teams from pain. First, size for the whole estate, not today’s app — a /16 costs nothing extra over a /28 (you pay for traffic and resources, never for address space), and running out of contiguous space later forces ugly secondary-CIDR workarounds. Second, never reuse the same CIDR across VPCs you might connect. If vpc-a and vpc-b are both 10.0.0.0/16, you can never peer them or attach them to the same Transit Gateway — overlapping ranges have no unambiguous route. Allocate a unique block per VPC up front; when this becomes hard to track by hand, that is exactly the problem VPC IPAM solves.
# Create a custom VPC with a /16 primary block
aws ec2 create-vpc \
--cidr-block 10.0.0.0/16 \
--tag-specifications 'ResourceType=vpc,Tags=[{Key=Name,Value=vpc-lab}]'
# Add a secondary IPv4 block later
aws ec2 associate-vpc-cidr-block --vpc-id vpc-0abc... --cidr-block 10.1.0.0/16
# Add an Amazon-provided IPv6 /56
aws ec2 associate-vpc-cidr-block --vpc-id vpc-0abc... --amazon-provided-ipv6-cidr-block
Subnets: public vs private, AZ placement, sizing, and reserved IPs
A subnet is a sub-range of the VPC CIDR that lives in exactly one Availability Zone. Resources attach to subnets, and the subnet’s route table determines whether it is public or private.
| Setting | What it is | Choices | Default | When / trade-off / gotcha |
|---|---|---|---|---|
| VPC | The parent network | Any VPC in the Region | — | The subnet’s CIDR must fall inside the VPC’s CIDR. |
| Availability Zone | Physical location of the subnet | Any AZ in the Region | AWS picks if unspecified | Pin it explicitly and spread workloads across ≥2 AZs for HA. A subnet cannot span or move AZs. |
| IPv4 CIDR block | The subnet’s address range | /16 to /28 within the VPC |
Required | /24 (256 addresses) is a comfortable default. Smaller than /28 is not allowed because of reserved IPs. |
| IPv6 CIDR | Optional /64 from the VPC’s /56 |
One /64 per subnet |
None | Required if the subnet hosts IPv6 resources. |
| Auto-assign public IPv4 | Give launched instances a public IP automatically | On / Off | Off | Turning this on is what people mean by a “public subnet” in practice — but it only matters alongside a route to an IGW. |
| Auto-assign IPv6 | Auto-assign an IPv6 address on launch | On / Off | Off | Enable for IPv6 subnets. |
Public vs private is purely about routing. A subnet is public when its route table has a 0.0.0.0/0 route pointing at an Internet Gateway (and, in practice, auto-assign public IP is on or instances carry Elastic IPs). It is private when it has no such route — instances reach the internet only outbound via a NAT gateway, or not at all. A common third tier is an isolated subnet with no internet route in either direction (for databases), reachable only inside the VPC and via endpoints.
The five reserved IP addresses
AWS reserves the first four and the last IP address in every subnet, so a /24 (256 addresses) gives you 251 usable, not 256. Memorise this — it is a classic exam question and it bites IP planning.
Address (in 10.0.1.0/24) |
Reserved for |
|---|---|
10.0.1.0 |
Network address |
10.0.1.1 |
VPC router (the implied default gateway) |
10.0.1.2 |
Amazon-provided DNS (the “.2 resolver” — VPC base +2) |
10.0.1.3 |
Reserved for future use |
10.0.1.255 |
Network broadcast (broadcast is not supported, but the address is still reserved) |
Because five addresses always disappear, the smallest permitted subnet is a /28 (16 addresses → 11 usable). The .2 resolver in particular matters later: it is the address the Amazon DNS server answers on, and several DNS features depend on it.
# Two subnets in two different AZs (HA), with auto-assign public IP on for the first
aws ec2 create-subnet --vpc-id vpc-0abc... --cidr-block 10.0.1.0/24 \
--availability-zone ap-south-1a \
--tag-specifications 'ResourceType=subnet,Tags=[{Key=Name,Value=public-1a}]'
aws ec2 modify-subnet-attribute --subnet-id subnet-0pub... --map-public-ip-on-launch
aws ec2 create-subnet --vpc-id vpc-0abc... --cidr-block 10.0.11.0/24 \
--availability-zone ap-south-1b \
--tag-specifications 'ResourceType=subnet,Tags=[{Key=Name,Value=private-1b}]'
Route tables: main vs custom, the local route, and priority
A route table is an ordered set of rules — destination CIDR → target — that the VPC router consults for every packet leaving a network interface. Each subnet is associated with exactly one route table at a time; if you do not associate one explicitly, the subnet uses the VPC’s main route table.
| Concept | What it is | Detail / gotcha |
|---|---|---|
| Main route table | The default table every new subnet implicitly uses | One per VPC; you can edit it, but the safer pattern is to leave it minimal (private) and attach custom tables to public subnets. |
| Custom route table | A table you create and explicitly associate with subnets | The recommended way to define “public” vs “private” — one custom table per tier. |
local route |
An automatic route for the entire VPC CIDR with target local |
Always present, cannot be deleted or edited. It is why every subnet can reach every other subnet with zero configuration. |
0.0.0.0/0 (IPv4) / ::/0 (IPv6) |
The “default route” — everything not matched elsewhere | Point it at an IGW (public), a NAT gateway (private outbound), a Transit Gateway, a peering connection, or an egress-only IGW (IPv6). |
| Subnet associations | Which subnets use this table | A subnet has one table; a table can serve many subnets. |
| Route propagation | Auto-learn routes from a VPN/Direct Connect gateway via BGP | Toggle per route table; avoids hand-maintaining on-prem prefixes. |
Route priority is longest-prefix match. When several routes could apply, the VPC router picks the most specific (longest prefix) one. A packet to 10.0.5.7 matches 10.0.0.0/16 → local over 0.0.0.0/0 → nat because /16 is more specific than /0. The local route therefore always wins for in-VPC traffic, which is precisely why you cannot accidentally route internal traffic out to the internet. Static routes beat propagated (BGP-learned) routes of the same prefix.
# A public route table: default route to the Internet Gateway, associated to the public subnet
aws ec2 create-route-table --vpc-id vpc-0abc... \
--tag-specifications 'ResourceType=route-table,Tags=[{Key=Name,Value=rtb-public}]'
aws ec2 create-route --route-table-id rtb-0pub... \
--destination-cidr-block 0.0.0.0/0 --gateway-id igw-0xyz...
aws ec2 associate-route-table --route-table-id rtb-0pub... --subnet-id subnet-0pub...
Internet Gateway: the door to the public internet
An Internet Gateway (IGW) is a horizontally scaled, redundant, highly available VPC component that allows communication between instances in your VPC and the internet. It does two jobs: it provides a target in your route tables for internet-bound traffic, and it performs one-to-one NAT between an instance’s private IPv4 address and its public IPv4 address (or Elastic IP).
Three conditions must all be true for an instance to be reachable from the internet over IPv4 — miss any one and connectivity silently fails, which is the most common “why can’t I reach my instance” support question:
- An IGW is attached to the VPC.
- The subnet’s route table has
0.0.0.0/0→ the IGW. - The instance has a public IPv4 address (auto-assigned, or an Elastic IP) and its security group / network ACL allow the traffic.
Key facts: a VPC can have only one IGW attached at a time; the IGW itself is free (you pay for data transfer and, since 2024, for public IPv4 addresses); and for IPv6 there is a separate egress-only internet gateway that allows outbound IPv6 only (the IPv6 equivalent of a NAT gateway, since IPv6 has no NAT).
aws ec2 create-internet-gateway \
--tag-specifications 'ResourceType=internet-gateway,Tags=[{Key=Name,Value=igw-lab}]'
aws ec2 attach-internet-gateway --internet-gateway-id igw-0xyz... --vpc-id vpc-0abc...
NAT Gateway vs NAT instance: outbound-only internet
Instances in a private subnet often still need outbound internet — to download patches, call a third-party API, or pull a container image — without being reachable inbound. That is Network Address Translation (NAT): many private addresses share one public address for outbound flows, and return traffic for those flows is allowed back, but nothing can initiate a connection to the private instances from outside.
You route the private subnet’s 0.0.0.0/0 to the NAT, and the NAT itself sits in a public subnet (it needs the IGW to reach the internet). There are two ways to provide NAT:
| Dimension | NAT Gateway (managed) | NAT instance (self-managed EC2) |
|---|---|---|
| What it is | A fully managed AWS service | An EC2 instance running NAT software |
| Availability | Highly available within one AZ; deploy one per AZ for zone resilience | Single instance = single point of failure; you build HA yourself |
| Throughput | Scales automatically 5 → 100 Gbps | Bounded by the instance type’s network/CPU |
| Management | Zero — no patching, no sizing | You patch, monitor, and size it |
| Source/dest check | N/A | Must disable source/destination check or it will not forward |
| Security groups | Cannot attach an SG (control via NACL / the private route) | Has a security group like any instance |
| Port forwarding / bastion | Not possible | Possible (it is a normal instance) |
| Cost | Hourly charge + per-GB data processing | Just the EC2 instance (often a small/free-tier type) |
| Use it when | Almost always — the default | Cost-sensitive dev/test, or you need features only an instance gives |
Default to the NAT Gateway — it is the managed, scalable, low-effort choice and the right answer in virtually every exam scenario. The two things to know cold: it is zonal, so a truly resilient design places one NAT gateway in each AZ and points each AZ’s private subnets at the NAT gateway in their own AZ (this also avoids paying cross-AZ data transfer); and its bill has two parts — an hourly rate plus a per-GB data-processing charge — which is exactly why pulling large objects from S3 through a NAT gateway is wasteful when a free gateway endpoint would keep that traffic off the NAT entirely (see the next section).
A NAT instance is the legacy approach. The detail interviewers love is that, because the instance forwards traffic for other hosts, you must disable the source/destination check (aws ec2 modify-instance-attribute --no-source-dest-check) — by default an instance drops packets whose source or destination is not itself.
# Allocate an Elastic IP and create a NAT gateway in the PUBLIC subnet
aws ec2 allocate-address --domain vpc # returns an AllocationId
aws ec2 create-nat-gateway --subnet-id subnet-0pub... --allocation-id eipalloc-0...
# Point the PRIVATE subnet's default route at the NAT gateway
aws ec2 create-route --route-table-id rtb-0priv... \
--destination-cidr-block 0.0.0.0/0 --nat-gateway-id nat-0...
DHCP option sets
When an instance boots, it gets its network configuration — DNS servers, domain name, NTP servers — via DHCP, and a DHCP option set is the VPC-level object that defines those values. Every VPC has one associated; the default one points at AmazonProvidedDNS and is fine for most cases.
| Option | What it controls | Default | When to change |
|---|---|---|---|
domain-name-servers |
Which DNS resolvers instances use | AmazonProvidedDNS (the .2 resolver) |
Point at custom resolvers (e.g. on-prem AD DNS, or Route 53 Resolver inbound endpoints) for hybrid name resolution. |
domain-name |
The domain suffix applied to hostnames | Region-specific (e.g. ap-south-1.compute.internal) |
Set a corporate suffix like corp.example.com. |
ntp-servers |
Time servers | Amazon Time Sync (169.254.169.123) |
Override only if you have a specific NTP requirement. |
netbios-name-servers / netbios-node-type |
Legacy Windows NetBIOS | None | Rarely needed; set node-type=2 for Windows estates that use it. |
The important gotchas: you cannot edit an option set in place — you create a new one and associate it with the VPC. After re-associating, existing instances pick up the change only when their DHCP lease renews (or on reboot), so do not expect it to take effect instantly. Replacing AmazonProvidedDNS with custom servers is the usual reason to touch this, and it is how you wire VPC DNS into a hybrid Active Directory environment.
DNS in the VPC: enableDnsSupport and enableDnsHostnames
Two VPC attributes control DNS, and confusing them is a perennial source of “my private endpoint resolves to a public IP” tickets. Both default differently for default vs custom VPCs.
| Attribute | What it does | Default (custom VPC) | If turned off |
|---|---|---|---|
enableDnsSupport |
Whether the Amazon DNS resolver (the .2 address) answers queries in the VPC |
On | Instances cannot resolve names via the AWS resolver; DNS-based features (including private DNS for endpoints) break. |
enableDnsHostnames |
Whether instances with a public IP get a public DNS hostname auto-assigned | Off | Instances get no public DNS name; private DNS names for interface endpoints will not resolve even if support is on. |
The rule to remember: enableDnsSupport must be on for any DNS to work at all, and enableDnsHostnames must also be on for interface (PrivateLink) endpoints’ private DNS names to resolve to the endpoint’s private IP. Both default to on in the default VPC and (in modern accounts) enableDnsSupport is on but enableDnsHostnames is off in custom VPCs — so when you adopt PrivateLink and find s3.ap-south-1.amazonaws.com still resolving to a public IP, the fix is almost always to turn on enableDnsHostnames.
aws ec2 modify-vpc-attribute --vpc-id vpc-0abc... --enable-dns-support '{"Value":true}'
aws ec2 modify-vpc-attribute --vpc-id vpc-0abc... --enable-dns-hostnames '{"Value":true}'
VPC endpoints: gateway vs interface (PrivateLink)
A VPC endpoint lets resources in your VPC reach supported AWS services (and third-party / your-own services) privately, over the AWS network, without an Internet Gateway, NAT gateway, or public IPs. There are two fundamentally different kinds, and knowing which is which — and when each is even available — is core SAA/ANS material.
| Dimension | Gateway endpoint | Interface endpoint (PrivateLink) |
|---|---|---|
| Supported services | Only Amazon S3 and DynamoDB | Most AWS services (SSM, EC2 API, ECR, CloudWatch, Secrets Manager, SQS, KMS, …) and partner/your-own services |
| How it works | A route added to your route table targeting the endpoint (a prefix list) |
An ENI with a private IP placed in your subnet(s) |
| What you point at it | Route tables | DNS — queries to the service name resolve to the ENI’s private IP (with private DNS enabled) |
| Cost | Free (no hourly or data charge) | Hourly per-endpoint, per-AZ charge + per-GB data processing |
| Cross-Region / on-prem reachable | No (stays in-Region, in-VPC) | Yes — reachable over peering, TGW, VPN, Direct Connect |
| Access control | Endpoint policy (a resource policy on the endpoint) | Endpoint policy + security group on the ENI |
| Use it for | Keeping S3/DynamoDB traffic off the NAT gateway — the classic cost win | Private access to every other AWS service API |
The decision tree is simple. Is it S3 or DynamoDB? Use a gateway endpoint — it is free and removes that traffic from your NAT bill entirely (a private subnet that only talks to S3 may not need a NAT gateway at all). Anything else? Use an interface endpoint, accepting the hourly cost in exchange for keeping API traffic private. Interface endpoints are built on AWS PrivateLink, the same technology you use to expose your own service privately to other VPCs — covered in depth in AWS PrivateLink for Service Providers and Consumers. Two gotchas: gateway endpoints are Region-local and route-based, so they do not work for on-prem or cross-Region callers (use an interface endpoint there); and interface-endpoint private DNS only works when both enableDnsSupport and enableDnsHostnames are on (see the DNS section above).
# Gateway endpoint for S3 (free) — attach to the private route table
aws ec2 create-vpc-endpoint --vpc-id vpc-0abc... \
--vpc-endpoint-type Gateway \
--service-name com.amazonaws.ap-south-1.s3 \
--route-table-ids rtb-0priv...
# Interface endpoint for SSM (PrivateLink) — ENIs in the private subnets, with private DNS
aws ec2 create-vpc-endpoint --vpc-id vpc-0abc... \
--vpc-endpoint-type Interface \
--service-name com.amazonaws.ap-south-1.ssm \
--subnet-ids subnet-0priv... \
--security-group-ids sg-0... \
--private-dns-enabled
Connecting VPCs: peering vs Transit Gateway (in brief)
A single VPC is rarely the whole story — you connect VPCs to each other and to on-premises. Two options dominate, and the line between them is a frequent interview question.
| Dimension | VPC peering | Transit Gateway (TGW) |
|---|---|---|
| Topology | One-to-one link between two VPCs | Hub-and-spoke; one TGW connects many VPCs (and VPN/Direct Connect) |
| Transitivity | Non-transitive — if A↔B and B↔C, A still cannot reach C | Transitive — all attached VPCs can route to each other |
| Scale | Connections explode as n(n-1)/2 (a full mesh of 10 VPCs = 45 peerings) |
Linear — each VPC attaches once |
| Routing control | Per-VPC route tables | Central TGW route tables; segmentation via multiple route tables |
| Cost | No hourly fee; pay data transfer | Hourly per-attachment fee + per-GB; more, but far simpler at scale |
| Cross-Region | Inter-Region peering supported | TGW peering across Regions |
| Use it when | A handful of VPCs, simple any-to-any | Many VPCs / accounts, central egress, hybrid connectivity |
The headline rule: peering is non-transitive and does not scale — it is fine for two or three VPCs, but a growing estate becomes an unmanageable mesh, at which point you move to a Transit Gateway, which is transitive, centrally routed, and the standard for multi-account networking. Peering also requires non-overlapping CIDRs (you cannot peer two 10.0.0.0/16 VPCs) and does not support edge-to-edge routing (you cannot use a peer’s IGW or NAT). The full hub-and-spoke design, segmentation, and centralised egress are covered in Designing Multi-Account VPC Connectivity with Transit Gateway.
VPC Flow Logs: seeing the traffic
You cannot debug — or secure — a network you cannot see. VPC Flow Logs capture metadata about the IP traffic going to and from network interfaces: source and destination IP and port, protocol, packet and byte counts, the action (ACCEPT or REJECT), and more. They do not capture packet contents — this is NetFlow-style metadata, not a packet capture.
| Setting | What it is | Choices | Notes |
|---|---|---|---|
| Scope | What the logs cover | VPC, subnet, or a single ENI | VPC-level captures everything beneath it; start there. |
| Filter | Which traffic to record | All, Accepted, or Rejected | Rejected is great for spotting blocked traffic / misconfigured security groups. |
| Destination | Where logs go | CloudWatch Logs, S3, or Kinesis Data Firehose | S3 is cheapest for archival/analytics (query with Athena); CloudWatch for alerting. |
| Format | Which fields | Default or custom | Add fields like vpc-id, subnet-id, pkt-srcaddr, tcp-flags for richer analysis. |
| Aggregation interval | How often records are emitted | 1 min or 10 min | 1-minute is more granular but higher volume. |
The single most important caveat for troubleshooting: flow logs show the result of security group and NACL evaluation, not the rules themselves. A REJECT tells you traffic was blocked but not by which layer — that you reason out from the stateful/stateless behaviour you will learn in the next lesson. Turn flow logs on for every production VPC; they are inexpensive (you pay for log storage/ingestion) and indispensable the day something breaks or a security review asks “what talked to what”.
aws ec2 create-flow-logs \
--resource-type VPC --resource-ids vpc-0abc... \
--traffic-type ALL \
--log-destination-type s3 \
--log-destination arn:aws:s3:::my-flow-logs-bucket/vpc/
The complete picture
The diagram below assembles every component into one production-shaped VPC: a /16 split into public and private subnets across two Availability Zones, an Internet Gateway on the public tier, a NAT Gateway per AZ for private-subnet egress, a free gateway endpoint pulling S3 traffic off the NAT, an interface endpoint for an AWS API, the route tables that tie it together, and flow logs watching it all.
Trace a packet through it and the whole lesson clicks: an instance in a private subnet hits the internet via its route table’s 0.0.0.0/0 → NAT gateway (in its AZ) → IGW; the same instance reaches S3 via the local-then-gateway-endpoint route with no NAT involved; and a packet to a peer subnet matches the local route and never leaves the VPC.
Hands-on lab
You will build a minimal but complete two-tier VPC — one public and one private subnet, an IGW, a NAT gateway, and an S3 gateway endpoint — entirely from the CLI, validate routing, then tear it all down. Everything here is AWS Free Tier eligible except the NAT gateway and the Elastic IP, so follow the Cost note and clean up promptly.
Prerequisites: the AWS CLI v2 configured (aws configure) with a region set (examples use ap-south-1).
Step 1 — Create the VPC and turn on DNS hostnames.
VPC=$(aws ec2 create-vpc --cidr-block 10.0.0.0/16 \
--query Vpc.VpcId --output text)
aws ec2 modify-vpc-attribute --vpc-id $VPC --enable-dns-hostnames '{"Value":true}'
aws ec2 create-tags --resources $VPC --tags Key=Name,Value=lab-vpc
echo "VPC=$VPC"
Step 2 — Create one public and one private subnet in the same AZ (for simplicity).
PUB=$(aws ec2 create-subnet --vpc-id $VPC --cidr-block 10.0.1.0/24 \
--availability-zone ap-south-1a --query Subnet.SubnetId --output text)
aws ec2 modify-subnet-attribute --subnet-id $PUB --map-public-ip-on-launch
PRIV=$(aws ec2 create-subnet --vpc-id $VPC --cidr-block 10.0.2.0/24 \
--availability-zone ap-south-1a --query Subnet.SubnetId --output text)
echo "PUB=$PUB PRIV=$PRIV"
Step 3 — Attach an Internet Gateway and make the public subnet public.
IGW=$(aws ec2 create-internet-gateway --query InternetGateway.InternetGatewayId --output text)
aws ec2 attach-internet-gateway --internet-gateway-id $IGW --vpc-id $VPC
RTPUB=$(aws ec2 create-route-table --vpc-id $VPC --query RouteTable.RouteTableId --output text)
aws ec2 create-route --route-table-id $RTPUB --destination-cidr-block 0.0.0.0/0 --gateway-id $IGW
aws ec2 associate-route-table --route-table-id $RTPUB --subnet-id $PUB
Step 4 — Create a NAT gateway in the public subnet and route the private subnet through it.
EIP=$(aws ec2 allocate-address --domain vpc --query AllocationId --output text)
NAT=$(aws ec2 create-nat-gateway --subnet-id $PUB --allocation-id $EIP \
--query NatGateway.NatGatewayId --output text)
# Wait until the NAT gateway is available (takes a couple of minutes)
aws ec2 wait nat-gateway-available --nat-gateway-ids $NAT
RTPRIV=$(aws ec2 create-route-table --vpc-id $VPC --query RouteTable.RouteTableId --output text)
aws ec2 create-route --route-table-id $RTPRIV --destination-cidr-block 0.0.0.0/0 --nat-gateway-id $NAT
aws ec2 associate-route-table --route-table-id $RTPRIV --subnet-id $PRIV
Step 5 — Add a free S3 gateway endpoint to the private route table.
aws ec2 create-vpc-endpoint --vpc-id $VPC --vpc-endpoint-type Gateway \
--service-name com.amazonaws.ap-south-1.s3 --route-table-ids $RTPRIV
Step 6 — Validate. Confirm the routing is exactly what you intended:
# Public route table should show 0.0.0.0/0 -> igw-...
aws ec2 describe-route-tables --route-table-ids $RTPUB \
--query 'RouteTables[].Routes' --output table
# Private route table should show 0.0.0.0/0 -> nat-... AND an S3 prefix-list -> vpce-...
aws ec2 describe-route-tables --route-table-ids $RTPRIV \
--query 'RouteTables[].Routes' --output table
Expected: the public table has a local route plus 0.0.0.0/0 → igw-…; the private table has local, 0.0.0.0/0 → nat-…, and a pl-… (S3) → vpce-… route. That last line is the gateway endpoint at work — S3 traffic now bypasses the NAT entirely.
Cleanup — delete in reverse dependency order (endpoints and NAT before the IGW and subnets, or the deletes will fail):
EPID=$(aws ec2 describe-vpc-endpoints --filters Name=vpc-id,Values=$VPC \
--query 'VpcEndpoints[0].VpcEndpointId' --output text)
aws ec2 delete-vpc-endpoints --vpc-endpoint-ids $EPID
aws ec2 delete-nat-gateway --nat-gateway-id $NAT
aws ec2 wait nat-gateway-deleted --nat-gateway-ids $NAT
aws ec2 release-address --allocation-id $EIP
aws ec2 detach-internet-gateway --internet-gateway-id $IGW --vpc-id $VPC
aws ec2 delete-internet-gateway --internet-gateway-id $IGW
aws ec2 delete-subnet --subnet-id $PUB
aws ec2 delete-subnet --subnet-id $PRIV
aws ec2 delete-route-table --route-table-id $RTPUB
aws ec2 delete-route-table --route-table-id $RTPRIV
aws ec2 delete-vpc --vpc-id $VPC
Cost note: the VPC, subnets, IGW, route tables, and the S3 gateway endpoint are free. The two charged items are the NAT gateway (an hourly rate plus per-GB processing) and the Elastic IP (free while attached to a running resource, charged when idle). Running this lab for under an hour and cleaning up costs a few US cents at most — but do not leave the NAT gateway running, as its hourly charge accrues around the clock.
Going deeper
You now know every component. This section is for the reader who has to operate a VPC in production — the limits, the less-obvious endpoint and routing features, and the tools that debug a network without touching it. None of this is required to build the lab above; all of it turns up in real incidents and Advanced Networking (ANS-C01) questions.
NAT Gateway SNAT ports — the connection limit nobody sees until it bites
A NAT gateway performs source NAT (SNAT): it rewrites each outbound flow’s private source IP to its own address and picks a source port so the return packet can be demultiplexed back to the right instance. Each NAT gateway sustains up to roughly 55,000 simultaneous connections to each unique destination — a unique destination being one combination of destination IP, destination port, and protocol. Exceed that against a single popular endpoint and new connections fail while the ErrorPortAllocation CloudWatch metric climbs and packets drop.
The trap is that the ceiling is per unique destination, not per NAT gateway overall. 55,000 connections spread across many destinations is fine; 60,000 connections hammering one api.example.com:443 from behind one NAT gateway will start failing. Three levers relieve it: (1) fan out the destination — if the far side answers on several IPs, the limit multiplies; (2) add NAT gateways and split private subnets across them; (3) kill connection churn with HTTP keep-alive / connection pooling so you are not burning a fresh source port per request. Bandwidth auto-scales 5 → 100 Gbps and rarely bottlenecks first, so a rising PacketsDropCount on an otherwise-idle NAT gateway is almost always port exhaustion.
| NAT gateway metric | What a non-zero / rising value means |
|---|---|
ErrorPortAllocation |
SNAT source-port exhaustion — the 55k-per-destination ceiling. Alarm on > 0. |
PacketsDropCount |
Packets dropped; with ErrorPortAllocation > 0 it confirms port exhaustion, not bandwidth. |
ActiveConnectionCount |
Concurrent connections in flight — trend it to see exhaustion coming. |
BytesOutToDestination |
Drives the per-GB data-processing charge — the cost half of the NAT bill. |
Ingress routing: gateway route tables and middlebox insertion (edge association)
Every route table so far attaches to a subnet. There is a second kind: you can associate a route table with the Internet Gateway itself (or a virtual private gateway). This is edge association, and the table is an ingress routing table — it redirects traffic entering the VPC through a middlebox (a firewall or IDS/IPS appliance) before it reaches the destination subnet.
The rules are deliberately narrow: a gateway route table may only hold routes whose destination is the VPC CIDR or a more-specific subset, and the target must be an ENI or a Gateway Load Balancer endpoint — never an IGW or NAT gateway. So the inbound path becomes: internet → IGW → (ingress table steers the subnet’s CIDR to the appliance) → appliance inspects → appliance forwards to the real subnet. That is how you insert inline inspection for inbound flows without changing anything on the workload instances themselves.
# Steer INBOUND traffic destined for the public subnet through an appliance ENI first
aws ec2 create-route --route-table-id rtb-0edge... \
--destination-cidr-block 10.0.1.0/24 \
--network-interface-id eni-0firewall...
# Associate the route table with the IGW (edge association), not a subnet
aws ec2 associate-route-table --route-table-id rtb-0edge... --gateway-id igw-0xyz...
A third endpoint type: Gateway Load Balancer endpoints
The gateway-vs-interface split you learned is the common case, but there is a third endpoint type: the Gateway Load Balancer endpoint (GWLBe). Like an interface endpoint it is powered by PrivateLink and lands in your VPC, but instead of fronting an AWS API it fronts a Gateway Load Balancer that spreads traffic across a fleet of third-party virtual appliances (next-gen firewalls, IDS/IPS) using the GENEVE protocol on UDP 6081. You reference a GWLBe as a route target — usually from an ingress or subnet route table — to transparently pass traffic through the appliance fleet and back.
| Endpoint type | Powered by | Fronts | Referenced from | Typical use |
|---|---|---|---|---|
| Gateway | route / prefix list | S3, DynamoDB only | route table | free private S3 / DynamoDB |
| Interface | PrivateLink (ENI) | most AWS APIs, partner / your services | DNS + SG on the ENI | private API access |
| Gateway Load Balancer (GWLBe) | PrivateLink | a Gateway Load Balancer + appliance fleet | route table | transparent inline inspection |
The full inspection architecture is in AWS Gateway Load Balancer for Inline Appliance Inspection.
Managed prefix lists — reference CIDRs by name
A prefix list is a named set of CIDR blocks you reference by its pl-... ID in route tables and security-group rules instead of pasting raw CIDRs everywhere. Two flavours: AWS-managed lists (AWS maintains them — the S3 and DynamoDB gateway endpoints you added earlier appear in the route table as a pl-... entry, and that is the S3 managed prefix list), and customer-managed lists you create and version yourself, e.g. “all our office ranges”. Reference one list in fifty security groups, then onboard a new branch office by editing the single list — every rule updates at once. The --max-entries you fix at creation counts against route-table and security-group entry limits, so size it deliberately (you can resize later).
aws ec2 create-managed-prefix-list --prefix-list-name corp-offices \
--max-entries 20 --address-family IPv4 \
--entries Cidr=203.0.113.0/24,Description=hq Cidr=198.51.100.0/24,Description=branch
Hybrid and private DNS: what the .2 resolver really is
The .2 resolver (the VPC base address +2, also reachable at the link-local 169.254.169.253) is the front door to Route 53 Resolver, the service that answers DNS inside every VPC. Left alone it resolves public names and the internal AWS names. Two building blocks extend it:
- Private hosted zones (PHZ) — a Route 53 zone (say
db.internal) associated with one or more VPCs, soprimary.db.internalresolves privately with no public record anywhere. This is also the mechanism behind interface-endpoint private DNS: enabling private DNS on an endpoint associates an AWS-managed PHZ that overrides the public service name with the endpoint’s private IP (which is exactly whyenableDnsHostnamesmust be on for it to work). - Resolver endpoints for hybrid DNS — an inbound endpoint gives on-premises resolvers a target to query VPC / PHZ names; an outbound endpoint plus resolver rules forward chosen domains (e.g.
corp.example.com) from the VPC out to your on-prem DNS. This is the modern replacement for hand-editing DHCP option sets to point at custom DNS servers.
The hybrid design in full — inbound/outbound endpoints, rules, and DNS Firewall — is covered in Route 53 Resolver, DNS Firewall & Hybrid Resolution.
IPv6 in practice: no NAT, everything routable
Enabling IPv6 (an Amazon-provided /56 to the VPC, a /64 per subnet) shifts several defaults worth stating outright:
- Every IPv6 address is globally routable — there is no NAT and no “private” IPv6 in the RFC 1918 sense. “Private” means simply not having an internet route, or using an egress-only internet gateway for outbound-only access (the IPv6 analogue of a NAT gateway — stateful, and free of the NAT gateway’s hourly + per-GB charge).
- Dual-stack subnets carry both families; you add a
::/0route (to the IGW, or to the egress-only IGW) alongside the IPv40.0.0.0/0. Security groups and NACLs need explicit IPv6 rules — an IPv40.0.0.0/0rule does not cover::/0. - Because there is no NAT to hide behind, IPv6 exposure is governed entirely by security groups, NACLs, and not routing inbound — audit those before enabling it on a subnet that already holds workloads.
Two connectivity patterns the basics miss: private NAT & shared subnets
Real estates routinely need two capabilities the public-NAT-plus-peering story never covers:
- Private NAT gateway. A NAT gateway created with connectivity type
privateperforms SNAT without any internet path — no Elastic IP, no IGW dependency. Its job is overlapping-CIDR communication: when two networks you must connect share an address range (a merger, a partner integration), you SNAT one side’s traffic to a distinct private range so the far side sees non-overlapping sources, routed over a Transit Gateway or VPN. Same SNAT engine, same source-port limits — just aimed at private space instead of the internet. - VPC sharing (AWS RAM). Instead of every account building its own VPC (and a peering/TGW mesh to join them), a central “network” account can share subnets into other accounts in the AWS Organization using AWS Resource Access Manager (RAM). Participant accounts launch instances, databases, and load balancers directly into the shared subnets, while the owner account alone controls the VPC, route tables, IGW, and NAT. It collapses VPC sprawl and centralises egress — participants cannot modify or delete the subnet or VPC, only place resources in it.
- Extra space from
100.64.0.0/10. When RFC 1918 ranges are exhausted across the estate, AWS lets you add a secondary CIDR outside RFC 1918 — the carrier-grade-NAT block100.64.0.0/10(RFC 6598) is the popular pick for internal-only, high-density subnets (think EKS pod IPs) precisely because it will not collide with a customer’s RFC 1918 network.
# Private NAT gateway — SNAT into private space, no Elastic IP, no internet path
aws ec2 create-nat-gateway --subnet-id subnet-0priv... \
--connectivity-type private
VPC Flow Logs: the fields that actually solve incidents
The default flow-log format is fine for a glance, but two facts make or break a real investigation:
srcaddr/dstaddrvspkt-srcaddr/pkt-dstaddr. In the default formatsrcaddris the address at the ENI — which, behind a NAT gateway or an intermediate appliance, is the NAT/appliance, not the origin. The custom fieldspkt-srcaddrandpkt-dstaddrcarry the original packet addresses, so they unmask “who really talked to what” through NAT and reveal asymmetric routing (the two differ when traffic takes different paths in and out).log-status. Each record carriesOK,NODATA(no traffic in the window — often the clue that a security group or NACL silently dropped everything), orSKIPDATA(records skipped due to an internal capacity constraint — a gap in your data, not a network event).
# Flow logs with a custom format that captures the ORIGINAL addresses + TCP flags
aws ec2 create-flow-logs --resource-type VPC --resource-ids vpc-0abc... \
--traffic-type ALL --log-destination-type s3 \
--log-destination arn:aws:s3:::my-flow-logs-bucket/vpc/ \
--log-format '${srcaddr} ${dstaddr} ${pkt-srcaddr} ${pkt-dstaddr} ${action} ${log-status} ${tcp-flags}'
Debugging routing without sending a packet
Config tables eventually need a debugger. Two built-in tools analyse your configuration (they do not send live traffic), so they work even when nothing can reach anything:
- VPC Reachability Analyzer traces a hop-by-hop path between a source and destination (instance, ENI, IGW, TGW, endpoint) and returns reachable or not reachable — and when not, the exact component that blocks it (a missing route, a security group, a NACL). It is the fastest answer to “why can’t A reach B?” without logging into anything.
- Network Access Analyzer runs the other direction: you declare a Network Access Scope (“nothing in the database tier is reachable from the internet”) and it surfaces every path that violates it — a guardrail for audits and config drift.
# Ask "can this instance reach that one on 443?" — analysed from configuration, no packets sent
aws ec2 create-network-insights-path \
--source i-0source... --destination i-0dest... \
--protocol tcp --destination-port 443
aws ec2 start-network-insights-analysis --network-insights-path-id nip-0abc...
Both tools get a full treatment in Reachability Analyzer & Network Access Analyzer for Connectivity Validation.
Common mistakes & troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Instance in a “public” subnet has no internet | Missing one of the three conditions (no IGW, no 0.0.0.0/0 route, or no public IP) |
Verify IGW attached, route table has 0.0.0.0/0 → igw, and the instance has a public/Elastic IP. |
| Private instance cannot reach the internet outbound | No NAT route, or NAT gateway is in a private subnet | Put the NAT gateway in a public subnet and point the private route table’s 0.0.0.0/0 at it. |
Cannot create subnet: CIDR not within VPC |
Subnet CIDR outside the VPC block, or overlaps another subnet | Choose a sub-range that fits inside the VPC CIDR and does not overlap. |
| Interface endpoint name still resolves to a public IP | enableDnsHostnames is off (or private DNS not enabled) |
Turn on enableDnsHostnames and enableDnsSupport; enable private DNS on the endpoint. |
| Cannot peer two VPCs | Overlapping CIDR ranges | Re-IP one VPC, or use distinct ranges from the start — overlapping blocks cannot be peered. |
| Self-built NAT instance forwards nothing | Source/destination check still enabled | Run modify-instance-attribute --no-source-dest-check. |
| S3 traffic is inflating the NAT bill | No gateway endpoint; S3 traffic flows through NAT | Add a free S3 gateway endpoint to the private route table. |
| App “A” cannot reach app “C” through a middle VPC | Peering is non-transitive | Add a direct peering A↔C, or move to a Transit Gateway. |
Common beginner mistakes
These are misconceptions — wrong mental models — rather than the symptom→fix entries in the troubleshooting table above. Fixing the model fixes a whole class of bugs at once.
-
“I’ll just tick the ‘public subnet’ box.” There is no such box. A subnet is public only because its route table sends
0.0.0.0/0to an Internet Gateway. Auto-assign-public-IP alone does nothing without that route; the route alone does nothing without a public IP on the instance. Public = route to IGW and a public/Elastic IP and permissive security group / NACL — all three, every time. -
“A /24 has 256 usable addresses.” It has 251. AWS reserves the first four and the last address in every subnet, so your DHCP pool, your ENIs, and any fixed IPs all come out of 251. Plan capacity on the usable count (
2^(32−prefix) − 5), not the raw block size. -
“One NAT gateway is highly available.” A NAT gateway is redundant within a single AZ, but it is zonal — if that AZ fails, every private subnet pointed at it loses egress, including subnets in healthy AZs. HA means one NAT gateway per AZ, each AZ’s private route table pointing at its own NAT gateway (which also avoids cross-AZ data-transfer charges).
-
“Endpoints and NAT do the same job.” A NAT gateway is a paid, general-purpose egress path to the whole internet. A gateway endpoint is a free, route-based private path to only S3 and DynamoDB. Routing terabytes of S3 traffic through NAT when a free endpoint exists is the single most common avoidable VPC cost.
-
“I created the endpoint but the service name still resolves to a public IP.” That is a DNS problem, not a routing one: interface-endpoint private DNS needs both
enableDnsSupportandenableDnsHostnameson the VPC, plus private-DNS enabled on the endpoint. The route was never the issue. -
“My two VPCs are peered, so everything can talk.” Peering is non-transitive and requires non-overlapping CIDRs. A↔B and B↔C does not let A reach C, and two
10.0.0.0/16VPCs can never peer at all. Plan unique CIDRs up front; reach for a Transit Gateway when the mesh grows. -
“Security groups protect my private subnet, so routing doesn’t matter.” Routing and filtering are different layers: a route decides where a packet may go, a security group decides whether it is allowed. A “private” subnet with an accidental
0.0.0.0/0 → IGWroute plus a public IP is exposed no matter how tidy the security group looks — audit routes, not only rules. -
“IPv6 is private like
10.xaddresses.” No — every AWS IPv6 address is globally routable. There is no NAT for IPv6; “private” means no inbound route (or an egress-only IGW for outbound). And an IPv40.0.0.0/0security-group rule does not cover::/0— dual-stack needs explicit IPv6 rules.
Best practices
- Plan CIDR for the whole estate. Allocate a unique, non-overlapping block per VPC, size production VPCs at
/16, and leave room for secondary blocks. Use IPAM once you have more than a handful. - Multi-AZ by default. At least two subnets in two AZs per tier; one NAT gateway per AZ so an AZ failure never takes out egress and you avoid cross-AZ data charges.
- Three subnet tiers. Public (load balancers, NAT), private-with-egress (app servers), and isolated (databases) — separated by their route tables.
- Leave the main route table private. Attach explicit custom route tables to public subnets so nothing becomes public by accident.
- Use endpoints aggressively. A free gateway endpoint for S3/DynamoDB and interface endpoints for the AWS APIs your private workloads call — this both saves NAT cost and keeps traffic off the public internet.
- Turn on flow logs everywhere (to S3 for cheap archival), and tag every network resource (
env,owner,tier) for cost allocation and automation.
Security notes
- Private by default. Build custom VPCs where nothing is internet-reachable unless a route deliberately makes it so; reserve public subnets for the few resources that truly need ingress.
- Endpoints reduce exposure. Reaching AWS services through interface/gateway endpoints keeps that traffic on the AWS network and off any IGW/NAT, shrinking your attack surface; pair with endpoint policies to restrict which resources can be reached.
- Flow logs are an audit and detection tool.
REJECTrecords surface scanning and misconfiguration; ship them to S3 and query with Athena, or to CloudWatch for alarms. - Routing is not a firewall. A route gets a packet to a destination; security groups and network ACLs decide whether it is allowed — design both layers, and read the next lesson for how stateful vs stateless filtering actually behaves.
- Mind the IGW NAT. The Internet Gateway’s one-to-one NAT means any instance with a public IP and a permissive security group is directly exposed — audit public IP assignment.
- Egress control at scale. For centralised, inspected egress across many VPCs, route through a Transit Gateway to an inspection VPC rather than per-VPC NAT.
Practice challenges
Work these in order — they climb from beginner counting to advanced design. Try each before opening the solution; the one-line why is the transferable lesson.
Challenge 1 (beginner) — Count the usable IPs. You create subnet 10.0.4.0/26. How many addresses can your instances actually use, and which specific addresses can they never use?
<details><summary>Solution</summary>
A /26 is 64 addresses (10.0.4.0–10.0.4.63); 59 are usable. AWS reserves 10.0.4.0 (network), 10.0.4.1 (VPC router), 10.0.4.2 (Amazon DNS resolver slot), 10.0.4.3 (future use), and 10.0.4.63 (broadcast).
Why: every subnet loses the same five addresses regardless of size — size on 2^(32−prefix) − 5.
</details>
Challenge 2 (beginner) — Make a subnet public, minimally. An instance with a public IP can’t reach the internet. The IGW igw-0xyz is already attached to the VPC and route table rtb-0pub is already associated with the subnet. Fix it in one command.
<details><summary>Solution</summary>
aws ec2 create-route --route-table-id rtb-0pub \
--destination-cidr-block 0.0.0.0/0 --gateway-id igw-0xyz
Why: the missing piece is the 0.0.0.0/0 → IGW route; a public IP plus an attached IGW do nothing until the route table points default traffic at the gateway.
</details>
Challenge 3 (intermediate) — Kill the S3 NAT bill. A private subnet’s instances pull terabytes from S3 monthly and the NAT gateway’s data-processing charge is now your biggest line item. Add the fix, and say what appears in the route table.
<details><summary>Solution</summary>
aws ec2 create-vpc-endpoint --vpc-id vpc-0abc \
--vpc-endpoint-type Gateway \
--service-name com.amazonaws.ap-south-1.s3 \
--route-table-ids rtb-0priv
A pl-… (S3 managed prefix list) → vpce-… route appears. S3 traffic now matches that more-specific route (longest-prefix match) and bypasses the NAT entirely — and the gateway endpoint is free.
Why: gateway endpoints move S3 / DynamoDB traffic off the paid NAT path at zero cost. </details>
Challenge 4 (intermediate) — Diagnose “private endpoint resolves public”. You added an interface endpoint for Secrets Manager with private DNS enabled, but secretsmanager.ap-south-1.amazonaws.com still resolves to a public IP from your instances. enableDnsSupport is already on. What’s wrong, and what’s the one-line fix?
<details><summary>Solution</summary>
aws ec2 modify-vpc-attribute --vpc-id vpc-0abc --enable-dns-hostnames '{"Value":true}'
Why: interface-endpoint private DNS needs both enableDnsSupport and enableDnsHostnames; the latter is off by default in custom VPCs, so the private hosted zone that overrides the public name never takes effect.
</details>
Challenge 5 (advanced) — Design per-AZ NAT resilience. Express, as route-table associations, a two-AZ design where losing either AZ never takes out the other AZ’s egress, with no cross-AZ NAT data-transfer charges. List the NAT gateways, route tables, and associations.
<details><summary>Solution</summary>
- NAT-A in public-subnet-A (AZ-a) and NAT-B in public-subnet-B (AZ-b) — two NAT gateways, one per AZ.
- rtb-priv-A:
0.0.0.0/0 → NAT-A, associated to private-subnet-A only. - rtb-priv-B:
0.0.0.0/0 → NAT-B, associated to private-subnet-B only. - One public route table (
0.0.0.0/0 → IGW) shared by both public subnets is fine.
Why: each AZ’s private subnet egresses through the NAT gateway in its own AZ, so an AZ loss is contained and no egress traffic crosses AZs (cross-AZ transfer is billed separately). </details>
Challenge 6 (advanced) — Explain a SNAT port-exhaustion incident. Instances behind a single NAT gateway open ~60,000 simultaneous connections to one partner API at 1.2.3.4:443. Some connections start failing even though bandwidth is far below 5 Gbps. Which CloudWatch metric confirms the cause, and give two fixes that don’t require reducing traffic?
<details><summary>Solution</summary>
- Metric:
ErrorPortAllocationrising above 0 (withPacketsDropCountclimbing) confirms SNAT source-port exhaustion — the ~55,000-connections-per-unique-destination ceiling, not bandwidth. - Two fixes without cutting traffic: (1) add NAT gateways and split the private subnets across them so flows to
1.2.3.4:443are shared; (2) enable HTTP keep-alive / connection pooling so requests reuse connections instead of each consuming a fresh source port. (A third, if the partner exposes multiple destination IPs: spread across them — the limit is per unique destination.)
Why: the limit is per destination IP + port + protocol; bandwidth auto-scales but source ports do not. </details>
Interview & exam questions
-
What makes a subnet “public”? Its route table has a
0.0.0.0/0route pointing at an Internet Gateway. (In practice you also enable auto-assign public IP or attach Elastic IPs.) There is no “public” flag on the subnet itself — it is purely routing. -
How many usable IPs are in a
/24subnet, and why not 256? 251. AWS reserves the first four addresses (network, VPC router, Amazon DNS, future use) and the last (broadcast) in every subnet. -
NAT Gateway vs NAT instance — give three differences. NAT Gateway is managed, auto-scales to 100 Gbps, and is HA within an AZ; the NAT instance is a self-managed EC2 (single point of failure, fixed throughput, you patch it) but can act as a bastion/port-forwarder and needs source/destination check disabled.
-
Why deploy one NAT gateway per Availability Zone? A NAT gateway is zonal; one per AZ removes the single-AZ dependency and avoids cross-AZ data-transfer charges by keeping each AZ’s egress local. If the AZ holding your only NAT gateway fails, all private-subnet egress fails.
-
Gateway endpoint vs interface endpoint — when do you use each? Gateway endpoints serve only S3 and DynamoDB, are free, and work via a route-table entry. Interface endpoints (PrivateLink) serve most other services, place an ENI with a private IP in your subnet, cost per hour + per GB, and are reachable cross-Region / on-prem.
-
Can you change a VPC’s primary CIDR after creation? No. The primary IPv4 CIDR is permanent. You can only add up to four (default) secondary CIDR blocks that do not overlap existing ranges.
-
An interface endpoint’s DNS name resolves to a public IP — what is wrong?
enableDnsHostnames(andenableDnsSupport) must be on, and private DNS must be enabled on the endpoint, for the service name to resolve to the endpoint’s private IP. -
Is VPC peering transitive? No. If A↔B and B↔C are peered, A cannot reach C through B. You add a direct A↔C peering or move to a Transit Gateway, which is transitive.
-
What is the
localroute and can you remove it? An automatic route for the entire VPC CIDR with targetlocalthat lets every subnet reach every other subnet. It cannot be deleted or modified and always wins for in-VPC traffic (longest-prefix match). -
You need a private instance to reach the internet outbound but never be reachable inbound — what do you build? A NAT gateway in a public subnet, with the private subnet’s
0.0.0.0/0route pointing at it. For IPv6, an egress-only internet gateway instead. -
How do route tables decide between two matching routes? Longest-prefix match — the most specific route wins (e.g.
/24over/0); static routes beat propagated BGP routes of the same prefix. -
How do you cut S3 data-transfer costs through a NAT gateway? Add a free S3 gateway endpoint to the private subnet’s route table so S3 traffic bypasses the NAT entirely.
Quick check
- True or false: a subnet can span two Availability Zones.
- Which two AWS services are supported by gateway endpoints?
- Which VPC attribute must be on for interface-endpoint private DNS names to resolve correctly?
- Where must a NAT gateway be placed — a public or a private subnet — and why?
- What is the smallest subnet size AWS allows, and what limits it?
Answers
- False — a subnet lives in exactly one AZ; use multiple subnets across AZs for HA.
- Amazon S3 and DynamoDB (only these two).
enableDnsHostnames(alongsideenableDnsSupport, which must also be on).- A public subnet — the NAT gateway needs a route to the Internet Gateway to reach the internet on behalf of private instances.
/28(16 addresses, 11 usable) — limited by the five reserved IPs AWS takes in every subnet.
Exercise
Design (on paper or in the console) a production-ready VPC for a three-tier web application in the ap-south-1 Region that must survive the loss of one Availability Zone:
- Choose a
/16CIDR and carve six subnets — public, private-app, and isolated-database tiers across two AZs. - Decide where IGWs and NAT gateways go, and how many NAT gateways you need for AZ resilience.
- Add a free S3 gateway endpoint and at least one interface endpoint (e.g. SSM, so you can manage instances without SSH/bastion).
- Sketch the route tables for each tier and confirm the database tier has no internet route.
- List which components are free and which incur cost, and estimate the dominant cost driver.
Bonus: explain what you would change to add a second VPC and connect the two, and at what point you would replace peering with a Transit Gateway.
Certification mapping
| Exam | Objective area this supports |
|---|---|
| SAA-C03 (Solutions Architect – Associate) | Design secure and resilient architectures — VPC/subnet/AZ design, public vs private routing, NAT for egress, gateway vs interface endpoints, and peering vs Transit Gateway trade-offs. |
| ANS-C01 (Advanced Networking – Specialty) | Network design and connectivity — CIDR/IPv6 planning, route-table behaviour and priority, PrivateLink/endpoints, DHCP option sets and hybrid DNS, and flow-log-based troubleshooting. |
| DVA-C02 (Developer – Associate) | Deployment and security — placing application resources in the right subnet tier and reaching AWS services privately via endpoints. |
| SOA-C02 (SysOps – Associate) | Networking and monitoring — operating NAT gateways, route tables, and VPC Flow Logs for day-to-day troubleshooting. |
Glossary
- VPC (Virtual Private Cloud) — a logically isolated, software-defined virtual network in one AWS Region.
- CIDR — Classless Inter-Domain Routing; the
/16-style notation defining an address block’s size. - Subnet — a sub-range of the VPC CIDR confined to a single Availability Zone.
- Availability Zone (AZ) — one or more discrete data centres in a Region with independent power and networking.
- Reserved IPs — the five addresses (first four + last) AWS reserves in every subnet.
- Route table — the ordered rules mapping destination CIDRs to targets; one per subnet (defaults to the main table).
localroute — the unremovable route for the VPC’s own CIDR that makes all subnets mutually reachable.- Internet Gateway (IGW) — the VPC’s door to the public internet; performs one-to-one NAT for public IPs.
- Egress-only internet gateway — the IPv6 equivalent of NAT: outbound-only IPv6 internet access.
- NAT (Network Address Translation) — letting many private addresses share a public address for outbound traffic.
- NAT Gateway / NAT instance — the managed vs self-managed ways to provide outbound-only internet to private subnets.
- Elastic IP (EIP) — a static public IPv4 address you allocate and attach.
- ENI (Elastic Network Interface) — the virtual NIC that resources attach to in a subnet.
- DHCP option set — VPC-level DNS/domain/NTP configuration handed to instances at boot.
enableDnsSupport/enableDnsHostnames— the two VPC attributes that control AWS DNS resolution and public/endpoint DNS names.- VPC endpoint — a private on-ramp to AWS (or partner) services; gateway (S3/DynamoDB, free, route-based) or interface (PrivateLink, ENI-based).
- PrivateLink — the technology behind interface endpoints; also exposes your own services privately.
- VPC peering — a one-to-one, non-transitive connection between two VPCs.
- Transit Gateway (TGW) — a transitive hub connecting many VPCs and on-prem links.
- VPC Flow Logs — metadata (not packet contents) about IP traffic on ENIs/subnets/VPCs.
- SNAT (source NAT) / SNAT ports — the NAT gateway rewriting a private source IP and choosing a source port per flow so returns route back; capped at ~55,000 connections per unique destination.
ErrorPortAllocation— the NAT gateway CloudWatch metric that signals SNAT source-port exhaustion.- Ingress routing / edge (gateway) route table — a route table associated with an IGW or virtual private gateway that steers inbound traffic to a middlebox (ENI or GWLBe) before it reaches the destination subnet.
- Gateway Load Balancer endpoint (GWLBe) — a third VPC endpoint type: a PrivateLink endpoint used as a route target to send traffic through a fleet of inline inspection appliances via GENEVE.
- GENEVE — the encapsulation protocol (UDP 6081) a Gateway Load Balancer uses to hand packets to appliances.
- Managed prefix list — a named, reusable set of CIDRs (AWS-managed like the S3 list, or customer-managed) referenced by
pl-…ID in routes and security groups. - Route 53 Resolver — the VPC DNS service behind the
.2resolver; inbound / outbound endpoints and rules enable hybrid on-prem resolution. - Private hosted zone (PHZ) — a Route 53 DNS zone associated with VPCs for private-only name resolution; the mechanism behind interface-endpoint private DNS.
- Egress-only internet gateway — the IPv6 analogue of a NAT gateway: stateful, outbound-only IPv6, no per-GB charge.
- Reachability Analyzer — a configuration-based tool that traces the path between two endpoints and names the component blocking connectivity.
- Network Access Analyzer — a tool that finds network paths violating a defined Network Access Scope.
Next steps
Continue the course with AWS Security Groups vs Network ACLs, In Depth — now that you control where traffic flows, learn the filtering layer that decides what is allowed, including the stateful-vs-stateless difference and the classic “return traffic blocked by a NACL” gotcha. Then deepen your networking with:
- Amazon VPC IPAM: Hierarchical CIDR Planning, Allocation & BYOIP at Scale — automate the address planning this lesson did by hand.
- Designing Multi-Account VPC Connectivity with Transit Gateway — replace the peering mesh with a transitive hub and centralised egress.
- AWS PrivateLink for Service Providers & Consumers — expose and consume private services across accounts using the technology behind interface endpoints.