AWS 267: AWS Site-to-Site VPN
Why this lesson matters
AWS Site-to-Site VPN creates managed IPsec tunnels between an AWS network and a customer network. It is often the fastest way to establish encrypted hybrid connectivity, a practical backup for Direct Connect, or a production path when internet variability and supported throughput fit the requirement. It is not “a private cable,” and a green tunnel icon does not prove that applications can communicate.
A working design needs correct components on both sides, compatible IKE/IPsec settings, non-overlapping prefixes, forward and return routes, firewall permission, an appropriate MTU, monitoring, two configured tunnels, and named owners. The AWS team owns the AWS endpoint; the customer owns the customer gateway device, internet path, routing policy, firewall, application path, and usually the hardest half of diagnosis.
This lesson builds the complete mental model from Linux routing concepts to architect-level choices and operations. It is read-only: you will inspect supplied or explicitly owned evidence and create a local design, not a billable VPN connection.
Outcomes
By the end, you can:
- distinguish a customer gateway resource from the physical or virtual customer gateway device;
- choose a virtual private gateway (VGW), Transit Gateway (TGW), Cloud WAN, acceleration, or a different connectivity service;
- explain IKE Phase 1, IPsec Phase 2, inside and outside addresses, PSK/certificate authentication, NAT traversal, DPD, and rekeying;
- design static or BGP routing and trace both forward and return paths;
- explain two-tunnel behavior, VGW active-path selection, TGW ECMP, and true device/site redundancy;
- select IPv4/IPv6, standard/large bandwidth, MTU, MSS, cryptographic, and logging options;
- inspect configuration, routes, telemetry, metrics, logs, and maintenance status with read-only commands;
- diagnose tunnel-down, BGP-down, one-way, intermittent, low-throughput, and failover faults;
- design a safe game day, credential rotation, maintenance, and rollback plan; and
- identify every cost component and prove that this no-create lesson changed nothing.
From Linux routing to a managed VPN
If you know Linux networking, map familiar concepts carefully:
| Linux/network concept | Site-to-Site VPN equivalent | Important difference |
|---|---|---|
| IPsec/IKE peer | AWS tunnel endpoint and customer gateway device | AWS manages its endpoint; you configure both tunnels on your device |
ip route | VPC route table, TGW route table, customer routing table | Several independent tables must agree |
| FRR/BIRD BGP neighbor | BGP peer over tunnel inside addresses | Tunnel can be IPsec-up while BGP is down |
iptables/nftables | Customer firewall plus security groups/NACLs | SGs are stateful; NACLs and many customer ACLs are stateless |
| Interface MTU/TCP MSS clamp | Tunnel overhead control | VPN has no jumbo frames and does not support Path MTU Discovery |
tcpdump and daemon logs | Customer packet capture plus AWS VPN logs/metrics | AWS logs negotiations and states, not customer payloads |
Encryption protects traffic in transit. It does not authorize an application, resolve DNS, fix overlapping addresses, inspect malware, or automatically advertise every network.
Components and terminology
| Component | What it represents | Common mistake |
|---|---|---|
| Customer gateway device | Router, firewall, SD-WAN appliance, or software endpoint in the customer network | Calling the AWS record the router |
| Customer gateway (CGW) | AWS resource containing the device's outside IP/certificate identity and BGP ASN | Expecting it to forward packets by itself |
| VPN connection | Two managed IPsec tunnels and their options | Configuring only Tunnel 1 |
| Virtual private gateway | Regional AWS VPN concentrator attached to one VPC | Expecting ECMP or IPv6 inner traffic |
| Transit Gateway | Regional routing hub for many VPC/on-premises attachments | Forgetting TGW association/propagation and VPC subnet routes |
| Tunnel outside addresses | Internet- or private-underlay endpoints carrying IKE/IPsec | Confusing them with routed workload addresses |
| Tunnel inside addresses | Point-to-point addresses used by tunnel/BGP peers | Advertising them as application networks |
| Local/remote network CIDRs | IKE Phase 2 proposal ranges | Treating them as enforced firewall rules; AWS uses route-based VPNs |
Each VPN connection supplies two tunnels for AWS-side endpoint resilience. That does not remove a single customer device, single power domain, single ISP, or single site as a failure domain. Higher availability commonly needs two customer devices and independent connections/underlays, with tested routing convergence.
Packet and control flow
application in on-premises prefix 10.20.0.0/16
-> local route/VRF
-> customer firewall and selected customer gateway device
-> IKE/IPsec security associations
-> encrypted outer packet across internet or private underlay
-> AWS tunnel endpoint
-> VGW or TGW route decision
-> VPC subnet route, NACL, security group, workload
-> return route through the same valid hybrid routing domain
Separate four proofs:
- IKE Phase 1: peers authenticate and create an IKE security association.
- IPsec Phase 2: peers negotiate child/IPsec security associations for protected traffic.
- Routing: static routes or BGP install usable prefixes through a healthy tunnel.
- Application: DNS, network controls, host firewall, listener, TLS, and application authorization succeed.
“Tunnel UP” mainly addresses the first two. “BGP ESTABLISHED” adds a routing adjacency but does not prove accepted prefixes, route-table selection, security controls, or the application.
Choose the AWS-side gateway
| Requirement | Better starting point | Reason/caveat |
|---|---|---|
| One VPC, simple IPv4 hybrid path | VGW | Lower conceptual overhead; one VGW attaches to a VPC at a time |
| Many VPCs/accounts or segmented routing domains | TGW | Central hub with attachment association and propagation controls |
| Multiple active tunnels for aggregate bandwidth | TGW with dynamic routing and ECMP | VGW does not support VPN ECMP |
| IPv6 inner or IPv6 outer tunnels | TGW or Cloud WAN | VGW supports the basic IPv4 outer/IPv4 inner model |
| Up to 5-Gbps large-bandwidth tunnel option | TGW or Cloud WAN | Large option is not available on VGW |
| Better internet-path consistency from a distant site | Accelerated VPN on TGW | Uses Global Accelerator; adds restrictions and cost |
| Private IP VPN over Direct Connect | TGW plus Direct Connect design | Underlay and overlay are separate; Direct Connect dependencies remain |
| Global managed core and policy | Cloud WAN evaluation | Broader service, policy, and cost decision than one VPN |
Accelerated VPN enters the AWS network at a nearby edge, but it is not available on VGW, requires NAT-T, must be selected when the connection is created, and requires the customer device to initiate IKE. It cannot simply be toggled on for an existing connection. Measure latency, loss, jitter, and cost before selecting it.
Static routing or BGP
Static routing lists customer prefixes in the VPN connection and configures explicit routes on the customer device. It suits a small, stable prefix set and devices without BGP. Every topology change becomes coordinated configuration, and liveness/failover depends on device mechanisms rather than route withdrawal.
Dynamic routing establishes BGP over each tunnel's inside addresses. The customer device advertises reachable on-premises prefixes and receives AWS-side prefixes according to the gateway design. BGP is usually preferred because adjacency and route withdrawal improve failure detection and because policy can influence paths. BGP does not guarantee correctness: wrong ASN, neighbor address, authentication, AS path, filters, maximum-prefix controls, or advertisements can still blackhole traffic.
Route decision rules to remember
- Longest prefix match is evaluated first.
- A VPC
localroute remains preferred over propagated hybrid routes, including an overlapping more-specific propagated route. - For an otherwise identical destination in a VPC table, listed static route targets take priority over a propagated VGW route.
- At a VGW, healthy path status matters before normal route attributes. For equal BGP prefixes, shorter AS path is preferred; MED can influence otherwise comparable paths.
- Direct Connect BGP, static VPN, and BGP VPN paths have documented VGW priorities; never infer priority from the order shown in a console table.
- A VGW selects one primary tunnel across its VPN connections; it does not perform VPN ECMP. TGW can use ECMP across dynamically routed equal-cost VPN paths.
- AWS can change the selected tunnel, including during endpoint maintenance. Customer devices should support asymmetric routing when possible.
For TGW, inspect two layers: the VPC subnet route sends remote traffic to the TGW, and the associated TGW route table selects the VPN attachment. Association chooses which table an attachment uses for its outbound lookup; propagation adds learned routes to selected tables. Mixing these terms is a common segmentation failure.
Address planning and IPv6
Inventory VPC, on-premises, branch, partner, container, service, and future acquisition prefixes before design. Overlap can make ordinary routing ambiguous. NAT can sometimes bridge overlap, but it adds state, translation ownership, logging, troubleshooting, and failover requirements; it is not a free checkbox in Site-to-Site VPN.
IPv4 tunnel inside addresses normally use unique /30 ranges from supported 169.254.0.0/16 space, excluding AWS-reserved blocks. They must be unique where the target gateway requires it and must not conflict on the customer device. IPv6 tunnel inside ranges use supported /126 ranges from fd00::/8 on TGW/Cloud WAN designs. Outside addresses can be IPv4 in the basic model; TGW and Cloud WAN also support IPv6 outside addresses and IPv4 or IPv6 inner traffic combinations. Validate the exact gateway and Region feature support before committing an IPv6 architecture.
The customer device usually needs a stable internet-routable address for public VPN. If NAT exists between the device and AWS, plan NAT-T and permit UDP 4500. If there is no NAT, IKE uses UDP 500 and encrypted data commonly uses ESP, IP protocol 50. Permit both AWS tunnel outside addresses in both directions. Do not open these protocols broadly to the internet when exact peers are known.
IKE, IPsec, authentication, and tunnel options
IKE negotiates algorithms, authenticates peers, performs Diffie-Hellman key exchange, and creates keying material. IPsec Encapsulating Security Payload protects data packets. Perfect Forward Secrecy uses fresh ephemeral key exchange so compromise of a long-term secret does not directly expose previous session keys.
| Option | Architecture question |
|---|---|
| IKEv1/IKEv2 | Which versions are permitted by security policy and device support? Prefer a deliberately selected compatible set rather than every default |
| Encryption/integrity/DH groups | What strong common proposal exists on both peers? Proposal order and mismatch are common Phase 1/2 faults |
| Phase 1/2 lifetime | Do peer values and AWS rekey behavior avoid interruption? Phase 2 lifetime must be lower than Phase 1 |
| Rekey margin/fuzz | Can both devices handle randomized AWS-initiated rekey without duplicate/stale SAs? |
| DPD timeout/action | Should AWS clear, do nothing, or restart after peer loss? Match this to routing convergence |
| Startup action | Does the customer initiate (Add, the default) or can AWS initiate (Start) for an IP-address CGW? |
| Replay window | Does the expected packet reordering require a tested window adjustment? |
| Local/remote CIDRs | Do proposals match the route-based design without pretending to be traffic enforcement? |
| Lifecycle control | Who reviews AWS Health maintenance and schedules eligible endpoint updates? |
PSKs are the default authentication method and are unique sensitive credentials per tunnel. Store, distribute, rotate, and redact them as secrets. AWS can store PSKs directly with the service or in Secrets Manager. Certificate authentication uses AWS Private CA and has CA-chain, certificate lifecycle, service-linked-role, compatibility, and cost implications. A certificate can improve rotation models, but it does not repair a broad routing or firewall design.
Never put PSKs, private keys, full downloaded device configurations, public endpoint inventories, or sensitive topology into tickets or course submissions. A downloaded vendor configuration is a starting template; validate algorithms, interface names, routing policy, NAT, HA behavior, software version, and local security standard before deployment.
MTU, MSS, throughput, and performance
IPsec adds headers, reducing payload size. Site-to-Site VPN has a maximum MTU of 1446 bytes and corresponding maximum IPv4 TCP MSS of 1406 bytes in the lowest-overhead supported case. NAT-T and some encryption/integrity combinations reduce them further. Jumbo frames and Path MTU Discovery are not supported by the service, so a silent MTU mismatch can produce the classic fault: ping works, TCP handshake works, but large transfers stall.
Use the official algorithm-specific table, then set tunnel/interface MTU and clamp TCP MSS on the customer device as appropriate. Test with do-not-fragment probes where the OS/path supports them, multiple payload sizes, TCP and UDP, and both directions. ICMP filtering can obscure useful evidence.
Current documented ceilings include up to 1.25 Gbps and 140,000 packets per second per standard tunnel, and up to 5 Gbps and 400,000 packets per second per large-bandwidth tunnel. They are ceilings, not promises. Packet size, traffic mix, encryption capacity of the customer device, single-flow limits, internet conditions, latency, loss, shaping, and application behavior determine observed throughput. TGW ECMP with dynamic routing can aggregate suitable flows across tunnels; one flow still follows one hash-selected path.
Measure baseline and failure-state throughput. A redundant design that meets capacity only while every path is healthy has no usable failure margin.
Security controls beyond encryption
The tunnel makes permitted networks reachable; it does not make them trusted. Apply least-privilege routes, security groups, NACLs where justified, customer firewall policy, workload host controls, and application authentication. Security groups can reference IP/CIDR sources across hybrid paths but do not magically learn business identity. Preserve source addresses unless the design intentionally translates them.
Use federated temporary AWS credentials and least-privilege EC2/VPC actions for VPN administration. Separate network design, change approval, credential access, and audit where risk requires it. CloudTrail records control-plane API changes; it does not log every packet. VPC Flow Logs show accepted/rejected inner flows at supported VPC interfaces, while VPN logs explain IKE, IPsec, DPD, and BGP state. Customer firewall/router logs complete the other side.
Hybrid DNS is separate. If on-premises clients must resolve private Route 53 names, or VPC workloads must resolve on-premises zones, design Route 53 Resolver inbound/outbound endpoints, rules, associations, firewalls, and return routes. An IPsec tunnel alone does not provide DNS forwarding.
Resilience and failover design
Minimum production posture is both tunnels configured, monitored, carrying correct routes, and tested. Better posture evaluates independent customer devices, power, edge switches, ISPs, carriers, public addresses, and sites. Draw failure domains rather than writing “HA” next to two lines.
| Failure | Expected control | Proof |
|---|---|---|
| One AWS tunnel endpoint | Second tunnel converges | timed packet loss, route change, application recovery |
| Customer router failure | Independent router/connection takes over | device isolation game day |
| ISP failure | Independent underlay remains | circuit/path evidence, not two tunnels on one ISP |
| BGP process/filter fault | adjacency alarm and alternate accepted routes | neighbor and route-table evidence |
| Bad route advertisement | prefix filters, max-prefix, rollback | rejected/withdrawn route evidence |
| PSK compromise | per-tunnel rotation procedure | staged rotation and redacted audit trail |
| AWS endpoint maintenance | both tunnels plus lifecycle/Health process | replacement status and maintenance test |
Tunnel endpoint lifecycle control can expose pending maintenance, the automatic-application deadline, and the last applied maintenance time. It gives scheduling control for eligible replacement events, not immunity from critical immediate updates. The normal resilience mechanism remains two functioning tunnels.
Read-only console and CLI investigation
In the VPC console, select the correct Region and inspect Site-to-Site VPN connections, Customer gateways, Virtual private gateways, and, where applicable, Transit Gateway attachments/route tables. Record the target gateway, routing type, category, outside address type, tunnel options, telemetry, static routes, tags, logging, and maintenance settings. Do not download or share the configuration because it can contain tunnel credentials.
Begin with a known identity and no mutation:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws ec2 describe-vpn-connections \
--query 'VpnConnections[].{Id:VpnConnectionId,State:State,Category:Category,Vgw:VpnGatewayId,Tgw:TransitGatewayId,Cgw:CustomerGatewayId,Static:Options.StaticRoutesOnly,Outside:Options.OutsideIpAddressType,Tunnels:VgwTelemetry}' \
--output json
aws ec2 describe-customer-gateways \
--query 'CustomerGateways[].{Id:CustomerGatewayId,State:State,Type:Type,Ip:IpAddress,Asn:BgpAsn,Certificate:CertificateArn}' \
--output table
aws ec2 describe-vpn-gateways \
--query 'VpnGateways[].{Id:VpnGatewayId,State:State,Asn:AmazonSideAsn,Attachments:VpcAttachments}' \
--output json
After selecting an explicitly owned connection, set its ID manually and review every field:
VPN_ID="vpn-replace-with-approved-id"
aws ec2 describe-vpn-connections --vpn-connection-ids "$VPN_ID" --output json
aws cloudwatch get-metric-data \
--metric-data-queries '[{"Id":"state","MetricStat":{"Metric":{"Namespace":"AWS/VPN","MetricName":"TunnelState","Dimensions":[{"Name":"VpnId","Value":"vpn-replace-with-approved-id"}]},"Period":300,"Stat":"Minimum"},"ReturnData":true}]' \
--start-time 2026-09-21T00:00:00Z --end-time 2026-09-21T01:00:00Z
Replace timestamps and IDs; do not copy the example unchanged. VgwTelemetry exposes each tunnel outside address, status, last status change, status message, and accepted route count. Compare that control-plane snapshot with CloudWatch TunnelState, TunnelDataIn, and TunnelDataOut, VPN logs, route tables, customer neighbor/routes, and an application probe. Small traffic metric values can occur from control traffic and are not application proof.
For lifecycle-controlled tunnels, use the approved connection ID and exact outside IP:
aws ec2 get-vpn-tunnel-replacement-status \
--vpn-connection-id "$VPN_ID" \
--vpn-tunnel-outside-ip-address "replace-with-approved-tunnel-ip"
For VGW designs inspect VPC route tables and propagation. For TGW designs inspect the VPC subnet route, attachment state, route-table association, propagation, and active/blackhole routes. Paginate results and preserve UTC timestamps.
Evidence-led troubleshooting
Start with the first failed layer and change one variable at a time.
| Symptom | Likely evidence path | Typical causes |
|---|---|---|
| Both tunnels down at Phase 1 | VPN log plus customer IKE/firewall capture | wrong outside IP/PSK/certificate, UDP 500/4500 or ESP blocked, proposal mismatch |
| Phase 1 up, Phase 2 down | child-SA proposals and traffic selectors | encryption/integrity/PFS/lifetime/local-remote CIDR mismatch |
| IPsec up, BGP down | inside IPs, ASN, neighbor state, TCP 179 inside tunnel | wrong peer/ASN, BGP policy, route-based device configuration |
| BGP up, zero/incorrect routes | advertised/received/accepted route tables | prefix filter, AS loop, max-prefix, missing advertisement |
| One-way traffic | forward and return route lookup at every hop | asymmetric firewall, missing return prefix, NAT, TGW table mismatch |
| Ping works, application fails | sized probes, MSS/MTU, DNS, SG/NACL, listener | fragmentation/PMTUD issue, blocked port, name resolution, host firewall |
| Tunnel flaps near regular interval | IKE/IPsec logs and lifetime timestamps | rekey/lifetime mismatch, duplicate IKE IDs, DPD behavior |
| Failover causes long outage | BGP timers/routes/session state/application retries | second tunnel unconfigured, stale route, asymmetric-intolerant device |
| Low throughput | per-tunnel metrics, PPS, CPU, loss, latency, flow count | device crypto ceiling, small packets, single flow, internet path, MTU |
| AWS says UP, user says down | inner path and application evidence | tunnel status proves too little; inspect route/security/DNS/workload |
Do not rotate PSKs, clear all security associations, restart both routers, disable route filters, or delete/recreate the VPN as the first diagnostic step. Those actions destroy state and evidence and can take down the remaining path. Capture timestamps, tunnel/BGP state, routes, metrics, logs, recent changes, and packet symptoms first.
No-create architecture exercise
Design connectivity for Northwind: VPC 10.40.0.0/16 in ap-south-1, data center 10.20.0.0/16, branch 10.30.0.0/16, two customer routers, and two independent ISPs. The VPC hosts a private API and Resolver endpoints. Requirements are 99.95% path availability, 500 Mbps normal throughput, 800 Mbps during one-router failure, dynamic failover, encrypted transport, 30-day negotiation logs, and no unapproved transitive routing between the data center and branch.
Produce:
- component/failure-domain diagram, including all four tunnel endpoints and underlays;
- IP/ASN/tunnel-inside allocation table with overlap proof;
- forward and return route table for each source/destination pair;
- VGW-versus-TGW decision and ECMP/path-policy explanation;
- IKE/IPsec/DPD/startup/rekey/MTU baseline and secret owner;
- security, DNS, logging, metric, alarm, retention, and escalation design;
- normal and degraded capacity calculation;
- tunnel, router, ISP, route-leak, MTU, and PSK-rotation game days;
- change/rollback plan with expected packet-loss and convergence budget; and
- current price estimate with every dependency.
Reject any design that relies on both AWS tunnels but one customer router/ISP while claiming end-to-end high availability. Also reject broad route advertisements without explicit branch/data-center segmentation evidence.
Cost and cleanup
Price by architecture and Region, not by memory. Potential charges include VPN connection-hours, public IPv4 addresses, data transfer out, TGW attachment-hours and data processing, CloudWatch Logs ingestion/storage/query, alarms, Secrets Manager secrets/API calls, Private CA, accelerated VPN's Global Accelerator hours and data-transfer premium, and large-bandwidth or VPN Concentrator options where selected. Customer router licenses, throughput tiers, internet circuits, support, monitoring, and staff are part of total cost even though they do not appear on the AWS bill.
The read-only path creates nothing, so cleanup evidence is an unchanged inventory. If a future approved lab creates a connection, hourly charges continue while it is provisioned; deleting only a route or shutting down a customer device does not stop all AWS charges. The cleanup plan must identify the VPN connection, CGW, VGW/TGW attachment dependencies, routes/propagations, logs, alarms, secrets/certificates, and retained evidence without deleting shared resources.
Knowledge check
- What is the difference between a customer gateway and a customer gateway device?
The customer gateway is an AWS control-plane representation; the device is the router/firewall that runs IKE, IPsec, and routing.
- Why are two tunnels not automatically end-to-end HA?
They protect AWS endpoints, but one customer device, ISP, power domain, or site can still fail both paths.
- Can an IPsec-up status prove application connectivity?
No. Routing, return path, security controls, DNS, host, listener, and application authorization remain separate.
- When can VPN ECMP aggregate paths?
With dynamic-routing Site-to-Site VPN attachments on TGW and suitable equal-cost routing; VGW does not provide VPN ECMP.
- Why can small pings work while HTTPS uploads stall?
IPsec overhead plus missing MTU/MSS handling can blackhole larger packets; the service does not support Path MTU Discovery.
- What evidence distinguishes a BGP failure from an IPsec failure?
IKE/Phase 2 state can be established while BGP neighbor state is not established; inspect both VPN logs/telemetry and customer BGP evidence.
Lesson acceptance
The lesson is complete only when the learner can:
- label every component, address, owner, and failure domain;
- trace a packet and return packet through all route/security layers;
- justify VGW/TGW, static/BGP, standard/large, accelerated/non-accelerated, and IPv4/IPv6 choices;
- explain and select tunnel authentication, crypto, DPD, startup, rekey, lifecycle, MTU, and logging options;
- prove both tunnels are configured and design genuine customer-side redundancy;
- interpret telemetry, CloudWatch metrics, VPN logs, BGP routes, and application tests without conflating them;
- diagnose all scenarios in the troubleshooting table using evidence before mutation;
- complete the Northwind design, capacity, game-day, rollback, and cost artifacts; and
- attest that no AWS or customer network resource was changed.
Official sources
- How AWS Site-to-Site VPN works
- Get started with Site-to-Site VPN
- Static and dynamic routing
- VPN route priority
- Tunnel options
- Tunnel authentication
- Customer gateway firewall rules
- Customer gateway device best practices
- Site-to-Site VPN logs
- CloudWatch VPN monitoring
- Tunnel endpoint lifecycle control
- Site-to-Site VPN quotas
- Accelerated Site-to-Site VPN
- AWS VPN pricing