Lesson 267 · AWS Learning Path

AWS 267: AWS Site-to-Site VPN

· Published · 17 min read

Labelled process diagram for AWS 267: On-premises prefixes to Two encrypted VPN tunnels to Virtual or transit gateway routing to VPC workload and return path, with decision, proof and rejection evidence.

Why this lesson matters

AWS Site-to-Site VPN creates managed IPsec tunnels between an AWS network and a customer network. It is often the fastest way to establish encrypted hybrid connectivity, a practical backup for Direct Connect, or a production path when internet variability and supported throughput fit the requirement. It is not “a private cable,” and a green tunnel icon does not prove that applications can communicate.

A working design needs correct components on both sides, compatible IKE/IPsec settings, non-overlapping prefixes, forward and return routes, firewall permission, an appropriate MTU, monitoring, two configured tunnels, and named owners. The AWS team owns the AWS endpoint; the customer owns the customer gateway device, internet path, routing policy, firewall, application path, and usually the hardest half of diagnosis.

This lesson builds the complete mental model from Linux routing concepts to architect-level choices and operations. It is read-only: you will inspect supplied or explicitly owned evidence and create a local design, not a billable VPN connection.

Outcomes

By the end, you can:

  • distinguish a customer gateway resource from the physical or virtual customer gateway device;
  • choose a virtual private gateway (VGW), Transit Gateway (TGW), Cloud WAN, acceleration, or a different connectivity service;
  • explain IKE Phase 1, IPsec Phase 2, inside and outside addresses, PSK/certificate authentication, NAT traversal, DPD, and rekeying;
  • design static or BGP routing and trace both forward and return paths;
  • explain two-tunnel behavior, VGW active-path selection, TGW ECMP, and true device/site redundancy;
  • select IPv4/IPv6, standard/large bandwidth, MTU, MSS, cryptographic, and logging options;
  • inspect configuration, routes, telemetry, metrics, logs, and maintenance status with read-only commands;
  • diagnose tunnel-down, BGP-down, one-way, intermittent, low-throughput, and failover faults;
  • design a safe game day, credential rotation, maintenance, and rollback plan; and
  • identify every cost component and prove that this no-create lesson changed nothing.

From Linux routing to a managed VPN

If you know Linux networking, map familiar concepts carefully:

Linux/network conceptSite-to-Site VPN equivalentImportant difference
IPsec/IKE peerAWS tunnel endpoint and customer gateway deviceAWS manages its endpoint; you configure both tunnels on your device
ip routeVPC route table, TGW route table, customer routing tableSeveral independent tables must agree
FRR/BIRD BGP neighborBGP peer over tunnel inside addressesTunnel can be IPsec-up while BGP is down
iptables/nftablesCustomer firewall plus security groups/NACLsSGs are stateful; NACLs and many customer ACLs are stateless
Interface MTU/TCP MSS clampTunnel overhead controlVPN has no jumbo frames and does not support Path MTU Discovery
tcpdump and daemon logsCustomer packet capture plus AWS VPN logs/metricsAWS logs negotiations and states, not customer payloads

Encryption protects traffic in transit. It does not authorize an application, resolve DNS, fix overlapping addresses, inspect malware, or automatically advertise every network.

Components and terminology

ComponentWhat it representsCommon mistake
Customer gateway deviceRouter, firewall, SD-WAN appliance, or software endpoint in the customer networkCalling the AWS record the router
Customer gateway (CGW)AWS resource containing the device's outside IP/certificate identity and BGP ASNExpecting it to forward packets by itself
VPN connectionTwo managed IPsec tunnels and their optionsConfiguring only Tunnel 1
Virtual private gatewayRegional AWS VPN concentrator attached to one VPCExpecting ECMP or IPv6 inner traffic
Transit GatewayRegional routing hub for many VPC/on-premises attachmentsForgetting TGW association/propagation and VPC subnet routes
Tunnel outside addressesInternet- or private-underlay endpoints carrying IKE/IPsecConfusing them with routed workload addresses
Tunnel inside addressesPoint-to-point addresses used by tunnel/BGP peersAdvertising them as application networks
Local/remote network CIDRsIKE Phase 2 proposal rangesTreating them as enforced firewall rules; AWS uses route-based VPNs

Each VPN connection supplies two tunnels for AWS-side endpoint resilience. That does not remove a single customer device, single power domain, single ISP, or single site as a failure domain. Higher availability commonly needs two customer devices and independent connections/underlays, with tested routing convergence.

Packet and control flow

application in on-premises prefix 10.20.0.0/16
  -> local route/VRF
  -> customer firewall and selected customer gateway device
  -> IKE/IPsec security associations
  -> encrypted outer packet across internet or private underlay
  -> AWS tunnel endpoint
  -> VGW or TGW route decision
  -> VPC subnet route, NACL, security group, workload
  -> return route through the same valid hybrid routing domain

Separate four proofs:

  1. IKE Phase 1: peers authenticate and create an IKE security association.
  2. IPsec Phase 2: peers negotiate child/IPsec security associations for protected traffic.
  3. Routing: static routes or BGP install usable prefixes through a healthy tunnel.
  4. Application: DNS, network controls, host firewall, listener, TLS, and application authorization succeed.

“Tunnel UP” mainly addresses the first two. “BGP ESTABLISHED” adds a routing adjacency but does not prove accepted prefixes, route-table selection, security controls, or the application.

Choose the AWS-side gateway

RequirementBetter starting pointReason/caveat
One VPC, simple IPv4 hybrid pathVGWLower conceptual overhead; one VGW attaches to a VPC at a time
Many VPCs/accounts or segmented routing domainsTGWCentral hub with attachment association and propagation controls
Multiple active tunnels for aggregate bandwidthTGW with dynamic routing and ECMPVGW does not support VPN ECMP
IPv6 inner or IPv6 outer tunnelsTGW or Cloud WANVGW supports the basic IPv4 outer/IPv4 inner model
Up to 5-Gbps large-bandwidth tunnel optionTGW or Cloud WANLarge option is not available on VGW
Better internet-path consistency from a distant siteAccelerated VPN on TGWUses Global Accelerator; adds restrictions and cost
Private IP VPN over Direct ConnectTGW plus Direct Connect designUnderlay and overlay are separate; Direct Connect dependencies remain
Global managed core and policyCloud WAN evaluationBroader service, policy, and cost decision than one VPN

Accelerated VPN enters the AWS network at a nearby edge, but it is not available on VGW, requires NAT-T, must be selected when the connection is created, and requires the customer device to initiate IKE. It cannot simply be toggled on for an existing connection. Measure latency, loss, jitter, and cost before selecting it.

Static routing or BGP

Static routing lists customer prefixes in the VPN connection and configures explicit routes on the customer device. It suits a small, stable prefix set and devices without BGP. Every topology change becomes coordinated configuration, and liveness/failover depends on device mechanisms rather than route withdrawal.

Dynamic routing establishes BGP over each tunnel's inside addresses. The customer device advertises reachable on-premises prefixes and receives AWS-side prefixes according to the gateway design. BGP is usually preferred because adjacency and route withdrawal improve failure detection and because policy can influence paths. BGP does not guarantee correctness: wrong ASN, neighbor address, authentication, AS path, filters, maximum-prefix controls, or advertisements can still blackhole traffic.

Route decision rules to remember

  • Longest prefix match is evaluated first.
  • A VPC local route remains preferred over propagated hybrid routes, including an overlapping more-specific propagated route.
  • For an otherwise identical destination in a VPC table, listed static route targets take priority over a propagated VGW route.
  • At a VGW, healthy path status matters before normal route attributes. For equal BGP prefixes, shorter AS path is preferred; MED can influence otherwise comparable paths.
  • Direct Connect BGP, static VPN, and BGP VPN paths have documented VGW priorities; never infer priority from the order shown in a console table.
  • A VGW selects one primary tunnel across its VPN connections; it does not perform VPN ECMP. TGW can use ECMP across dynamically routed equal-cost VPN paths.
  • AWS can change the selected tunnel, including during endpoint maintenance. Customer devices should support asymmetric routing when possible.

For TGW, inspect two layers: the VPC subnet route sends remote traffic to the TGW, and the associated TGW route table selects the VPN attachment. Association chooses which table an attachment uses for its outbound lookup; propagation adds learned routes to selected tables. Mixing these terms is a common segmentation failure.

Address planning and IPv6

Inventory VPC, on-premises, branch, partner, container, service, and future acquisition prefixes before design. Overlap can make ordinary routing ambiguous. NAT can sometimes bridge overlap, but it adds state, translation ownership, logging, troubleshooting, and failover requirements; it is not a free checkbox in Site-to-Site VPN.

IPv4 tunnel inside addresses normally use unique /30 ranges from supported 169.254.0.0/16 space, excluding AWS-reserved blocks. They must be unique where the target gateway requires it and must not conflict on the customer device. IPv6 tunnel inside ranges use supported /126 ranges from fd00::/8 on TGW/Cloud WAN designs. Outside addresses can be IPv4 in the basic model; TGW and Cloud WAN also support IPv6 outside addresses and IPv4 or IPv6 inner traffic combinations. Validate the exact gateway and Region feature support before committing an IPv6 architecture.

The customer device usually needs a stable internet-routable address for public VPN. If NAT exists between the device and AWS, plan NAT-T and permit UDP 4500. If there is no NAT, IKE uses UDP 500 and encrypted data commonly uses ESP, IP protocol 50. Permit both AWS tunnel outside addresses in both directions. Do not open these protocols broadly to the internet when exact peers are known.

IKE, IPsec, authentication, and tunnel options

IKE negotiates algorithms, authenticates peers, performs Diffie-Hellman key exchange, and creates keying material. IPsec Encapsulating Security Payload protects data packets. Perfect Forward Secrecy uses fresh ephemeral key exchange so compromise of a long-term secret does not directly expose previous session keys.

OptionArchitecture question
IKEv1/IKEv2Which versions are permitted by security policy and device support? Prefer a deliberately selected compatible set rather than every default
Encryption/integrity/DH groupsWhat strong common proposal exists on both peers? Proposal order and mismatch are common Phase 1/2 faults
Phase 1/2 lifetimeDo peer values and AWS rekey behavior avoid interruption? Phase 2 lifetime must be lower than Phase 1
Rekey margin/fuzzCan both devices handle randomized AWS-initiated rekey without duplicate/stale SAs?
DPD timeout/actionShould AWS clear, do nothing, or restart after peer loss? Match this to routing convergence
Startup actionDoes the customer initiate (Add, the default) or can AWS initiate (Start) for an IP-address CGW?
Replay windowDoes the expected packet reordering require a tested window adjustment?
Local/remote CIDRsDo proposals match the route-based design without pretending to be traffic enforcement?
Lifecycle controlWho reviews AWS Health maintenance and schedules eligible endpoint updates?

PSKs are the default authentication method and are unique sensitive credentials per tunnel. Store, distribute, rotate, and redact them as secrets. AWS can store PSKs directly with the service or in Secrets Manager. Certificate authentication uses AWS Private CA and has CA-chain, certificate lifecycle, service-linked-role, compatibility, and cost implications. A certificate can improve rotation models, but it does not repair a broad routing or firewall design.

Never put PSKs, private keys, full downloaded device configurations, public endpoint inventories, or sensitive topology into tickets or course submissions. A downloaded vendor configuration is a starting template; validate algorithms, interface names, routing policy, NAT, HA behavior, software version, and local security standard before deployment.

MTU, MSS, throughput, and performance

IPsec adds headers, reducing payload size. Site-to-Site VPN has a maximum MTU of 1446 bytes and corresponding maximum IPv4 TCP MSS of 1406 bytes in the lowest-overhead supported case. NAT-T and some encryption/integrity combinations reduce them further. Jumbo frames and Path MTU Discovery are not supported by the service, so a silent MTU mismatch can produce the classic fault: ping works, TCP handshake works, but large transfers stall.

Use the official algorithm-specific table, then set tunnel/interface MTU and clamp TCP MSS on the customer device as appropriate. Test with do-not-fragment probes where the OS/path supports them, multiple payload sizes, TCP and UDP, and both directions. ICMP filtering can obscure useful evidence.

Current documented ceilings include up to 1.25 Gbps and 140,000 packets per second per standard tunnel, and up to 5 Gbps and 400,000 packets per second per large-bandwidth tunnel. They are ceilings, not promises. Packet size, traffic mix, encryption capacity of the customer device, single-flow limits, internet conditions, latency, loss, shaping, and application behavior determine observed throughput. TGW ECMP with dynamic routing can aggregate suitable flows across tunnels; one flow still follows one hash-selected path.

Measure baseline and failure-state throughput. A redundant design that meets capacity only while every path is healthy has no usable failure margin.

Security controls beyond encryption

The tunnel makes permitted networks reachable; it does not make them trusted. Apply least-privilege routes, security groups, NACLs where justified, customer firewall policy, workload host controls, and application authentication. Security groups can reference IP/CIDR sources across hybrid paths but do not magically learn business identity. Preserve source addresses unless the design intentionally translates them.

Use federated temporary AWS credentials and least-privilege EC2/VPC actions for VPN administration. Separate network design, change approval, credential access, and audit where risk requires it. CloudTrail records control-plane API changes; it does not log every packet. VPC Flow Logs show accepted/rejected inner flows at supported VPC interfaces, while VPN logs explain IKE, IPsec, DPD, and BGP state. Customer firewall/router logs complete the other side.

Hybrid DNS is separate. If on-premises clients must resolve private Route 53 names, or VPC workloads must resolve on-premises zones, design Route 53 Resolver inbound/outbound endpoints, rules, associations, firewalls, and return routes. An IPsec tunnel alone does not provide DNS forwarding.

Resilience and failover design

Minimum production posture is both tunnels configured, monitored, carrying correct routes, and tested. Better posture evaluates independent customer devices, power, edge switches, ISPs, carriers, public addresses, and sites. Draw failure domains rather than writing “HA” next to two lines.

FailureExpected controlProof
One AWS tunnel endpointSecond tunnel convergestimed packet loss, route change, application recovery
Customer router failureIndependent router/connection takes overdevice isolation game day
ISP failureIndependent underlay remainscircuit/path evidence, not two tunnels on one ISP
BGP process/filter faultadjacency alarm and alternate accepted routesneighbor and route-table evidence
Bad route advertisementprefix filters, max-prefix, rollbackrejected/withdrawn route evidence
PSK compromiseper-tunnel rotation procedurestaged rotation and redacted audit trail
AWS endpoint maintenanceboth tunnels plus lifecycle/Health processreplacement status and maintenance test

Tunnel endpoint lifecycle control can expose pending maintenance, the automatic-application deadline, and the last applied maintenance time. It gives scheduling control for eligible replacement events, not immunity from critical immediate updates. The normal resilience mechanism remains two functioning tunnels.

Read-only console and CLI investigation

In the VPC console, select the correct Region and inspect Site-to-Site VPN connections, Customer gateways, Virtual private gateways, and, where applicable, Transit Gateway attachments/route tables. Record the target gateway, routing type, category, outside address type, tunnel options, telemetry, static routes, tags, logging, and maintenance settings. Do not download or share the configuration because it can contain tunnel credentials.

Begin with a known identity and no mutation:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws ec2 describe-vpn-connections \
  --query 'VpnConnections[].{Id:VpnConnectionId,State:State,Category:Category,Vgw:VpnGatewayId,Tgw:TransitGatewayId,Cgw:CustomerGatewayId,Static:Options.StaticRoutesOnly,Outside:Options.OutsideIpAddressType,Tunnels:VgwTelemetry}' \
  --output json
aws ec2 describe-customer-gateways \
  --query 'CustomerGateways[].{Id:CustomerGatewayId,State:State,Type:Type,Ip:IpAddress,Asn:BgpAsn,Certificate:CertificateArn}' \
  --output table
aws ec2 describe-vpn-gateways \
  --query 'VpnGateways[].{Id:VpnGatewayId,State:State,Asn:AmazonSideAsn,Attachments:VpcAttachments}' \
  --output json

After selecting an explicitly owned connection, set its ID manually and review every field:

VPN_ID="vpn-replace-with-approved-id"
aws ec2 describe-vpn-connections --vpn-connection-ids "$VPN_ID" --output json

aws cloudwatch get-metric-data \
  --metric-data-queries '[{"Id":"state","MetricStat":{"Metric":{"Namespace":"AWS/VPN","MetricName":"TunnelState","Dimensions":[{"Name":"VpnId","Value":"vpn-replace-with-approved-id"}]},"Period":300,"Stat":"Minimum"},"ReturnData":true}]' \
  --start-time 2026-09-21T00:00:00Z --end-time 2026-09-21T01:00:00Z

Replace timestamps and IDs; do not copy the example unchanged. VgwTelemetry exposes each tunnel outside address, status, last status change, status message, and accepted route count. Compare that control-plane snapshot with CloudWatch TunnelState, TunnelDataIn, and TunnelDataOut, VPN logs, route tables, customer neighbor/routes, and an application probe. Small traffic metric values can occur from control traffic and are not application proof.

For lifecycle-controlled tunnels, use the approved connection ID and exact outside IP:

aws ec2 get-vpn-tunnel-replacement-status \
  --vpn-connection-id "$VPN_ID" \
  --vpn-tunnel-outside-ip-address "replace-with-approved-tunnel-ip"

For VGW designs inspect VPC route tables and propagation. For TGW designs inspect the VPC subnet route, attachment state, route-table association, propagation, and active/blackhole routes. Paginate results and preserve UTC timestamps.

Evidence-led troubleshooting

Start with the first failed layer and change one variable at a time.

SymptomLikely evidence pathTypical causes
Both tunnels down at Phase 1VPN log plus customer IKE/firewall capturewrong outside IP/PSK/certificate, UDP 500/4500 or ESP blocked, proposal mismatch
Phase 1 up, Phase 2 downchild-SA proposals and traffic selectorsencryption/integrity/PFS/lifetime/local-remote CIDR mismatch
IPsec up, BGP downinside IPs, ASN, neighbor state, TCP 179 inside tunnelwrong peer/ASN, BGP policy, route-based device configuration
BGP up, zero/incorrect routesadvertised/received/accepted route tablesprefix filter, AS loop, max-prefix, missing advertisement
One-way trafficforward and return route lookup at every hopasymmetric firewall, missing return prefix, NAT, TGW table mismatch
Ping works, application failssized probes, MSS/MTU, DNS, SG/NACL, listenerfragmentation/PMTUD issue, blocked port, name resolution, host firewall
Tunnel flaps near regular intervalIKE/IPsec logs and lifetime timestampsrekey/lifetime mismatch, duplicate IKE IDs, DPD behavior
Failover causes long outageBGP timers/routes/session state/application retriessecond tunnel unconfigured, stale route, asymmetric-intolerant device
Low throughputper-tunnel metrics, PPS, CPU, loss, latency, flow countdevice crypto ceiling, small packets, single flow, internet path, MTU
AWS says UP, user says downinner path and application evidencetunnel status proves too little; inspect route/security/DNS/workload

Do not rotate PSKs, clear all security associations, restart both routers, disable route filters, or delete/recreate the VPN as the first diagnostic step. Those actions destroy state and evidence and can take down the remaining path. Capture timestamps, tunnel/BGP state, routes, metrics, logs, recent changes, and packet symptoms first.

No-create architecture exercise

Design connectivity for Northwind: VPC 10.40.0.0/16 in ap-south-1, data center 10.20.0.0/16, branch 10.30.0.0/16, two customer routers, and two independent ISPs. The VPC hosts a private API and Resolver endpoints. Requirements are 99.95% path availability, 500 Mbps normal throughput, 800 Mbps during one-router failure, dynamic failover, encrypted transport, 30-day negotiation logs, and no unapproved transitive routing between the data center and branch.

Produce:

  1. component/failure-domain diagram, including all four tunnel endpoints and underlays;
  2. IP/ASN/tunnel-inside allocation table with overlap proof;
  3. forward and return route table for each source/destination pair;
  4. VGW-versus-TGW decision and ECMP/path-policy explanation;
  5. IKE/IPsec/DPD/startup/rekey/MTU baseline and secret owner;
  6. security, DNS, logging, metric, alarm, retention, and escalation design;
  7. normal and degraded capacity calculation;
  8. tunnel, router, ISP, route-leak, MTU, and PSK-rotation game days;
  9. change/rollback plan with expected packet-loss and convergence budget; and
  10. current price estimate with every dependency.

Reject any design that relies on both AWS tunnels but one customer router/ISP while claiming end-to-end high availability. Also reject broad route advertisements without explicit branch/data-center segmentation evidence.

Cost and cleanup

Price by architecture and Region, not by memory. Potential charges include VPN connection-hours, public IPv4 addresses, data transfer out, TGW attachment-hours and data processing, CloudWatch Logs ingestion/storage/query, alarms, Secrets Manager secrets/API calls, Private CA, accelerated VPN's Global Accelerator hours and data-transfer premium, and large-bandwidth or VPN Concentrator options where selected. Customer router licenses, throughput tiers, internet circuits, support, monitoring, and staff are part of total cost even though they do not appear on the AWS bill.

The read-only path creates nothing, so cleanup evidence is an unchanged inventory. If a future approved lab creates a connection, hourly charges continue while it is provisioned; deleting only a route or shutting down a customer device does not stop all AWS charges. The cleanup plan must identify the VPN connection, CGW, VGW/TGW attachment dependencies, routes/propagations, logs, alarms, secrets/certificates, and retained evidence without deleting shared resources.

Knowledge check

  1. What is the difference between a customer gateway and a customer gateway device?

The customer gateway is an AWS control-plane representation; the device is the router/firewall that runs IKE, IPsec, and routing.

  1. Why are two tunnels not automatically end-to-end HA?

They protect AWS endpoints, but one customer device, ISP, power domain, or site can still fail both paths.

  1. Can an IPsec-up status prove application connectivity?

No. Routing, return path, security controls, DNS, host, listener, and application authorization remain separate.

  1. When can VPN ECMP aggregate paths?

With dynamic-routing Site-to-Site VPN attachments on TGW and suitable equal-cost routing; VGW does not provide VPN ECMP.

  1. Why can small pings work while HTTPS uploads stall?

IPsec overhead plus missing MTU/MSS handling can blackhole larger packets; the service does not support Path MTU Discovery.

  1. What evidence distinguishes a BGP failure from an IPsec failure?

IKE/Phase 2 state can be established while BGP neighbor state is not established; inspect both VPN logs/telemetry and customer BGP evidence.

Lesson acceptance

The lesson is complete only when the learner can:

  • label every component, address, owner, and failure domain;
  • trace a packet and return packet through all route/security layers;
  • justify VGW/TGW, static/BGP, standard/large, accelerated/non-accelerated, and IPv4/IPv6 choices;
  • explain and select tunnel authentication, crypto, DPD, startup, rekey, lifecycle, MTU, and logging options;
  • prove both tunnels are configured and design genuine customer-side redundancy;
  • interpret telemetry, CloudWatch metrics, VPN logs, BGP routes, and application tests without conflating them;
  • diagnose all scenarios in the troubleshooting table using evidence before mutation;
  • complete the Northwind design, capacity, game-day, rollback, and cost artifacts; and
  • attest that no AWS or customer network resource was changed.

Official sources

Advertisement