AWS 293: Route 53 and Global Accelerator failover
Why this lesson matters
Route 53 and AWS Global Accelerator can both direct clients away from an unhealthy endpoint, but they act at different layers and times. Route 53 changes DNS answers. Resolvers and clients can continue using an answer they already cached. Global Accelerator exposes static anycast addresses and steers new network connections through the AWS global network. Existing connections do not magically become new healthy sessions.
Neither service makes the recovery Region ready. Traffic steering can turn a partial incident into a complete outage if the secondary has stale data, missing secrets, insufficient capacity, blocked partner traffic, or a health endpoint that says “OK” while orders fail.
An architect must therefore connect client behavior, protocol, DNS, network sessions, health evidence, application dependencies, data authority, recovery capacity, failover authorization, and controlled failback. This lesson builds that complete chain from Linux DNS knowledge.
What you will be able to do
By the end, you can:
- trace a client through recursive DNS or a Global Accelerator listener to a Regional endpoint;
- distinguish DNS control from anycast connection steering;
- choose appropriate Route 53 routing and health-evaluation patterns;
- explain Global Accelerator listeners, endpoint groups, traffic dials, endpoint weights, and health behavior;
- account for TTL, caches, connection reuse, client IP, TLS, firewall, and protocol requirements;
- design health signals that represent user success without causing avoidable failovers;
- block failover when application, dependency, data, or capacity readiness is absent;
- define active-passive, warm-standby, and active-active operating procedures;
- test failover and failback with objective evidence; and
- compare cost, quotas, security, and operational ownership.
Before you start
- This is a no-create lesson. Do not alter DNS, health checks, accelerator dials, endpoint weights, or endpoints.
- Use the reserved names
api.example.comandtcp.example.comin your design. Do not query or test someone else's domain or IP address. - Existing-account inspection requires explicit authorization. Hosted-zone names, records, accelerator IPs, endpoint IDs, and health data can be sensitive.
- Use a non-root federated identity and redact account-specific values before sharing evidence.
- Do not confuse a tabletop or local DNS cache experiment with proof of production recovery.
- Record service Region support, quotas, prices, and endpoint requirements at design time because these can change.
1. Start with the client-observed path
Route 53 path
application
-> operating-system/stub cache
-> recursive resolver cache
-> Route 53 authoritative name servers
-> routing policy plus recorded health state
-> DNS answer and TTL
-> client opens a connection to the answer
-> Regional edge/load balancer/application/data
Route 53 does not proxy the application connection. It answers a DNS query. A recursive resolver might serve thousands of clients from one cached answer, and some clients, frameworks, proxies, or appliances can cache beyond the intended behavior. Route 53 health checks run independently of each query. A DNS answer change affects lookups after relevant caches expire; it does not move an established TCP connection.
Global Accelerator path
application resolves accelerator DNS name or uses static address
-> nearest healthy AWS edge advertisement
-> listener matches protocol and port
-> client location, endpoint-group dial, endpoint weight and health
-> AWS global network
-> Regional ALB/NLB/EC2/EIP endpoint
-> application and data
A standard accelerator supplies static anycast IP addresses and a DNS name. AWS advertises the same addresses from multiple edge locations. The listener handles configured TCP or UDP port ranges. Endpoint groups are Regional; registered endpoints within them receive traffic according to health and weight. Changes and health events primarily affect new connections. Application reconnect behavior therefore remains part of RTO.
Global Accelerator is not a cache, CDN, DNS authority, web application firewall, database replication system, or application load balancer. It can sit in front of suitable Regional endpoints and complement those services.
2. Recovery truth comes before traffic steering
Define the failure scope and recovery contract first:
| Question | Required answer |
|---|---|
| Failure scope | Instance, Availability Zone, Region, dependency, network path, bad release, data corruption, or control plane? |
| Service objective | Availability target, maximum detection time, RTO, RPO, and acceptable degraded behavior? |
| Data authority | Which Region accepts writes before, during, and after failover? |
| Recovery shape | Backup/restore, pilot light, warm standby, active-passive, or active-active? |
| Capacity | How quickly can secondary capacity handle expected and reconnect traffic? |
| Dependencies | Identity, secrets, keys, certificates, queues, databases, third parties, observability, and operations? |
| Authority | Who can declare disaster, steer traffic, freeze writes, and initiate failback? |
Traffic must not move merely because the primary endpoint is unreachable from one probe. Define a readiness gate such as:
regional_ready = application_ready
AND critical_dependencies_ready
AND data_within_rpo
AND write_authority_granted
AND capacity_ready
AND security_and_observability_ready
Some of these signals can be automated. Data authority and business risk may require a human decision. Separate detection automation from failover authorization where a false positive can create split-brain writes or material loss.
3. Route 53 records and routing policies
A public hosted zone holds internet DNS records; a private hosted zone answers through associated VPC DNS paths. The same record name and type can have multiple records when the routing policy permits it. Record identifiers distinguish members of that set.
Relevant policies are:
- Failover: primary and secondary, intended for active-passive routing.
- Weighted: split answers by relative weight, useful for controlled exposure or active-active distribution.
- Latency: choose among configured Regional records based on measured AWS-network latency from the resolver's apparent location.
- Geolocation: choose from the query origin's geographic location, with a default record for unmatched locations.
- Geoproximity: shift geographic boundaries around resource locations using bias; effect is relative and should be changed cautiously.
- IP-based: map client CIDR groups to answers when known source networks require deterministic policy.
- Multivalue answer: return up to eight healthy records selected approximately at random; it is not a replacement for a load balancer.
- Simple: one resource or a set without policy-level health routing; do not use it when health-based DR behavior is required.
Routing policy answers “which eligible record?” Health configuration answers “is this record eligible?” Data and business readiness remain outside DNS unless deliberately represented by a trustworthy signal.
Alias and non-alias health
For a supported AWS alias target, EvaluateTargetHealth can inherit health from the referenced target or record tree. It is not the same as attaching an arbitrary Route 53 endpoint health check. For a non-alias record, associate a Route 53 health check when appropriate.
Trace complex alias trees from leaf to parent. A parent may appear healthy because at least one child remains healthy even when the specific capability a user needs is broken. Avoid circular, opaque, or over-nested policies.
Route 53 failover has availability-preserving behavior that beginners often miss. If primary and secondary are both considered unhealthy, Route 53 can return the primary. If no health check is configured for the secondary, Route 53 treats it as an always-available fallback when the primary is unhealthy. A secondary must therefore be independently monitored and operationally gated even when policy semantics would return it.
4. Route 53 health-check designs
Route 53 supports three broad health sources:
- Endpoint health checks over supported protocols and settings from distributed health checkers.
- Calculated health checks that combine child health checks against a threshold.
- CloudWatch-alarm-based health checks for a metric and alarm in the required account/Region context.
An endpoint probe should use a dedicated path and validate a meaningful but bounded dependency chain. A shallow liveness check asks whether a process can respond. Readiness asks whether it should receive user traffic. Deep health might test data and dependencies, but it can add load, expose sensitive failure detail, and cause correlated failover when one shared noncritical dependency fails.
Good design commonly separates:
- load balancer target health for one replica;
- Regional readiness for the whole service;
- business synthetic transactions for end-to-end evidence; and
- data-lag and write-authority gates controlled outside a public endpoint.
Protect health paths without blocking legitimate AWS checkers. Use HTTPS and correct certificate names where supported, restrict returned information, prevent caching, and keep the check inexpensive and idempotent. Monitor the health-check system itself. A firewall change, TLS expiry, or DNS error can produce a false outage.
Set interval and failure threshold from a detection budget, not impatience. Faster detection can increase sensitivity to transient loss. Recovery time is approximately:
detection + health aggregation + DNS answer change + cache expiry
+ client reconnect/retry + secondary startup/capacity + application recovery
Low TTL affects only the cache portion and increases authoritative queries. Lower it before planned traffic changes and wait for previous TTLs to age. Raising it during an incident does not remove already cached answers.
5. Global Accelerator objects and decisions
A standard accelerator contains:
- static IP addresses and an accelerator DNS name;
- one or more listeners with TCP or UDP port ranges and optional client affinity;
- Regional endpoint groups; and
- endpoints with health and relative weights.
Supported standard endpoint types and exact requirements must be verified in current documentation. Common targets include Application Load Balancers, Network Load Balancers, EC2 instances, and Elastic IP addresses. Prefer load balancer endpoints for multi-instance Regional availability unless a direct endpoint requirement is justified.
Traffic dial versus endpoint weight
The traffic dial limits the share of traffic that Global Accelerator would already direct to an endpoint group, usually a Region. It is not a percentage of all global listener traffic. Reducing one Region's dial causes eligible new traffic to use other endpoint groups.
An endpoint weight divides traffic among endpoints inside one endpoint group. Weights are relative, not percentages. A zero weight can remove normal traffic, but availability behavior and health rules still need review.
Changing a dial affects new connections; it does not terminate existing ones. For long-lived WebSocket, TCP, database, gaming, or UDP application state, define drain, reconnect, retry, and session migration behavior. A quick control-plane update can still produce a long user transition.
Health and last-resort behavior
For ALB and NLB endpoints, Global Accelerator uses the load balancer's health information. For EC2 and Elastic IP endpoints, endpoint-group health-check configuration is relevant. Check protocol, port, path, interval, threshold, and security access.
If no healthy endpoints are available, a standard accelerator can route to all endpoints to preserve possible availability. Do not assume “all unhealthy” means traffic is dropped. Design an explicit maintenance, isolation, or application-level rejection mechanism if sending traffic to an unhealthy fleet is unsafe.
Client IP and security
Client IP preservation depends on endpoint type and configuration. Supported cases include ALB, EC2, and NLB with security groups; Elastic IP and NLB without security groups do not support preservation in the same way. Verify the current matrix before relying on source addresses for authorization, logging, rate limits, or allowlists.
Never use source IP as the only user identity. When preservation is enabled, security groups and network ACLs must allow the intended client traffic model. When it is not, applications need the supported method to obtain client context, if available, and must trust only controlled intermediaries.
TLS usually terminates at the selected endpoint architecture, not at the accelerator. Certificates must cover the client hostname, and every Regional endpoint must implement equivalent protocol, cipher, SNI, and renewal behavior.
6. Route 53 versus Global Accelerator
| Requirement | Route 53 direction | Global Accelerator direction |
|---|---|---|
| DNS-based routing across many endpoint kinds | Strong fit | Only supported accelerator endpoints |
| Static client-facing IP allowlist | DNS answers can change | Strong fit with static anycast addresses |
| TCP or UDP steering with rapid new-connection reaction | Limited by DNS/client behavior | Strong fit for configured listener ports |
| HTTP content caching | Use CloudFront or another cache | Not a cache |
| Geographic, resolver-based, IP-based, or DNS weighted policy | Rich DNS policies | Proximity, health, Regional dials, endpoint weights |
| Private DNS failover | Supported policies in private hosted zones | Standard accelerator is an internet-facing global entry model; evaluate exact design |
| Existing connection migration | Neither guarantees transparent migration | New connections steer; existing sessions need application behavior |
| Non-AWS arbitrary endpoint records | DNS can represent them with valid health design | Endpoint types are constrained |
| Client uses literal approved IPs | DNS cannot help that client | Static accelerator addresses can fit |
They can be combined. For example, Route 53 can alias a friendly application name to Global Accelerator. Do not create two independent health systems that fight each other. State which layer owns Regional failover and which owns only naming.
Choose CloudFront instead when the core need is HTTP/S edge caching, origin shielding, web delivery features, or edge functions. Choose an ALB when balancing HTTP/S only within a Region. Choose an NLB when Regional transport-layer behavior or static zonal addresses fit. These services can be layers in one design, but each must have a clear responsibility.
7. Active-passive and active-active designs
Active-passive
The secondary may be cold, pilot light, or warm standby. A DNS secondary record or healthy accelerator endpoint does not prove adequate capacity. Define activation steps, scaling time, data promotion, write-authority fencing, secret/certificate readiness, partner allowlists, and verification before traffic eligibility.
Avoid automatic failover when data promotion is manual or when both Regions could accept conflicting writes. Automation can detect and prepare; a controlled gate grants write authority and triggers steering.
Active-active
Both Regions serve traffic, so failover is a capacity redistribution rather than a cold start. Verify that either remaining Region can handle displaced demand plus reconnect surge. Define global data consistency, home-Region ownership, conflict resolution, idempotency, session behavior, asynchronous lag, and degradation when one Region is absent.
Active-active compute does not make a single-Region database active-active. Draw the data path for every write and read, including stale-read and partition behavior.
8. Prevent false positives and cascading failure
Use these controls:
- multiple independent observations before declaring a Regional failure;
- thresholds and evaluation windows aligned to the RTO;
- a health endpoint that excludes optional dependencies;
- synthetic business checks observed separately from routing probes;
- data lag, capacity, and write-authority interlocks;
- suppression during approved maintenance or controlled deployment;
- circuit breakers and load shedding at the application layer;
- alarms for flapping and frequent health transitions;
- a minimum stable period before failback; and
- manual approval for high-consequence steering where justified.
Failover can overload a healthy secondary, saturate cross-Region data paths, exhaust database connections, trigger autoscaling too slowly, or cause retry storms. Capacity testing must include the failure load and reconnect behavior, not just steady traffic.
9. Failover and failback runbook
Define each task with owner, command or console path, expected evidence, timeout, failure action, and rollback counterpart.
Failover gates
- Confirm user impact from at least two signals.
- Classify whether the failure is endpoint, AZ, Region, shared dependency, deployment, data, or client-specific.
- Verify secondary application, dependencies, data lag, security, observability, capacity, quota, and partner readiness.
- Fence primary writes or otherwise establish one write authority.
- Record the RPO boundary and business acceptance.
- Change the approved Route 53 record or accelerator dial/endpoint state.
- Observe DNS answers or new connection distribution from multiple client networks.
- Validate critical business transactions and resulting data.
- Watch saturation, errors, latency, backlog, security, cost, and data replication.
Failback gates
Failback is a planned migration, not an automatic mirror image. Repair and test the original Region, synchronize changes, decide which copy is authoritative, test capacity and dependencies, choose a low-risk window, shift a controlled portion of new traffic, validate, then increase. Preserve a return path to the DR Region until stability criteria pass.
Automatic failback can flap traffic between marginal Regions. Require a stable observation period and explicit approval. Reconcile data before changing write authority.
10. Read-only inspection and Linux evidence
Only inspect owned resources. Route 53 is global in the CLI. Global Accelerator API calls use its documented home endpoint; the CLI examples specify us-west-2.
aws sts get-caller-identity --query Arn --output text
aws route53 list-hosted-zones \
--query 'HostedZones[].{Name:Name,Private:Config.PrivateZone}'
aws route53 list-health-checks \
--query 'HealthChecks[].{Id:Id,Type:HealthCheckConfig.Type,Disabled:HealthCheckConfig.Disabled}'
aws globalaccelerator list-accelerators --region us-west-2 \
--query 'Accelerators[].{Name:Name,Enabled:Enabled,Status:Status,IpSets:IpSets}'
For an explicitly owned accelerator, use its ARN to list listeners, then each listener ARN to list endpoint groups, and each group ARN to describe endpoint configurations. Do not paste a placeholder into a command and assume the result is evidence.
Local DNS tools reveal different layers:
dig api.example.com A
dig api.example.com AAAA
dig api.example.com A +trace
resolvectl query api.example.com
getent ahosts api.example.com
dig against a chosen resolver shows its answer and reported TTL. +trace follows public delegation and bypasses the usual recursive path, so it does not represent an application's cache. getent uses the operating system's name-service configuration and is closer to what many Linux applications see. None proves every user's resolver behavior.
11. Guided design and failure workshop
Design recovery for two services:
Service A: public HTTPS ordering API at api.example.com, ALBs in ap-south-1 and eu-west-1, warm standby, Aurora-compatible Regional data design supplied by the exercise, 10,000 requests per second peak, 15-minute RPO, 20-minute RTO, partners that allowlist source and destination IPs, and mobile clients with connection pooling.
Service B: long-lived TCP telemetry at tcp.example.com, NLBs in both Regions, devices configured with either a DNS name or two fixed IPs, sessions lasting up to six hours, reconnect backoff, and duplicate telemetry tolerance only when an idempotency key is present.
Complete these artifacts:
- Client-to-data paths for Route 53 and Global Accelerator candidates.
- Protocol, IP, TLS, cache, connection, client-IP, and endpoint requirements.
- Route 53 failover, weighted, and latency policy comparisons for Service A.
- A standard accelerator model with listeners, groups, dials, endpoints, weights, and health.
- A data-authority and Regional-readiness state machine.
- Detection and RTO budget math.
- Health-check design with false-positive and shared-dependency analysis.
- Capacity calculation for normal, failover, and reconnect surge.
- Primary and secondary unhealthy behavior, including last-resort routing.
- Stepwise failover and failback runbooks with authorization.
- Tests from at least three resolver/client networks and two connection ages.
- Failure injections: DNS cache beyond TTL, TLS expiry, blocked health checker, stale secondary data, zero healthy accelerator endpoints, overloaded database, and partner allowlist omission.
- Cost and quota inventory.
- Final service-by-service decision with rejected alternatives.
There is no mandatory single answer. Service A may use Route 53, Global Accelerator, CloudFront plus Regional origins, or layered controls depending on the supplied constraints. Service B's fixed-IP and transport requirements often favor Global Accelerator, but session duration and reconnect behavior remain unsolved unless the application handles them.
12. Cost, quotas, security, and ownership
Route 53 cost can include hosted zones, DNS query categories, health checks, optional features, traffic policies, and logs or monitoring around the design. Global Accelerator cost includes each accelerator and data transfer through it, plus Regional endpoints, load balancers, compute, transfer, logs, and security services. Price both normal and disaster traffic.
Review quotas for records, health checks, accelerators, listeners, port ranges, endpoint groups, endpoints, and related load-balancer or network resources. A DR design that depends on a quota increase requested during the incident is not ready.
Protect DNS and accelerator changes with least privilege, MFA-backed emergency access, change control, CloudTrail, AWS Config where supported, and alerts on configuration changes. Separate routine observation from permission to steer production traffic. Protect the registrar and domain delegation, which sit outside an individual record change.
Assign owners for DNS, accelerator, network, load balancers, application health, data replication, capacity, security, business acceptance, incident command, cost, and failback. Test access and backup contacts during game days.
Diagnose a failing design
| Observation | Likely cause | Evidence and correction |
|---|---|---|
| Some users remain on primary | Recursive/client cache or reused connection | Compare resolver TTL, OS behavior, process cache, and connection age; design retry/drain |
| DNS changed but API still fails | Secondary application, dependency, TLS, security, or data not ready | Trace a full transaction and Regional-readiness gate |
| Health check fails while users succeed | Probe blocked, certificate/name mismatch, wrong path, or threshold too sensitive | Inspect checker reachability and exact response; repair probe design |
| Health check passes while orders fail | Check is too shallow or misses critical dependency | Separate liveness, readiness, and synthetic business signals |
| GA dial 50 does not yield half of all traffic | Dial applies only to traffic already directed to that group and only new connections | Inspect Regional distribution and connection lifecycle |
| Weight zero endpoint receives traffic in an outage | Availability-preserving behavior or no healthy alternative | Review endpoint health/weight semantics and explicit isolation |
| Failover creates duplicate writes | Both Regions retained write authority or retries lack idempotency | Fence writes, use stable keys, and reconcile |
| Failback immediately fails again | Original cause or data/capacity readiness was not proven | Require repair evidence, synchronization, canary shift, and stable period |
Knowledge check
- Does Route 53 move an existing TCP connection?
No. It changes eligible DNS answers; clients must resolve and connect again.
- What does a low TTL guarantee?
Only the intended cache lifetime for compliant caching layers, not instant universal change.
- What is the Global Accelerator traffic dial percentage applied to?
Traffic that the accelerator would already direct to that endpoint group, primarily for new connections.
- Why must the secondary have its own health and readiness evidence?
Routing-policy fallback can return it even when its real application or data is unready.
- What happens when a standard accelerator has no healthy endpoints?
It can route to all endpoints to preserve possible availability, so unsafe endpoints need explicit isolation.
- Is client IP preservation universal in Global Accelerator?
No. It depends on endpoint type and configuration and must be verified.
- What must happen before Regional write traffic moves?
Data must meet the RPO, the old writer must be fenced as designed, and one authority must grant target writes.
- Why is failback usually manual and staged?
Data, dependencies, capacity, and the original failure must be repaired and proven without causing traffic flapping.
Lesson acceptance
You may continue when your submission contains:
- exact DNS and accelerator request paths;
- a recovery contract and Regional-readiness state machine;
- correct Route 53 policy, alias, health, cache, and all-unhealthy reasoning;
- correct Global Accelerator listener, group, dial, weight, health, connection, and client-IP reasoning;
- a service-specific choice rather than a universal winner;
- data-authority, capacity, false-positive, and cascading-failure controls;
- timed failover and staged failback runbooks with owners;
- evidence from cache, resolver, new/old connection, application, data, and business layers;
- failure-injection results and corrected design;
- complete cost and quota inventory; and
- least-privilege change authority and a game-day schedule.