AWS 165: Route 53 health checks and failover
Why this lesson matters
Design endpoint, calculated, or CloudWatch-alarm health checks and understand checker reachability, thresholds, DNS failover, and false positives.
DNS failover removes unhealthy answers for new resolution; it does not terminate existing connections, repair data or guarantee zero downtime. Health signal design must avoid both routing users to a broken system and failing away from a healthy system because the checker cannot reach it.
What you will be able to do
By the end, you can:
- explain route 53 health checks and failover in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Design endpoint, calculated, or CloudWatch-alarm health checks and understand checker reachability, thresholds, DNS failover, and false positives. |
| Scope and boundary | The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Route 53 health checks and failover. |
| Evidence of success | Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Route 53 health checks and failover. |
| Cost model | Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design. |
| Safe rejection rule | Avoid probing a page that stays green while dependencies fail or using very low TTLs without considering query cost and resolver behavior. |
How the request flows
+---------------------------+
| Health checker or alarm |
+---------------------------+
|
v
+----------------------+
| Health state |
+----------------------+
|
v
+-----------------------------+
| Failover record selection |
+-----------------------------+
|
v
+------------------------------------------+
| Cached DNS answer and application test |
+------------------------------------------+
For Route 53 health checks and failover, the important boundary is this: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Route 53 health checks and failover. Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Route 53 health checks and failover. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use health-based DNS failover for endpoints that can operate independently and have a meaningful externally observable health signal. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid probing a page that stays green while dependencies fail or using very low TTLs without considering query cost and resolver behavior. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
Health signal types and path
Endpoint health checks originate from Route 53 checker IP ranges and test an IP/domain, port and HTTP/HTTPS/TCP behavior, optionally path/string/latency. Firewalls must permit checker sources without opening the application broadly. HTTPS certificate/hostname/SNI behavior and redirects/content can produce false results. The check should validate the minimum critical serving path but avoid slow/nonessential dependencies that cause cascading failover.
Calculated health checks combine child checks by threshold and can model quorum/maintenance; inversion and disabled checks need governance. CloudWatch-alarm health uses a metric/alarm in its Region and has explicit behavior when data is insufficient. Evaluate Target Health on an alias uses supported AWS resource health (for example load-balancer targets) without a standalone checker. Private endpoints cannot be probed directly by public checkers; use an application metric/alarm or calculated design.
Health-check interval, failure threshold and checker consensus shape detection time, but DNS TTL, resolver behavior, connection pools and application retry add recovery delay. Recovery/failback should use a longer stability window than failover to prevent flapping. A secondary must have deployed code/config/secrets/certificates, scaled capacity, current data within RPO and tested dependencies before DNS sends users.
Route 53 first applies routing policy then health rules according to record type/group. If all eligible records are unhealthy, it may return unhealthy records rather than no answer; design application-level circuit breakers. Failover records need matching name/type and PRIMARY/SECONDARY identifiers. Alias chains inherit health in specific ways - draw every record, not only the top alias.
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Use the Console service search and open Route 53, Health checks and Hosted zone failover records; confirm the account and Region before reading the page.
- Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
- Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
- Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws route53 list-health-checks --query 'HealthChecks[].{Id:Id,Type:HealthCheckConfig.Type,FQDN:HealthCheckConfig.FullyQualifiedDomainName,Port:HealthCheckConfig.Port,Path:HealthCheckConfig.ResourcePath}' --output table
aws route53 get-health-check-status --health-check-id replace-with-health-check-id --output json
Expected interpretation
Checker observations prove the configured probe from Route 53 locations, not a full user transaction. DNS failover is also affected by TTL and recursive caching.
Practical work
Design active-passive DNS for two regional endpoints. Define health path, expected string, interval, threshold, alarm, evaluate-target-health choice, TTL, test, false-positive guard, and failback approval.
Build a UTC timeline: fault, first failed probe, unhealthy consensus, authoritative answer change, resolver cache expiry, first secondary request, data validation, recovery and approved failback. Test checker-IP denial, TLS hostname failure, endpoint returns 200 with broken dependency, alarm insufficient data, cached primary, secondary under-capacity/stale data, all-unhealthy behavior and flapping. Compare Route 53 DNS failover, Global Accelerator connection routing and ARC readiness/routing controls.
Diagnose this topic from its own evidence
Inspect health checker status by checker/Region, failure reason, endpoint response/TLS and CloudWatch alarm before DNS answers. If health is unhealthy but manual curl works, compare source network, Host header, path, TLS and expected string. If DNS still returns primary, query authoritative servers and inspect TTL/cache. If secondary receives traffic but fails, the error is DR readiness, not Route 53. Never force failover until data and capacity owners accept the consequence.
Cost and cleanup
Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Design endpoint, calculated, or CloudWatch-alarm health checks and understand checker reachability, thresholds, DNS failover, and false positives.
- Which scope or ownership boundary must be proved first?
Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Route 53 health checks and failover.
- What evidence is strong enough to accept the result?
Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Route 53 health checks and failover.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid probing a page that stays green while dependencies fail or using very low TTLs without considering query cost and resolver behavior.
- Which cost dimensions and retained resources need an owner?
Expected direction: Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Lesson acceptance
Pass when the learner distinguishes endpoint/calculated/alarm/evaluate-target health, quantifies probe+TTL+connection timing, tests false positive/all-unhealthy behavior and validates secondary data/capacity/failback. Fail if health check equals application recovery, private endpoint is assigned a public probe, or failover is called zero-RPO.