AWS 016: High availability, fault tolerance, RTO and RPO
The problem
A manager asks for "zero downtime and zero data loss" without describing business impact or budget. The team adds backups and calls the system highly available, but has never restored one.
Availability, fault tolerance, backup, and disaster recovery address related but different outcomes.
Learning outcomes
You will be able to:
- distinguish reliability, resilience, high availability, fault tolerance, backup, and disaster recovery;
- calculate availability and downtime;
- define Recovery Time Objective and Recovery Point Objective;
- select a recovery approach from business requirements;
- identify dependencies and tests needed to prove recovery.
Terms
Reliability
The ability of a workload to perform its intended function correctly and consistently over time.
Resilience
The ability to withstand, respond to, and recover from disruption.
High availability
A design that reduces service interruption through redundancy, detection, failover, repair, and operation across appropriate failure domains.
Fault tolerance
The ability to continue the required function despite specified component faults, often with little or no visible interruption. It generally requires more redundancy and coordination than recovery after failure.
Backup
A recoverable copy or recovery point. A backup is useful only when it is protected, retained, and restorable.
Disaster recovery
The strategy and procedures that restore a workload after a defined disaster. It includes infrastructure, configuration, applications, data, identity, networking, dependencies, people, and communication.
Availability calculation
AWS defines availability as time available for use divided by total time.
availability = available time / total time
For a 30-day month:
30 x 24 x 60 = 43,200 minutes
At 99.9 percent availability:
unavailable fraction = 0.001
43,200 x 0.001 = 43.2 minutes
At 99.99 percent:
43,200 x 0.0001 = 4.32 minutes
State the measurement period and what counts as available. A site returning errors or unusable responses is not truly available merely because a process is running.
RTO and RPO
Recovery Time Objective is the maximum acceptable delay from interruption to restored service.
Recovery Point Objective is the maximum acceptable age of the recovered data, expressed as how much time of data loss the business can tolerate.
last usable recovery point incident restored service
| | |
|<------ possible data loss: RPO ------>| |
|<--- downtime: RTO ------->|
Example:
- incident at 14:00;
- RPO 15 minutes means recovery must reach a usable point no earlier than 13:45;
- RTO 60 minutes means required service must be restored by 15:00.
RPO does not mean backup frequency alone. Replication lag, backup success, corruption, retention, and restore capability affect the achieved result.
Recovery strategies
AWS describes a spectrum:
| Strategy | Normal recovery environment | General trade-off |
|---|---|---|
| Backup and restore | data and deployable configuration retained | lower steady cost, longer recovery |
| Pilot light | core data and critical components active | faster recovery, more operation |
| Warm standby | scaled-down functional workload | still faster, higher steady cost |
| Multi-site active/active | multiple sites actively serve | shortest interruption potential, greatest consistency and operational complexity |
Do not attach generic RTO numbers without testing your workload.
High availability is not disaster recovery
A Multi-AZ design can handle an Availability Zone failure. It might not handle:
- destructive application writes;
- credential compromise;
- accidental deletion replicated everywhere;
- Region-wide requirements;
- a faulty deployment;
- dependency failure;
- business-location disruption.
Backups can protect data but do not automatically keep the service running. Use multiple controls for different failure modes.
Practical business-impact exercise
Create:
mkdir -p "$HOME/nitwings-aws/evidence/aws-016"
Create recovery-objectives.md for:
The portal delivers lessons all day. During a two-hour final exam window, an outage blocks submissions. Progress is written continually. Video can be regenerated, but submitted answers cannot. Instructors can tolerate four hours of publishing downtime outside exams.
Include:
- critical user journeys;
- time-dependent impact;
- availability target and measurement period;
- RTO and RPO for lesson browsing;
- RTO and RPO for exam submission;
- RTO and RPO for publishing;
- failure events covered;
- recovery strategy per component;
- dependencies;
- restore and failover test;
- owner and evidence;
- cost trade-off.
Expected: exam submission receives stricter objectives than publishing. Irreplaceable answers need a smaller RPO than regenerable media.
Test, do not assume
A recovery test verifies:
- detection;
- decision authority;
- backup or replica usability;
- infrastructure and configuration deployment;
- identity and secrets;
- network and DNS;
- application function;
- data correctness;
- actual RTO and RPO;
- failback or steady-state decision.
A successful snapshot job is not a successful restore test.
Common misconceptions
- Redundancy without health detection and failover is not high availability.
- Two copies in one failure domain do not protect against that domain.
- Replication can copy corruption.
- A service SLA is not the application's achieved availability.
- Zero RTO/RPO claims require precise scope and substantial design evidence.
- More availability usually adds cost and operational complexity.
Knowledge check
- What does RTO measure?
- What does RPO measure?
- Does a backup prove recovery?
- Can a highly available system still need disaster recovery?
- Why should objectives differ by business function?
Expected answers: restoration delay; tolerable data-loss time; no; yes; impact and value differ.
Completion gate
Pass when calculations are correct and recovery-objectives.md defines differentiated, testable objectives, strategies, dependencies, failure scope, owners, and cost consequences.
No AWS resources were created.