Lesson 016 · AWS Learning Path

AWS 016: High availability, fault tolerance, RTO and RPO

· Published · 4 min read

A service spans resilient locations with recovery checkpoints and timelines representing RTO and RPO

The problem

A manager asks for "zero downtime and zero data loss" without describing business impact or budget. The team adds backups and calls the system highly available, but has never restored one.

Availability, fault tolerance, backup, and disaster recovery address related but different outcomes.

Learning outcomes

You will be able to:

  1. distinguish reliability, resilience, high availability, fault tolerance, backup, and disaster recovery;
  2. calculate availability and downtime;
  3. define Recovery Time Objective and Recovery Point Objective;
  4. select a recovery approach from business requirements;
  5. identify dependencies and tests needed to prove recovery.

Terms

Reliability

The ability of a workload to perform its intended function correctly and consistently over time.

Resilience

The ability to withstand, respond to, and recover from disruption.

High availability

A design that reduces service interruption through redundancy, detection, failover, repair, and operation across appropriate failure domains.

Fault tolerance

The ability to continue the required function despite specified component faults, often with little or no visible interruption. It generally requires more redundancy and coordination than recovery after failure.

Backup

A recoverable copy or recovery point. A backup is useful only when it is protected, retained, and restorable.

Disaster recovery

The strategy and procedures that restore a workload after a defined disaster. It includes infrastructure, configuration, applications, data, identity, networking, dependencies, people, and communication.

Availability calculation

AWS defines availability as time available for use divided by total time.

availability = available time / total time

For a 30-day month:

30 x 24 x 60 = 43,200 minutes

At 99.9 percent availability:

unavailable fraction = 0.001
43,200 x 0.001 = 43.2 minutes

At 99.99 percent:

43,200 x 0.0001 = 4.32 minutes

State the measurement period and what counts as available. A site returning errors or unusable responses is not truly available merely because a process is running.

RTO and RPO

Recovery Time Objective is the maximum acceptable delay from interruption to restored service.

Recovery Point Objective is the maximum acceptable age of the recovered data, expressed as how much time of data loss the business can tolerate.

last usable recovery point       incident               restored service
          |                         |                           |
          |<------ possible data loss: RPO ------>|           |
                                    |<--- downtime: RTO ------->|

Example:

  • incident at 14:00;
  • RPO 15 minutes means recovery must reach a usable point no earlier than 13:45;
  • RTO 60 minutes means required service must be restored by 15:00.

RPO does not mean backup frequency alone. Replication lag, backup success, corruption, retention, and restore capability affect the achieved result.

Recovery strategies

AWS describes a spectrum:

StrategyNormal recovery environmentGeneral trade-off
Backup and restoredata and deployable configuration retainedlower steady cost, longer recovery
Pilot lightcore data and critical components activefaster recovery, more operation
Warm standbyscaled-down functional workloadstill faster, higher steady cost
Multi-site active/activemultiple sites actively serveshortest interruption potential, greatest consistency and operational complexity

Do not attach generic RTO numbers without testing your workload.

High availability is not disaster recovery

A Multi-AZ design can handle an Availability Zone failure. It might not handle:

  • destructive application writes;
  • credential compromise;
  • accidental deletion replicated everywhere;
  • Region-wide requirements;
  • a faulty deployment;
  • dependency failure;
  • business-location disruption.

Backups can protect data but do not automatically keep the service running. Use multiple controls for different failure modes.

Practical business-impact exercise

Create:

mkdir -p "$HOME/nitwings-aws/evidence/aws-016"

Create recovery-objectives.md for:

The portal delivers lessons all day. During a two-hour final exam window, an outage blocks submissions. Progress is written continually. Video can be regenerated, but submitted answers cannot. Instructors can tolerate four hours of publishing downtime outside exams.

Include:

  • critical user journeys;
  • time-dependent impact;
  • availability target and measurement period;
  • RTO and RPO for lesson browsing;
  • RTO and RPO for exam submission;
  • RTO and RPO for publishing;
  • failure events covered;
  • recovery strategy per component;
  • dependencies;
  • restore and failover test;
  • owner and evidence;
  • cost trade-off.

Expected: exam submission receives stricter objectives than publishing. Irreplaceable answers need a smaller RPO than regenerable media.

Test, do not assume

A recovery test verifies:

  1. detection;
  2. decision authority;
  3. backup or replica usability;
  4. infrastructure and configuration deployment;
  5. identity and secrets;
  6. network and DNS;
  7. application function;
  8. data correctness;
  9. actual RTO and RPO;
  10. failback or steady-state decision.

A successful snapshot job is not a successful restore test.

Common misconceptions

  • Redundancy without health detection and failover is not high availability.
  • Two copies in one failure domain do not protect against that domain.
  • Replication can copy corruption.
  • A service SLA is not the application's achieved availability.
  • Zero RTO/RPO claims require precise scope and substantial design evidence.
  • More availability usually adds cost and operational complexity.

Knowledge check

  1. What does RTO measure?
  2. What does RPO measure?
  3. Does a backup prove recovery?
  4. Can a highly available system still need disaster recovery?
  5. Why should objectives differ by business function?

Expected answers: restoration delay; tolerable data-loss time; no; yes; impact and value differ.

Completion gate

Pass when calculations are correct and recovery-objectives.md defines differentiated, testable objectives, strategies, dependencies, failure scope, owners, and cost consequences.

No AWS resources were created.

Official sources

Advertisement