Lesson 318 · AWS Learning Path

AWS 318: Reliability pillar

· Published · 9 min read

Labelled process diagram for AWS 318: Business objectives and dependencies to Reliability pillar review to Failure and recovery improvements to Measured availability, RPO, RTO, and learning, with decision, proof and...

Why this lesson matters

Reliability is the ability of a workload to perform its intended function correctly and consistently when expected. The AWS Well-Architected Reliability pillar groups practices into Foundations, Workload Architecture, Change Management, and Failure Management.

Reliability is not “multi-AZ” written beside every service. A workload can run in three Availability Zones and still fail because all nodes share one database limit, DNS record, deployment, identity dependency, corrupt data set, or human procedure. Redundancy without independent failure boundaries and tested recovery creates confidence rather than resilience.

This lesson starts with user outcomes and failure budgets, then proves capacity, distributed-system behavior, change safety, backup/restore, disaster recovery, and game-day evidence.

What you will be able to do

By the end, you can:

  • define availability, durability, resilience, fault tolerance, RTO, RPO, MTPD, and degraded service;
  • map critical user journeys and dependency failure domains;
  • manage service quotas and network constraints with failover headroom;
  • choose service boundaries and distributed-system patterns;
  • design timeouts, retries, backoff, jitter, idempotency, queues, and load shedding;
  • scale for normal, peak, backlog, and failure conditions;
  • deploy changes progressively with compatible rollback;
  • select backup, pilot-light, warm-standby, and active patterns from objectives;
  • test component, AZ, Region, dependency, data-corruption, and control-plane failures;
  • distinguish infrastructure recovery from business recovery; and
  • conduct an evidence-based Reliability pillar review.

Before you start

  • Complete operational/security prerequisites. Recovery that bypasses authorization or cannot be operated is not reliable.
  • Use a fictional or approved workload. Do not inject production failure in this T0 review.
  • Record account, Region, version, time, source, and owner for evidence.
  • Never publish topology, backup locations, recovery credentials, customer impact, or incident details.
  • Historical uptime does not prove future recovery. Direct tests are required.

1. Define the outcome and objectives

For each critical user journey define:

  • correct result and maximum latency;
  • demand and peak;
  • availability target and measurement window;
  • acceptable degraded behavior;
  • MTPD or maximum tolerable outage;
  • RTO from disruption to restored outcome;
  • RPO or maximum acceptable data loss measured in time/business records;
  • consistency and one-writer requirements;
  • dependency assumptions; and
  • owner/decision authority.
business process
 -> user journey
 -> service-level indicator
 -> target/error budget
 -> architecture and recovery investment

Availability is successful service time/requests under a defined scope. Durability is probability data remains intact. RTO is recovery time objective; RPO is recovery point objective. Neither is proven by a service marketing number. Measure actual recovery time capability and recovery point capability in exercises.

2. Foundations: quotas and network

Inventory quotas in every account and Region used for normal operation and failover. Include resource count, API rate, throughput, concurrency, payload, routes, IP addresses, NAT ports, DNS, certificates, KMS, IAM, and third-party limits.

Set alarms before exhaustion and maintain failure headroom. If Region B normally runs at 60 percent and must absorb Region A, its quota/capacity may be insufficient. Quota increase lead time belongs in recovery readiness.

Network topology must provide redundant paths where objectives require them. Analyze VPC CIDRs, subnets/AZs, routes, NAT, load balancers, DNS, Transit Gateway, Direct Connect/VPN, firewalls, resolver endpoints, and external providers in both directions.

A second VPN tunnel that terminates on the same customer router/power/circuit is not independent. Multi-AZ subnets using one central appliance can retain a single failure domain.

aws service-quotas list-services --max-results 20 --output table
aws service-quotas list-service-quotas --service-code ec2 --output table
aws ec2 describe-availability-zones --all-availability-zones --output table
aws ec2 describe-route-tables --output table

3. Design workload service architecture

Segment by business domain and failure isolation, not fashionable size. A monolith can be reliable when well-designed; hundreds of synchronous microservices can multiply failure.

For each service contract define:

  • request/response or event schema/version;
  • timeout and latency budget;
  • idempotency and duplicate behavior;
  • ordering/consistency;
  • authentication/authorization;
  • retryable/permanent errors;
  • throttling/backpressure;
  • availability/degraded behavior;
  • observability; and
  • compatibility/deprecation.

Reduce hard synchronous dependencies. Use queues for temporal decoupling where delay is acceptable. Keep essential purchase path separate from optional recommendations/reviews so noncritical failure degrades rather than blocks.

4. Apply distributed-system resilience patterns

Timeouts

Every remote call needs a bounded connect and request timeout based on end-to-end latency budget. An absent or excessive timeout consumes threads/connections until cascade.

Retries

Retry only transient/idempotent work, with exponential backoff, jitter, maximum attempts, and total deadline. Retries multiply load during an outage. If five layers retry three times, one request can amplify dramatically. Choose one responsible layer.

Idempotency

Use a stable operation key and store outcome so duplicate requests do not duplicate payment, message, or mutation. Idempotency needs scope and expiry.

Circuit breaking and load shedding

Stop calling a dependency that is predictably failing, then probe recovery. Reject lower-priority or excess work before all capacity is exhausted. Return clear degraded responses.

Queues and poison messages

Set visibility/acknowledgement, retry count, dead-letter path, retention, age alarm, deduplication/ordering as required, replay authority, and downstream idempotency. A DLQ without owner/runbook is delayed data loss.

Static stability

Pre-provision critical dependencies needed during impairment and avoid recovery that depends on an unavailable control plane. Cache only with expiry and consistency rules.

5. Scale for demand and failure

Capacity model includes:

normal peak + growth + failure absorption + deployment overlap
+ retry/replay/backlog catch-up + safety margin

Measure the bottleneck end to end: client, edge, load balancer, compute, thread/connection pool, database, cache, queue, stream partition, storage, downstream API, and quota.

Autoscaling is delayed feedback. Startup, warm-up, metric lag, cooldown, and dependency capacity matter. Scale on demand/backlog/latency signals that predict user harm, not CPU alone. Test scale-in so in-flight work drains safely.

Protect stateful systems with connection limits, pooling/proxies, partition/hot-key design, read/write separation where correct, and bounded concurrency. More application instances can overwhelm a fixed database faster.

6. Manage change reliably

All changes can fail: code, configuration, infrastructure, schema, data, certificate, DNS, policy, dependency, and manual operation.

Use version control, automated tests, immutable artifacts, staged environments, canary/linear/blue-green deployment, health gates, and automatic stop. Preserve previous compatible version.

Database evolution uses expand-and-contract:

  1. add backward-compatible schema;
  2. deploy code that handles old/new;
  3. migrate/backfill with checkpoints;
  4. verify;
  5. stop old writers;
  6. remove old schema later.

Rollback has a deadline. After irreversible writes or external effects, compensation or fix-forward may be safer. Define this before deployment.

Monitor change failure rate, rollback success, time to detect, and user outcome. Correlation is strongest with precise deployment/configuration timelines.

7. Design for failure

Map failure scopes:

  • process/container/instance;
  • rack/AZ;
  • Region;
  • account/identity/control plane;
  • network/DNS/certificate;
  • software/configuration deployment;
  • dependency/provider;
  • overload/abuse;
  • data corruption/deletion/ransomware; and
  • operator/procedure.

For each, specify detection, automatic behavior, degraded state, containment, recovery, data authority, failback, owner, and tested evidence.

Multi-AZ is the default for many production objectives, but inspect service semantics. AZ-independent compute does not help if storage, NAT, directory, or database is single-AZ.

Multi-Region adds replication lag, consistency, conflict, DNS/client caching, secrets/certificates, quotas, deployment, observability, support, and cost. Use it only when business objectives justify operational complexity.

8. Back up and restore trustworthy data

Backup design includes scope, frequency, application consistency, retention, encryption, cross-account/Region isolation, immutability, legal hold, catalog, monitoring, and deletion authority.

RPO is constrained by the last usable clean recovery point, not the newest backup job. Continuous replication can replicate corruption or ransomware. Preserve point-in-time history and detection-aware recovery.

Restore test steps:

  1. select approved point;
  2. restore into isolated environment;
  3. verify keys/roles/network/dependencies;
  4. validate schema/integrity and malware/corruption;
  5. start application in controlled order;
  6. perform business reconciliation;
  7. measure recovery point/time;
  8. destroy test environment safely; and
  9. remediate gaps.

9. Select disaster-recovery pattern

PatternCharacteristics
Backup and restoreLowest steady cost, longer recovery, strongest need for restore automation
Pilot lightCritical data/core kept ready, application capacity created during event
Warm standbyScaled-down functional environment, scale and route traffic during event
Active-active/multi-siteMultiple serving locations, highest consistency/operations complexity

Pattern names do not determine RTO/RPO. Prove full timeline: declare, authorize, restore/promote, scale, configure, validate, route, client converge, business reconcile, communicate. Failback is a separate risky migration.

Prevent split brain with one-writer authority, fencing, quorum, or conflict rules. DNS failover changes where new lookups go; it does not stop old sessions or prove data readiness.

10. Test resilience

Progress:

  1. architecture/tabletop review;
  2. component failure in development;
  3. isolated restore;
  4. staging game day;
  5. production-safe fault injection;
  6. full recovery/failback exercise where justified.

An experiment states steady state, hypothesis, targets, blast radius, stop conditions, alarms, roles, rollback, and evidence. Never inject failure without independent safety controls.

Test instance loss, AZ dependency, database failover, DNS, expired certificate, quota exhaustion, downstream latency, queue backlog, corrupt deploy, operator absence, backup restore, and Region recovery according to risk.

11. Diagnose from evidence

SymptomEvidenceResponse
Multi-AZ workload still failsfull dependency/failure-domain mapRemove shared DNS/network/state/control dependency.
Retry stormcall tree, retry layers/attempts, latency, rateStop amplification, centralize bounded retry with jitter.
Autoscaling adds nodes but errors risedatabase connections, downstream quota, hot keyProtect bottleneck and shed load.
Failover succeeds but users still failDNS/cache/sessions, data readiness, auth/cert, journey SLIValidate complete client-to-data path.
Latest backup cannot restorejob logs, catalog, key, corruption point, restore testUse earlier clean point and remediate backup/control gap.
Rollback cannot rundata/schema/external effects, artifact, IAM/control planeInvoke compensation/fix-forward plan and preserve evidence.
Queue age grows after recoveryconsumer throughput, poison messages, retries, quotaQuarantine poison work and scale safe catch-up.

12. Reliability workshop

Review a three-tier order system with RTO 30 minutes, RPO 5 minutes, two-AZ normal operation, and warm standby in another Region.

Submit:

  1. journeys/objectives/error budgets;
  2. dependency/failure-domain map;
  3. quota and failover-capacity model;
  4. bidirectional network proof;
  5. service contracts;
  6. timeout/retry/idempotency/load-shedding policy;
  7. queue/backlog/poison/replay design;
  8. autoscaling and database protection;
  9. deployment/schema/rollback design;
  10. failure matrix;
  11. backup/clean restore evidence plan;
  12. DR declaration/failover/failback runbook;
  13. one-writer/fencing design;
  14. seven diagnostic cases;
  15. game-day plan; and
  16. Well-Architected findings/backlog.

Cost and cleanup

Reliability costs include redundant capacity, cross-AZ/Region transfer, replication, backups, warm environments, tests, observability, support, and engineering. Compare cost with quantified outage impact and objectives. Do not remove tested recovery capability for nominal savings without business approval.

T0 creates nothing. Later experiments must delete temporary restore instances, volumes, snapshots copied for tests, DNS records, logs, fault templates, alarms, and credentials only after evidence/retention approval.

Knowledge check

  1. Four areas? Foundations, Workload Architecture, Change Management, Failure Management.
  2. Reliability versus durability? Correct service over time versus data remaining intact.
  3. Why failover quota headroom? Recovery location must absorb failed-location demand.
  4. Why can retries reduce reliability? They amplify load and latency during failure.
  5. What makes an operation idempotent? Repeating the same identified request does not duplicate its effect.
  6. Why multi-AZ can still fail? Shared dependencies can remain single failure domains.
  7. Backup versus recovery proof? Successful restore and business validation within RTO/RPO.
  8. Why plan failback separately? It moves current authority/data/traffic again and creates new risk.

Lesson acceptance

Pass when each user objective maps to quantified capacity, dependency, failure, change, restore, and recovery evidence, including degraded behavior and owners. Fail if multi-AZ labels replace path analysis, quotas lack failover headroom, retries are unbounded, autoscaling ignores bottlenecks, rollback ignores data, backups are untested, or DR cannot prevent dual writers.

Official sources

Advertisement