Lesson 295 · AWS Learning Path

AWS 295: Failure testing, runbooks and recovery validation

· Published · 16 min read

Labelled process diagram for AWS 295: Known steady state to Bounded approved fault to Observed detection and recovery to Measured objective, cleanup, and improvement, with decision, proof and rejection evidence.

Why this lesson matters

An architecture diagram expresses an intention. A backup-success message proves that one job completed. A recovery runbook records an expected procedure. None proves that a user can complete a critical transaction after a real failure.

Failure testing turns resilience assumptions into bounded experiments. It observes whether the system detects a fault, limits impact, recovers within objectives, preserves data, alerts the right people, and leaves trustworthy evidence. The result can disprove the design, and that is useful: discovering a gap during an approved game day is safer than discovering it during an uncontrolled incident.

This lesson teaches a progressive method from tabletop to production-like fault injection. AWS Fault Injection Service (AWS FIS) is one mechanism, not the purpose. Some risks are better tested with load tools, restoration exercises, application fault switches, dependency simulators, security exercises, or a tabletop.

What you will be able to do

By the end, you can:

  • distinguish functional, performance, recovery, resilience, chaos, and game-day testing;
  • write a measurable steady state and falsifiable hypothesis;
  • choose the safest environment and smallest useful blast radius;
  • explain FIS templates, actions, targets, selection modes, roles, stop conditions, reports, logs, and safety levers;
  • prevent broad tag selection and multi-account/AZ-name mistakes;
  • design stop conditions without treating them as instant rollback;
  • measure detection, response, RPO, RTO, recovery, and user impact from one timeline;
  • validate technical and human runbooks under controlled pressure;
  • convert findings into owned engineering work and regression tests; and
  • produce three reviewable experiment plans without changing AWS resources.

Before you start

  • This is a no-create, no-fault lesson. Do not start an FIS experiment or alter a resource.
  • Fault injection requires explicit workload-owner, operations, security, and change approval. Production experiments require stronger approval, customer-impact controls, and demonstrated lower-environment evidence.
  • Do not target resources because their names look familiar. Resolve exact account, Region, VPC, application, environment, owner, and resource IDs before approval.
  • Do not share account IDs, private ARNs, alarm names, resource tags, logs containing customer data, or incident contacts.
  • A stopping mechanism can take time and cannot always undo an action already applied. Preserve independent recovery controls.
  • Verify current FIS action behavior, supported resources, quotas, and pricing before execution.

1. Choose the kind of evidence you need

Test typeQuestionExample
FunctionalDoes the intended feature work?Submit and retrieve an order
Performance/loadDoes it meet capacity and latency under demand?Sustain peak requests with headroom
Backup/restoreCan protected data be reconstructed?Restore to an isolated account and reconcile
Recovery procedureCan the service return after a declared scenario?Regional recovery and business acceptance
ResilienceDoes it continue or degrade acceptably during a known fault?Lose one replica or dependency path
Chaos experimentDoes a hypothesis hold under a controlled adverse condition?Introduce realistic latency and observe behavior
Game dayDo technology, people, process, access, communications, and decisions work together?Simulated AZ impairment with command center
TabletopIs the plan logically complete when execution is unsafe or unavailable?Discuss total identity-provider failure

Use the least risky method that can produce the required evidence. A tabletop can expose missing authority and contacts but cannot prove automation or timing. A lower environment can validate mechanisms but may not reproduce production scale, data, topology, and organizational pressure. Increase realism only after earlier stages pass.

2. Write the experiment charter

Every experiment starts with a one-page charter:

  1. Business service and critical user journey.
  2. Risk or past incident being tested.
  3. Steady-state metrics and normal range.
  4. Falsifiable hypothesis.
  5. Exact fault and why it is realistic.
  6. Environment, account, Region, AZ IDs, target inventory, and selection mode.
  7. Expected blast radius and maximum tolerated impact.
  8. Preconditions and readiness gates.
  9. Stop conditions, safety lever, manual abort, and recovery procedure.
  10. Timeline, duration, owners, approvers, observers, and incident commander.
  11. Evidence sources and synchronized clocks.
  12. Success, partial success, and failure criteria.
  13. Cleanup, verification, and follow-up process.

A useful hypothesis is falsifiable:

If one of four stateless application instances becomes unavailable for 10 minutes at 40 percent normal peak traffic, then successful checkout remains at least 99.5 percent, p95 latency stays below 800 ms, no order is duplicated or lost, the service alarm enters ALARM within 2 minutes, and capacity returns within 8 minutes without operator intervention.

“The platform is highly available” cannot be disproved because it defines no event or threshold.

3. Define steady state from the user backward

Steady state is the measurable business and technical behavior expected before, during, and after the fault. Use several layers:

  • business: successful checkout, claims processed, files accepted, payment completion;
  • edge/API: request success, latency percentiles, throttles, TLS and DNS success;
  • application: queue age, worker throughput, retries, circuit-breaker state, thread/connection saturation;
  • data: transaction commits, replica lag, conflict/duplicate counts, RPO markers;
  • infrastructure: healthy targets, capacity, CPU/memory/disk/network, AZ distribution;
  • operations: alarm, page delivery, acknowledgement, escalation, runbook start; and
  • security: audit continuity, denied unauthorized paths, secret/key access, findings.

Define the source, query, aggregation period, dimensions, threshold, missing-data treatment, normal baseline window, and dashboard owner. An average can hide a failed tenant or AZ. A technical 200 response can hide a wrong business result.

Capture a pre-experiment baseline under comparable demand. Confirm telemetry is current and time synchronized. If you cannot observe the stated steady state, do not inject the fault.

4. Build a progression of confidence

Use a maturity ladder:

  1. Review architecture, failure modes, ownership, and runbooks.
  2. Tabletop the decision and communication flow.
  3. Test one component in a sandbox with synthetic data.
  4. Test the integrated service in a production-like environment under representative load.
  5. Test a narrow production scope during an approved low-risk window.
  6. Automate proven experiments as controlled regression where appropriate.

Each stage must have promotion criteria. Do not jump to production because a sandbox passed. Compare differences in instance count, scaling, quotas, network, IAM, data volume, integrations, observability, on-call staffing, and traffic.

Start with one failure and one recovery mechanism. Compound failures are valuable only after individual behaviors are understood. Otherwise diagnosis becomes ambiguous and blast radius grows without proportional learning.

5. Understand an AWS FIS experiment

An FIS experiment template is reusable configuration. An experiment is one execution and resolves its targets at the beginning.

Actions and sequencing

Actions inject conditions or invoke supported automation. Each action has an action ID, parameters, targets where required, and optional dependencies on other actions. Independent actions can run in parallel; dependencies create sequence. Duration syntax and action-specific recovery behavior must be verified from the current action documentation.

Map every action to:

  • intended physical failure;
  • affected layer and resources;
  • beginning and ending behavior;
  • whether stopping the experiment reverses it;
  • service/API permissions;
  • recovery verification; and
  • any persistent side effect.

Do not assume stopping FIS means the workload instantly returns. A stopped instance might need restart, failed connections need retry, autoscaling needs time, caches need refill, and data reconciliation can remain.

Targets and selection modes

Targets identify eligible resources by explicit ARNs/IDs, tags, and supported filters. A selection mode can choose all, a count, or a percentage depending on the target definition. FIS resolves targets at experiment start and uses that selected set for the execution.

Broad tags are dangerous. Environment=production can span unrelated services. Require multiple stable dimensions such as application, environment, owner, experiment eligibility, VPC, and AZ ID, then perform a read-only preflight that prints the exact resources. For a first experiment, explicit IDs or a dedicated opt-in tag are safer than discovery over a broad estate.

Set behavior for empty target resolution deliberately. Skipping an empty target can be useful in orchestration, but it can also produce a misleading “successful” experiment that injected no intended fault. Evidence must list the resolved targets and action results.

Experiment role and human authorization

FIS assumes an IAM experiment role to perform actions. Grant only the APIs and resources required by the approved template. Constrain trust with the source account and experiment ARN pattern where applicable. Separate template author, approver, experiment starter, safety operator, and evidence reviewer when consequence warrants it.

FIS permissions are technical capability, not change authorization. Require a change record, named start authority, maintenance or customer communication, and an independent person who can stop the experiment.

Stop conditions

FIS stop conditions use CloudWatch alarms. When a referenced alarm triggers, FIS stops the running experiment. Design alarms from maximum acceptable impact, not merely the hypothesis success threshold.

For example:

  • hypothesis: p95 latency remains below 800 ms;
  • experiment failure: p95 exceeds 800 ms for two evaluation periods;
  • stop condition: checkout success falls below 97 percent for one period or data-integrity alarm triggers.

This preserves learning time while protecting customers. Test alarm state transitions and missing-data behavior before the experiment. The alarm must be in the correct Region/account context and must not depend on the impaired path it monitors. Combine automatic stops with a manual abort and direct service recovery.

Safety lever

Each account and Region has an FIS safety lever. Engaging it stops running experiments in that scope and prevents new ones from starting. A stopped or cancelled experiment cannot resume; a new execution is required later.

The safety lever is broader than one template, so name who may engage it, how multi-account experiments are covered, and what evidence confirms every targeted account/Region. Test read access and the communication process. Do not treat it as an application rollback button.

Logs, reports, and CloudTrail

Configure experiment logs and report outputs where required. Record template/version, experiment ID, resolved targets, actions, timestamps, action results, stop condition changes, safety-lever state, CloudTrail events, application telemetry, incident communication, and cleanup. Protect sensitive output and set retention/cost controls.

6. Multi-account and Availability Zone safety

A multi-account FIS experiment runs from an orchestrator account and uses target-account configurations with account IDs and roles. Target accounts can receive AWS Health awareness, but that does not replace direct stakeholder communication. Review trust, least privilege, consistent opt-in tags, logging, quotas, stop alarms, and safety levers in every account.

Availability Zone names such as ap-south-1a can map to different physical zones in different accounts. Use AZ IDs when one physical zone must be targeted consistently across accounts. Otherwise a “single-AZ” experiment can affect different zones and fail to test the intended dependency.

Start multi-account work only after the same action and safeguards pass in one account. Verify which account owns shared load balancers, databases, networks, observability, and DNS; the application boundary rarely follows account boundaries exactly.

7. Design safety as independent layers

Use all applicable controls:

  • approved environment and maintenance window;
  • experiment opt-in tag with owner and expiration;
  • exact preflight target list and expected count;
  • smallest selection mode and shortest useful duration;
  • least-privilege experiment role;
  • tested CloudWatch stop alarms;
  • account/Region safety lever and authorized operator;
  • manual abort and direct resource recovery procedure;
  • maximum customer-impact and time budget;
  • freeze on unrelated changes;
  • current backups and restore proof;
  • healthy standby capacity and service quotas;
  • incident bridge, on-call staff, AWS Support plan, and vendor contacts; and
  • prohibition on progression when baseline or telemetry is unhealthy.

Define abort behavior for unexpected target count, alarm malfunction, telemetry loss, operator loss, data-integrity signal, unrelated incident, customer threshold, action overrun, or inability to recover. “We can stop if necessary” is not a control.

8. Measure one authoritative timeline

Use synchronized UTC timestamps:

T0 fault action begins
T1 fault takes effect
T2 monitoring first detects it
T3 alarm enters ALARM
T4 page delivered
T5 human acknowledges
T6 recovery automation/operator action begins
T7 technical service restored
T8 data reconciled
T9 business journey accepted
T10 fault and temporary resources cleaned up

Derive:

  • detection time: T2 - T1;
  • alarm delay: T3 - T1;
  • acknowledgement time: T5 - T4;
  • recovery action delay: T6 - T1;
  • technical recovery: T7 - T1;
  • effective RTO: T9 - T1; and
  • experiment closure: T10 - T0.

RPO needs a durable business marker. Write a known sequence or transaction before and during the event, then identify the newest accepted durable result after recovery. Do not infer RPO from server uptime, replica “healthy” status, or backup job time alone.

Record user impact area, duration, affected count, lost/duplicate/corrupt transactions, degraded features, and SLO error-budget consumption. An experiment can confirm the hypothesis while revealing slow paging or poor operator access; preserve both results.

9. Validate the runbook, not only the platform

A runbook step must state trigger, prerequisite, exact authorized action, expected result, evidence, timeout, owner, escalation, and recovery/rollback. During the experiment, observers record whether:

  • the correct person found the correct version;
  • access worked without borrowing credentials;
  • commands were safe and parameters validated;
  • expected output matched reality;
  • parallel actions and decision authority were clear;
  • communication reached business and technical audiences;
  • stop and recovery steps worked; and
  • cleanup and handoff were completed.

Do not coach silently around a missing step and then mark the runbook passed. Record the intervention as a failure and improve the artifact.

Automation should be versioned, reviewed, idempotent where possible, bounded by timeouts, and produce evidence. A Systems Manager Automation runbook can orchestrate approved steps, but its document permissions, targets, parameters, output, rollback, and regional availability need the same review as FIS.

10. Three experiment designs

Experiment A: one stateless instance is lost

Hypothesis: an ALB-backed Auto Scaling service maintains the stated checkout steady state when one explicitly selected application instance stops; unhealthy detection, connection draining/retry, and replacement restore desired capacity within thresholds.

Prove target count, remaining AZ/capacity headroom, ALB health, Auto Scaling process, session externalization, idempotent client retry, stop alarm, restart/replacement behavior, and no data loss. A passing EC2 recovery with failed checkout is a failed experiment.

Experiment B: dependency latency increases

Hypothesis: when a noncritical recommendation dependency adds 1.5 seconds latency, the order API enforces a 500 ms timeout, opens its circuit breaker, serves orders without recommendations, and does not exhaust threads or connections.

Prefer an approved application fault switch, proxy/dependency simulator, or currently supported FIS network action that precisely represents the path. Do not impair an entire subnet merely because the tool can. Observe timeout budget, retries, retry amplification, pool saturation, fallback correctness, recovery hysteresis, and user output.

Experiment C: one AZ is impaired

Hypothesis: when application resources in one physical AZ become unavailable, new traffic uses remaining AZs, database availability follows the documented mechanism, queue processing stays within age threshold, and the service meets RTO without exceeding remaining capacity.

Inventory every zonal dependency: subnets, NAT, load balancers, endpoints, compute, database, cache, storage, queue access, DNS resolver endpoints, and observability. Use AZ IDs across accounts. Begin with one layer or a managed scenario and expand only after proving safety. An AZ experiment that touches only EC2 does not prove AZ independence.

11. Guided game-day workshop

Use a supplied two-AZ order service with ALB, four EC2 application instances, Auto Scaling, RDS Multi-AZ, SQS workers, ElastiCache, one NAT gateway per AZ, interface endpoints, and external payment/recommendation APIs. The business objective is 99.5 percent successful checkout during a single-instance event, degraded checkout without recommendations during dependency latency, and restoration within 20 minutes during an AZ scenario with no lost accepted order.

Produce these artifacts:

  1. System, dependency, identity, data, network, failure, and observability diagrams.
  2. User-focused steady-state definition and baseline query catalog.
  3. Risk-ranked failure-mode inventory from design and past incidents.
  4. Experiment charter for each of A, B, and C.
  5. Exact target inventory and a simulated FIS resolution/preflight report.
  6. Template design containing actions, ordering, targets, role, alarms, logs, reports, and options without creating it.
  7. IAM trust/permission boundaries and separation of duties.
  8. Three independent safety layers plus direct recovery for every action.
  9. Multi-account and AZ-ID analysis.
  10. Minute-level execution, observation, communication, and decision timeline.
  11. Runbook excerpts that a different operator can follow.
  12. Evidence matrix for business, application, data, infrastructure, security, and operations.
  13. RPO/RTO, detection, acknowledgement, technical recovery, and business recovery calculations.
  14. Failure injections from supplied results: missing alarm data, six resolved targets instead of one, expired opt-in tag, replacement blocked by quota, retry storm, stale dashboard, operator without permission, and cleanup failure.
  15. Final findings, risk acceptance, engineering backlog, owner/due date, and regression schedule.

No experiment passes merely because FIS reaches completed. The hypothesis, customer impact, data correctness, recovery, cleanup, and evidence all need separate results.

12. Read-only inspection

Only in an authorized account:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text

aws fis list-experiment-templates \
  --query 'experimentTemplates[].{Id:id,Description:description,Created:creationTime}'

aws fis list-experiments \
  --query 'experiments[].{Id:id,Template:experimentTemplateId,State:state.status,Created:creationTime}'

aws fis get-safety-lever --id default \
  --query '{Status:state.status,Reason:state.reason}'

aws ssm list-documents --filters Key=Owner,Values=Self \
  --query 'DocumentIdentifiers[].{Name:Name,Type:DocumentType,Version:DocumentVersion}'

Listing proves only control-plane inventory. For one owned template, review the complete template, exact target filters, action dependencies, selection modes, role, stop alarms, logs, report configuration, experiment options, and previous execution results. Do not start or stop anything.

13. Close the learning loop

Classify every outcome:

  • hypothesis sustained with complete evidence;
  • hypothesis rejected;
  • inconclusive because observability or injection was invalid;
  • experiment stopped by guardrail;
  • aborted manually; or
  • setup failed before a valid fault.

Do not relabel stopped or invalid experiments as resilience success. Preserve evidence and record surprises, user impact, data impact, operator interventions, safeguard performance, runbook gaps, cost, and cleanup.

For each finding create an owner, severity, corrective action, due date, acceptance evidence, and retest. Update architecture diagrams, alarms, runbooks, automation, quotas, training, and risk register. A game day without tracked remediation is theater.

Turn stable, low-blast-radius experiments into scheduled or pipeline regression only after target safeguards, data isolation, concurrency control, change freezes, stop alarms, and ownership are reliable. Review templates whenever architecture or tags change.

Cost, quotas, and cleanup

Include FIS action charges under current pricing, affected-resource runtime, load generation, observability ingestion/query/retention, snapshots/backups, replacement capacity, network/data transfer, support, staff time, incident communication, temporary environments, and remediation. Failure testing can trigger scaling and cross-AZ traffic that exceeds the template's direct charge.

Check FIS quotas for actions, targets, templates, concurrent experiments, duration, stop conditions, and multi-account scope, plus quotas of every affected service. Quota exhaustion may be the very failure discovered; ensure the experiment itself cannot consume emergency reserve unexpectedly.

Cleanup verifies action end state, instance/task/pod health, Auto Scaling desired capacity, traffic, alarms, temporary fault switches, SSM commands, security rules, test data, snapshots, logs, tags, open incidents, and steady-state recovery. Keep evidence according to policy, not forever by default.

Diagnose a weak experiment

SymptomHidden problemCorrection
FIS says completed but no metric changedEmpty/wrong target or ineffective actionProve resolved targets and fault effect
More resources fail than expectedBroad tag/filter or ALL selectionUse opt-in tags, exact preflight, count, VPC and AZ constraints
Stop alarm never firesWrong dimensions/Region, missing-data policy, delayed metric, or impaired telemetryTest alarm before fault and provide independent abort
Stopping experiment leaves outageAction is not instantly reversible or system needs recoveryDocument action semantics and direct restoration
CPU and hosts recover but checkout failsSteady state was infrastructure-onlyAdd user journey, dependency, data, and business evidence
Single-AZ test crosses physical zonesAZ names differ across accountsUse consistent AZ IDs and map all accounts
RTO reported from instance statusBusiness/data validation excludedEnd clock at accepted user service
Same defect returns next quarterNo owner, retest, or template regressionTrack remediation to evidence-backed closure

Knowledge check

  1. What makes a hypothesis falsifiable?

A defined fault, scope, duration, metrics, thresholds, and recovery behavior can prove it wrong.

  1. When does FIS resolve targets?

At experiment start; the selected set is then used for the execution.

  1. Is a CloudWatch stop condition a rollback?

No. It stops FIS actions as supported, but workload recovery and persistent effects still require handling.

  1. What does the safety lever do?

In one account and Region it stops running experiments and prevents new ones from starting.

  1. Why use AZ IDs in multi-account experiments?

AZ names can refer to different physical zones in different accounts.

  1. Why can an empty-target skip be dangerous?

The experiment may appear successful without injecting the intended fault.

  1. When does effective RTO end?

When the required business service and data are accepted, not when infrastructure returns.

  1. What completes a game day?

Evidence, cleanup, owned remediation, updated artifacts, and a scheduled retest.

Lesson acceptance

You may continue when your submission contains:

  • a risk-based choice of test type and realism level;
  • measurable steady state and falsifiable hypotheses;
  • complete FIS component and target-resolution reasoning;
  • least privilege, stop alarms, safety lever, manual abort, and direct recovery;
  • correct multi-account and AZ-ID controls;
  • one authoritative detection-to-business-recovery timeline;
  • runbook, access, communication, and decision validation;
  • three complete instance, latency, and AZ experiment designs;
  • business, data, security, infrastructure, and operations evidence;
  • cost, quota, cleanup, and evidence-retention ownership; and
  • a remediation backlog with owners, due dates, acceptance evidence, and retests.

Official sources

Advertisement