Lesson 389 · AWS Learning Path

AWS 389: Fault Injection Service experiments, stop conditions, and blast-radius control

· Published · 4 min read

Labelled process diagram for AWS 389: Versioned intent to Automated validation to Controlled AWS change to Observed result and retained evidence, with decision, proof and rejection evidence.

Why this lesson matters

AWS Fault Injection Service performs controlled fault actions against selected resources. Safe chaos engineering begins with a falsifiable hypothesis, known steady state, exact targets, least-privilege role, bounded actions, independent stop conditions, observers and abort authority, recovery verification, and learning. Random breakage is not an experiment.

Experiment contract

FieldRequired content
HypothesisUnder fault X, user SLI stays above threshold and recovers by time Y
Steady stateBaseline traffic, health, capacity, data, alarms and dependencies
TargetAccount, Region, resource type, IDs/tags/filters and resolved inventory
ActionFault, intensity, duration, sequence/parallel dependencies
Blast radiusCount/percent, environment, AZ, tenant/data and time bounds
Stop conditionsIndependent CloudWatch alarms with tested actions/state semantics
RecoveryAutomatic action reversal plus operator runbook and verification
EvidenceTemplate/version, experiment ID, selected targets, timeline and SLI

FIS resolves targets at experiment start. Selection modes such as ALL, COUNT(n), or PERCENT(n) can select unexpectedly when the eligible set changes. Preview/independently inventory exact candidates, require experiment-specific safe tags plus account/VPC/AZ filters, fail on empty targets where appropriate, and prohibit production IDs through policy. Percent rounding and randomness must be accepted in the hypothesis.

IAM, actions, and sequencing

The experiment role is assumed by FIS and needs only APIs/resources/conditions required by selected actions and discovery. The human or pipeline starting the experiment needs separate authority and PassRole where applicable. Multi-account experiments add target-account roles and a larger governance boundary. Protect template creation/update and role policies from the experiment operator changing their own limits.

Actions can run in parallel or after named actions. Understand whether stopping an experiment reverses the service action, merely stops further injection, or leaves manual recovery. CPU/network/API faults delivered through SSM/EKS require healthy management channels and cleanup. Stopping an instance, changing capacity, or injecting errors can trigger Auto Scaling, alarms, failover, retries, and cost cascades.

Stop conditions are last-resort safety controls, not the primary result metric. Use alarms independent from the faulted component and test them before execution. Define ALARM, INSUFFICIENT_DATA, disabled actions, evaluation delay, dashboard/notification, and human abort. FIS stops when a condition triggers, but recovery may continue and already-observed side effects are not undone.

Safe execution procedure

  1. Obtain owner/change/security approval and conflict check; confirm no incident or risky deployment.
  2. Record account/Region/caller, template revision, role, exact candidate targets, quotas and current service health.
  3. Test stop alarms and recovery automation; brief observers and abort commander.
  4. Begin in nonproduction or one disposable target with representative traffic.
  5. Start one immutable template execution; watch FIS status, selected targets, stop alarms, system and user SLIs.
  6. Abort on any threshold, unexpected target, evidence loss, security/data risk, or operator uncertainty.
  7. Verify action reversal, capacity/data/queues/alarms, exact user recovery, and no residual SSM/network/resource change.
  8. Export report/logs/timeline, create actions, and clean experiment-only resources.
aws fis get-experiment-template --id TEMPLATE_ID --region ap-south-1
aws fis list-experiments --region ap-south-1
aws fis get-experiment --id EXPERIMENT_ID --region ap-south-1
aws cloudwatch describe-alarms --alarm-names STOP_ALARM --region ap-south-1
aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventSource,AttributeValue=fis.amazonaws.com --region ap-south-1

Workshop and failure matrix

Design a progressive experiment for a three-AZ stateless API: terminate one canary instance, impair 10 percent of API calls, then test one-AZ capacity only after prior gates. Specify hypothesis, sample/traffic, target query and resolved list, actions, role, stop alarms, SLI, expected Auto Scaling/load-balancer response, recovery deadlines, and abort authority.

Analyze 20 failures: wrong account, broad tag matches production, ALL used, percent selects too many small-set targets, empty target silently succeeds, experiment role broad, PassRole abuse, SSM agent unavailable, parallel actions amplify, stop alarm disabled, missing data, alarm depends on faulted path, alarm delay, operator cannot stop, action not reversible, ASG replaces targets and expands cost, data write duplicated, regional dependency impacted, experiment ends but service unhealthy, and residual fault remains.

Cost and acceptance

Price FIS actions/experiments under current pricing, target compute/load, replacement/surge, monitoring/logs/reports, SSM/Lambda, data transfer, recovery resources, and customer risk. This lesson performs no live fault.

Submit hypothesis/steady-state contract, exact targeting proof, IAM chain, action graph, stop-alarm test, observer/abort plan, progressive timeline, 20-case matrix, recovery verification, cost, and learning backlog. Pass requires bounded targets, independent tested stops, authorized human control, no assumption that experiment completion equals recovery, and measured user outcome.

Official sources

Advertisement