AWS 389: Fault Injection Service experiments, stop conditions, and blast-radius control
Why this lesson matters
AWS Fault Injection Service performs controlled fault actions against selected resources. Safe chaos engineering begins with a falsifiable hypothesis, known steady state, exact targets, least-privilege role, bounded actions, independent stop conditions, observers and abort authority, recovery verification, and learning. Random breakage is not an experiment.
Experiment contract
| Field | Required content |
|---|---|
| Hypothesis | Under fault X, user SLI stays above threshold and recovers by time Y |
| Steady state | Baseline traffic, health, capacity, data, alarms and dependencies |
| Target | Account, Region, resource type, IDs/tags/filters and resolved inventory |
| Action | Fault, intensity, duration, sequence/parallel dependencies |
| Blast radius | Count/percent, environment, AZ, tenant/data and time bounds |
| Stop conditions | Independent CloudWatch alarms with tested actions/state semantics |
| Recovery | Automatic action reversal plus operator runbook and verification |
| Evidence | Template/version, experiment ID, selected targets, timeline and SLI |
FIS resolves targets at experiment start. Selection modes such as ALL, COUNT(n), or PERCENT(n) can select unexpectedly when the eligible set changes. Preview/independently inventory exact candidates, require experiment-specific safe tags plus account/VPC/AZ filters, fail on empty targets where appropriate, and prohibit production IDs through policy. Percent rounding and randomness must be accepted in the hypothesis.
IAM, actions, and sequencing
The experiment role is assumed by FIS and needs only APIs/resources/conditions required by selected actions and discovery. The human or pipeline starting the experiment needs separate authority and PassRole where applicable. Multi-account experiments add target-account roles and a larger governance boundary. Protect template creation/update and role policies from the experiment operator changing their own limits.
Actions can run in parallel or after named actions. Understand whether stopping an experiment reverses the service action, merely stops further injection, or leaves manual recovery. CPU/network/API faults delivered through SSM/EKS require healthy management channels and cleanup. Stopping an instance, changing capacity, or injecting errors can trigger Auto Scaling, alarms, failover, retries, and cost cascades.
Stop conditions are last-resort safety controls, not the primary result metric. Use alarms independent from the faulted component and test them before execution. Define ALARM, INSUFFICIENT_DATA, disabled actions, evaluation delay, dashboard/notification, and human abort. FIS stops when a condition triggers, but recovery may continue and already-observed side effects are not undone.
Safe execution procedure
- Obtain owner/change/security approval and conflict check; confirm no incident or risky deployment.
- Record account/Region/caller, template revision, role, exact candidate targets, quotas and current service health.
- Test stop alarms and recovery automation; brief observers and abort commander.
- Begin in nonproduction or one disposable target with representative traffic.
- Start one immutable template execution; watch FIS status, selected targets, stop alarms, system and user SLIs.
- Abort on any threshold, unexpected target, evidence loss, security/data risk, or operator uncertainty.
- Verify action reversal, capacity/data/queues/alarms, exact user recovery, and no residual SSM/network/resource change.
- Export report/logs/timeline, create actions, and clean experiment-only resources.
aws fis get-experiment-template --id TEMPLATE_ID --region ap-south-1
aws fis list-experiments --region ap-south-1
aws fis get-experiment --id EXPERIMENT_ID --region ap-south-1
aws cloudwatch describe-alarms --alarm-names STOP_ALARM --region ap-south-1
aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventSource,AttributeValue=fis.amazonaws.com --region ap-south-1
Workshop and failure matrix
Design a progressive experiment for a three-AZ stateless API: terminate one canary instance, impair 10 percent of API calls, then test one-AZ capacity only after prior gates. Specify hypothesis, sample/traffic, target query and resolved list, actions, role, stop alarms, SLI, expected Auto Scaling/load-balancer response, recovery deadlines, and abort authority.
Analyze 20 failures: wrong account, broad tag matches production, ALL used, percent selects too many small-set targets, empty target silently succeeds, experiment role broad, PassRole abuse, SSM agent unavailable, parallel actions amplify, stop alarm disabled, missing data, alarm depends on faulted path, alarm delay, operator cannot stop, action not reversible, ASG replaces targets and expands cost, data write duplicated, regional dependency impacted, experiment ends but service unhealthy, and residual fault remains.
Cost and acceptance
Price FIS actions/experiments under current pricing, target compute/load, replacement/surge, monitoring/logs/reports, SSM/Lambda, data transfer, recovery resources, and customer risk. This lesson performs no live fault.
Submit hypothesis/steady-state contract, exact targeting proof, IAM chain, action graph, stop-alarm test, observer/abort plan, progressive timeline, 20-case matrix, recovery verification, cost, and learning backlog. Pass requires bounded targets, independent tested stops, authorized human control, no assumption that experiment completion equals recovery, and measured user outcome.