AWS 391: Inject a controlled failure and prove automated recovery
Why this lesson matters
This lab proves a production-shaped but isolated service can detect one instance failure, remove the target, replace capacity, route healthy traffic, notify operators, and return to steady state inside a measured objective. An Auto Scaling event alone is not recovery proof.
Lab contract and safety
Use an owner-approved sandbox, never production. Set budget and teardown deadlines. The evidence-only track analyzes supplied FIS, Auto Scaling, target-health, alarm, synthetic and CloudTrail records. A live track uses two Availability Zones, an ALB, target group, launch template, Auto Scaling group desired capacity four, stateless release endpoint, CloudWatch/SNS, synthetic probe, and a least-privilege FIS template that terminates exactly one tagged canary instance.
| Contract item | Required value/evidence |
|---|---|
| Hypothesis | One instance loss causes no more than agreed failed requests and full recovery within 8 minutes |
| Steady state | Four healthy targets, stable SLI/capacity, exact release ID |
| Target | One disposable instance selected from explicit safe tags and resolved list |
| Stop | User error/latency, healthy-target floor, capacity and data alarms |
| Recovery | ASG replacement, bootstrap, target health, warm-up and user SLI |
| Abort | Named commander with console/CLI path and rollback authority |
| Evidence | UTC timeline linking experiment, instance, activities, target and requests |
Confirm target candidates independently immediately before start. The FIS role may terminate only tagged lab instances. The starter cannot broaden that role. Test stop alarms and SNS destination, verify no deployment/incident overlaps, and use a synthetic request with release/target/request IDs. Do not inject until baseline traffic is sufficient to measure harm.
Execution procedure
- Record caller/account/Region, architecture, IaC revision, launch-template/AMI version, FIS template revision, quotas and cost timer.
- Run a ten-minute baseline: target count, request good/total, latency, errors, ASG capacity, alarms and notification path.
- Resolve and approve the single target; confirm it has no local state and another AZ has capacity.
- Start the experiment once. Record experiment ID and exact
StartTime. - Observe target unhealthy/deregistration, connection behavior, ASG activity, launch, lifecycle/bootstrap, registration, health,
InService, warm-up, and desired capacity restoration. - Continue user probes through a bake window. Record detection, containment, replacement, target-ready, SLI-restored and alert-received times.
- Stop immediately for unexpected target, breached floor, stop alarm, data/security risk, missing evidence, or operator uncertainty.
- Verify no residual fault, queue/backlog, unhealthy target, pending instance, alarm, notification error, or version mismatch.
- Export redacted evidence and delete all owned resources in dependency order.
aws fis get-experiment --id EXPERIMENT_ID --region ap-south-1
aws autoscaling describe-scaling-activities --auto-scaling-group-name ASG --region ap-south-1
aws autoscaling describe-auto-scaling-instances --instance-ids INSTANCE_ID --region ap-south-1
aws elbv2 describe-target-health --target-group-arn TARGET_GROUP --region ap-south-1
aws cloudwatch describe-alarm-history --alarm-name ALARM --region ap-south-1
Evidence timeline and calculations
Build one UTC table with event source, timestamp, resource/release/request ID, observation, meaning, and confidence. Calculate detection time, alert latency, deregistration/containment, replacement launch, bootstrap, target-health, capacity restoration, user recovery, total recovery, failed request count, latency impact, and availability during window.
Separate correlation from causation. Use FIS/CloudTrail for injected action, ASG activities for replacement reason, target reason codes for routing, instance/application logs for readiness, and synthetic/API records for customer outcome. Reconcile clock skew and ingestion delay; do not infer event order from dashboard display alone.
Variations and failure analysis
Repeat only in evidence/simulation for: bootstrap failure, one-AZ capacity shortage, health endpoint false positive, notification delivery failure, stop alarm missing data, long-lived request exceeding drain, scaling policy concurrent action, and replacement runs wrong AMI. For each state whether automation contains, retries, expands damage, or needs human action.
Analyze 16 lab hazards: wrong account, target query broad, role broad, no baseline traffic, alarm disabled, stop too slow, ASG health type wrong, grace/warm-up mismatch, subnet IP shortage, user data secret failure, target registers early, desired capacity restored but SLI bad, stale dashboard, duplicate notification, teardown deletes shared resource, and orphaned ENI/log/snapshot.
Cost, cleanup, and acceptance
Price ALB, EC2/EBS surge, FIS, CloudWatch/Synthetics/logs, SNS, NAT/data transfer, and idle time. Stop experiment, delete ASG, wait for instances/targets/ENIs, then load balancer/listeners/target group, launch template, alarms/canary/topic, roles/policies, logs, networking, and images only if exclusively owned. Verify tag/resource and later billing inventories.
Submit contract, approvals, target proof, baseline, complete timeline, calculations, exact user outcome, eight variations, sixteen hazards, cost, and two-pass cleanup. Pass requires bounded injection, tested stops, measured SLI and recovery, capacity plus release identity, and no unowned resource.