Lesson 391 · AWS Learning Path

AWS 391: Inject a controlled failure and prove automated recovery

· Published · 4 min read

Labelled process diagram for AWS 391: Approved hypothesis and healthy baseline to Bounded failure with stop conditions to Detection, containment and automated recovery to User validation, measured objectives...

Why this lesson matters

This lab proves a production-shaped but isolated service can detect one instance failure, remove the target, replace capacity, route healthy traffic, notify operators, and return to steady state inside a measured objective. An Auto Scaling event alone is not recovery proof.

Lab contract and safety

Use an owner-approved sandbox, never production. Set budget and teardown deadlines. The evidence-only track analyzes supplied FIS, Auto Scaling, target-health, alarm, synthetic and CloudTrail records. A live track uses two Availability Zones, an ALB, target group, launch template, Auto Scaling group desired capacity four, stateless release endpoint, CloudWatch/SNS, synthetic probe, and a least-privilege FIS template that terminates exactly one tagged canary instance.

Contract itemRequired value/evidence
HypothesisOne instance loss causes no more than agreed failed requests and full recovery within 8 minutes
Steady stateFour healthy targets, stable SLI/capacity, exact release ID
TargetOne disposable instance selected from explicit safe tags and resolved list
StopUser error/latency, healthy-target floor, capacity and data alarms
RecoveryASG replacement, bootstrap, target health, warm-up and user SLI
AbortNamed commander with console/CLI path and rollback authority
EvidenceUTC timeline linking experiment, instance, activities, target and requests

Confirm target candidates independently immediately before start. The FIS role may terminate only tagged lab instances. The starter cannot broaden that role. Test stop alarms and SNS destination, verify no deployment/incident overlaps, and use a synthetic request with release/target/request IDs. Do not inject until baseline traffic is sufficient to measure harm.

Execution procedure

  1. Record caller/account/Region, architecture, IaC revision, launch-template/AMI version, FIS template revision, quotas and cost timer.
  2. Run a ten-minute baseline: target count, request good/total, latency, errors, ASG capacity, alarms and notification path.
  3. Resolve and approve the single target; confirm it has no local state and another AZ has capacity.
  4. Start the experiment once. Record experiment ID and exact StartTime.
  5. Observe target unhealthy/deregistration, connection behavior, ASG activity, launch, lifecycle/bootstrap, registration, health, InService, warm-up, and desired capacity restoration.
  6. Continue user probes through a bake window. Record detection, containment, replacement, target-ready, SLI-restored and alert-received times.
  7. Stop immediately for unexpected target, breached floor, stop alarm, data/security risk, missing evidence, or operator uncertainty.
  8. Verify no residual fault, queue/backlog, unhealthy target, pending instance, alarm, notification error, or version mismatch.
  9. Export redacted evidence and delete all owned resources in dependency order.
aws fis get-experiment --id EXPERIMENT_ID --region ap-south-1
aws autoscaling describe-scaling-activities --auto-scaling-group-name ASG --region ap-south-1
aws autoscaling describe-auto-scaling-instances --instance-ids INSTANCE_ID --region ap-south-1
aws elbv2 describe-target-health --target-group-arn TARGET_GROUP --region ap-south-1
aws cloudwatch describe-alarm-history --alarm-name ALARM --region ap-south-1

Evidence timeline and calculations

Build one UTC table with event source, timestamp, resource/release/request ID, observation, meaning, and confidence. Calculate detection time, alert latency, deregistration/containment, replacement launch, bootstrap, target-health, capacity restoration, user recovery, total recovery, failed request count, latency impact, and availability during window.

Separate correlation from causation. Use FIS/CloudTrail for injected action, ASG activities for replacement reason, target reason codes for routing, instance/application logs for readiness, and synthetic/API records for customer outcome. Reconcile clock skew and ingestion delay; do not infer event order from dashboard display alone.

Variations and failure analysis

Repeat only in evidence/simulation for: bootstrap failure, one-AZ capacity shortage, health endpoint false positive, notification delivery failure, stop alarm missing data, long-lived request exceeding drain, scaling policy concurrent action, and replacement runs wrong AMI. For each state whether automation contains, retries, expands damage, or needs human action.

Analyze 16 lab hazards: wrong account, target query broad, role broad, no baseline traffic, alarm disabled, stop too slow, ASG health type wrong, grace/warm-up mismatch, subnet IP shortage, user data secret failure, target registers early, desired capacity restored but SLI bad, stale dashboard, duplicate notification, teardown deletes shared resource, and orphaned ENI/log/snapshot.

Cost, cleanup, and acceptance

Price ALB, EC2/EBS surge, FIS, CloudWatch/Synthetics/logs, SNS, NAT/data transfer, and idle time. Stop experiment, delete ASG, wait for instances/targets/ENIs, then load balancer/listeners/target group, launch template, alarms/canary/topic, roles/policies, logs, networking, and images only if exclusively owned. Verify tag/resource and later billing inventories.

Submit contract, approvals, target proof, baseline, complete timeline, calculations, exact user outcome, eight variations, sixteen hazards, cost, and two-pass cleanup. Pass requires bounded injection, tested stops, measured SLI and recovery, capacity plus release identity, and no unowned resource.

Official sources

Advertisement