AWS 385: Deployment health gates, Auto Scaling replacement, and load-balancer deregistration
Why this lesson matters
An instance can be running while its application is unready, registered while serving errors, or terminated while requests remain. Safe fleet replacement aligns bootstrap, lifecycle hooks, EC2/Auto Scaling/load-balancer/application health, warm-up, capacity, draining, alarms, checkpoints, and rollback.
Health states and timers
| Control | Purpose | Common confusion |
|---|---|---|
| EC2 status checks | Host/instance reachability | Not application correctness |
| ASG health-check grace | Delay replacement evaluation after service begins | Not instance refresh warm-up |
| Target health | Route only to targets passing configured check | Shallow endpoint may lie |
| Lifecycle launch hook | Hold Pending:Wait for preparation/evidence | Must heartbeat/complete before timeout |
| Instance warm-up | Delay refresh progress/scaling metric inclusion | Starts around InService, not boot start |
| Deregistration delay | Stop new requests and drain existing connections | Does not guarantee app process waits |
| Termination hook/policy | Run shutdown/export action | Hook alone can time out and termination proceed |
| Bake/checkpoint/alarm | Observe release before expansion/completion | Needs meaningful traffic and missing-data policy |
Launch flow: EC2 boot, user data/agent, launch hook, deep local test and release identity, lifecycle complete, target registration/health, InService, warm-up, then refresh progresses. Exact ordering varies with configuration; prove from events rather than assuming.
Termination flow: selection, target deregistration, connection draining, termination hook and graceful application stop, then instance termination. Long-lived WebSocket/gRPC/streaming or background jobs need explicit maximum duration, handoff/idempotency, process signal handling, and client retry. Deregistration delay only covers load-balancer connections it knows.
Capacity and replacement mathematics
For desired 60, minimum healthy 90 percent, maximum healthy 120 percent, calculate documented rounding and theoretical floor/ceiling before release. Then test per-AZ capacity, subnet IP/ENI, EC2 quota, target limit, instance weights, Spot/On-Demand policy, licenses, database connections, NAT, and budget. A global healthy percentage can still overload one AZ.
Default instance warm-up should represent time until metrics are trustworthy, not merely the health endpoint. Grace should be long enough for target checks but short enough to replace genuine failure. Aggressive liveness/target checks can create replacement storms; overly tolerant checks send traffic to broken code.
Instance maintenance policies can provide default minimum and maximum healthy percentages, while a specific refresh may override them. Document which source won. Scaling policies continue to respond during many replacements, so distinguish demand scale-out from deployment surge. Check whether scale-in protection, standby instances, warm-pool state, health replacement, capacity rebalance, or weighted capacity can block or alter refresh progress.
Health endpoint design should report process readiness and safe critical local dependencies without making every optional external dependency remove every target. Add synthetic user transactions and release-scoped error/latency/business alarms outside target health. Define alarm missing-data and evaluation windows relative to traffic and rollout interval.
Before termination, stop accepting new background work, advertise unready, drain the load balancer, complete or checkpoint bounded work, flush telemetry, and send lifecycle completion. Make every step idempotent because events and callbacks can repeat. On timeout, the default lifecycle result and instance lifecycle policy determine safety; test handler outage rather than assuming the shutdown script always runs.
Read-only inspection and game day
aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names ASG --region ap-south-1
aws autoscaling describe-instance-refreshes --auto-scaling-group-name ASG --region ap-south-1
aws autoscaling describe-scaling-activities --auto-scaling-group-name ASG --region ap-south-1
aws autoscaling describe-lifecycle-hooks --auto-scaling-group-name ASG --region ap-south-1
aws elbv2 describe-target-health --target-group-arn TARGET_GROUP --region ap-south-1
aws elbv2 describe-target-group-attributes --target-group-arn TARGET_GROUP --region ap-south-1
Design a three-AZ fleet for 9,000 requests/second with desired 60 and weighted mixed instances. Provide capacity math, per-AZ floor, launch/termination state machines, hook heartbeat/default results, checks/timers, refresh preferences/checkpoints/bake, scaling interaction, alarms, and rollback requirements.
Inject 18 failures: AMI boot, user-data partial success, hook handler down, heartbeat timeout, wrong health path, shallow 200, target flapping, warm-up too short, grace too long, subnet IP exhaustion, one-AZ capacity loss, Spot rebalance overlap, connection exceeds drain, process ignores termination, background duplicate, alarm missing data, rollback AMI unavailable, and database schema incompatible. Preserve instance, activity, lifecycle, target, alarm, release, and user evidence.
Cost and acceptance
Price surge EC2/EBS, load balancer capacity, cross-AZ/NAT traffic, logs/metrics/traces, hooks/Lambda/EventBridge, warm pools, retained AMIs/snapshots, and overprovisioned bake time. This lesson creates nothing.
Submit timing/state diagram, capacity math, health contract, hook runbooks, draining design, refresh/scale interaction, release dashboard, eighteen failures, cost, and rollback/forward-fix criteria. Pass requires measured readiness, per-AZ headroom, graceful long-lived work, version-scoped alarms, and proof of restored user health.