Lesson 385 · AWS Learning Path

AWS 385: Deployment health gates, Auto Scaling replacement, and load-balancer deregistration

· Published · 4 min read

Labelled process diagram for AWS 385: Versioned intent to Automated validation to Controlled AWS change to Observed result and retained evidence, with decision, proof and rejection evidence.

Why this lesson matters

An instance can be running while its application is unready, registered while serving errors, or terminated while requests remain. Safe fleet replacement aligns bootstrap, lifecycle hooks, EC2/Auto Scaling/load-balancer/application health, warm-up, capacity, draining, alarms, checkpoints, and rollback.

Health states and timers

ControlPurposeCommon confusion
EC2 status checksHost/instance reachabilityNot application correctness
ASG health-check graceDelay replacement evaluation after service beginsNot instance refresh warm-up
Target healthRoute only to targets passing configured checkShallow endpoint may lie
Lifecycle launch hookHold Pending:Wait for preparation/evidenceMust heartbeat/complete before timeout
Instance warm-upDelay refresh progress/scaling metric inclusionStarts around InService, not boot start
Deregistration delayStop new requests and drain existing connectionsDoes not guarantee app process waits
Termination hook/policyRun shutdown/export actionHook alone can time out and termination proceed
Bake/checkpoint/alarmObserve release before expansion/completionNeeds meaningful traffic and missing-data policy

Launch flow: EC2 boot, user data/agent, launch hook, deep local test and release identity, lifecycle complete, target registration/health, InService, warm-up, then refresh progresses. Exact ordering varies with configuration; prove from events rather than assuming.

Termination flow: selection, target deregistration, connection draining, termination hook and graceful application stop, then instance termination. Long-lived WebSocket/gRPC/streaming or background jobs need explicit maximum duration, handoff/idempotency, process signal handling, and client retry. Deregistration delay only covers load-balancer connections it knows.

Capacity and replacement mathematics

For desired 60, minimum healthy 90 percent, maximum healthy 120 percent, calculate documented rounding and theoretical floor/ceiling before release. Then test per-AZ capacity, subnet IP/ENI, EC2 quota, target limit, instance weights, Spot/On-Demand policy, licenses, database connections, NAT, and budget. A global healthy percentage can still overload one AZ.

Default instance warm-up should represent time until metrics are trustworthy, not merely the health endpoint. Grace should be long enough for target checks but short enough to replace genuine failure. Aggressive liveness/target checks can create replacement storms; overly tolerant checks send traffic to broken code.

Instance maintenance policies can provide default minimum and maximum healthy percentages, while a specific refresh may override them. Document which source won. Scaling policies continue to respond during many replacements, so distinguish demand scale-out from deployment surge. Check whether scale-in protection, standby instances, warm-pool state, health replacement, capacity rebalance, or weighted capacity can block or alter refresh progress.

Health endpoint design should report process readiness and safe critical local dependencies without making every optional external dependency remove every target. Add synthetic user transactions and release-scoped error/latency/business alarms outside target health. Define alarm missing-data and evaluation windows relative to traffic and rollout interval.

Before termination, stop accepting new background work, advertise unready, drain the load balancer, complete or checkpoint bounded work, flush telemetry, and send lifecycle completion. Make every step idempotent because events and callbacks can repeat. On timeout, the default lifecycle result and instance lifecycle policy determine safety; test handler outage rather than assuming the shutdown script always runs.

Read-only inspection and game day

aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names ASG --region ap-south-1
aws autoscaling describe-instance-refreshes --auto-scaling-group-name ASG --region ap-south-1
aws autoscaling describe-scaling-activities --auto-scaling-group-name ASG --region ap-south-1
aws autoscaling describe-lifecycle-hooks --auto-scaling-group-name ASG --region ap-south-1
aws elbv2 describe-target-health --target-group-arn TARGET_GROUP --region ap-south-1
aws elbv2 describe-target-group-attributes --target-group-arn TARGET_GROUP --region ap-south-1

Design a three-AZ fleet for 9,000 requests/second with desired 60 and weighted mixed instances. Provide capacity math, per-AZ floor, launch/termination state machines, hook heartbeat/default results, checks/timers, refresh preferences/checkpoints/bake, scaling interaction, alarms, and rollback requirements.

Inject 18 failures: AMI boot, user-data partial success, hook handler down, heartbeat timeout, wrong health path, shallow 200, target flapping, warm-up too short, grace too long, subnet IP exhaustion, one-AZ capacity loss, Spot rebalance overlap, connection exceeds drain, process ignores termination, background duplicate, alarm missing data, rollback AMI unavailable, and database schema incompatible. Preserve instance, activity, lifecycle, target, alarm, release, and user evidence.

Cost and acceptance

Price surge EC2/EBS, load balancer capacity, cross-AZ/NAT traffic, logs/metrics/traces, hooks/Lambda/EventBridge, warm pools, retained AMIs/snapshots, and overprovisioned bake time. This lesson creates nothing.

Submit timing/state diagram, capacity math, health contract, hook runbooks, draining design, refresh/scale interaction, release dashboard, eighteen failures, cost, and rollback/forward-fix criteria. Pass requires measured readiness, per-AZ headroom, graceful long-lived work, version-scoped alarms, and proof of restored user health.

Official sources

Advertisement