AWS 366: EC2 and Auto Scaling deployment patterns
Why this lesson matters
EC2 fleets combine image/revision delivery with Auto Scaling replacement, load-balancer registration, lifecycle hooks, warm-up, scaling policies, mixed instances, and state. A rollout can pass instance health while deploying the wrong AMI or exhausting one subnet.
Pattern choices
| Pattern | Mechanism | Strength | Main risk |
|---|---|---|---|
| Golden AMI + instance refresh | New numbered launch-template version and rolling replacement | Immutable/reproducible hosts | Image pipeline/capacity time |
| CodeDeploy in-place | Hooks update existing instances | Lower temporary capacity | Drift, mixed/partial host state |
| CodeDeploy blue-green | Replacement ASG/instances and traffic reroute | Fast traffic return while old retained | Duplicate capacity and data compatibility |
| User-data pull | Instance fetches code at boot | Simple bootstrap | Mutable “latest”; skip-matching cannot see code change |
| Root-volume replacement refresh | Supported mixed-instances configuration replaces root volume | Preserves selected instance attachments/state | Strict requirements and no ELB support boundary |
Prefer immutable images for stable fleet state; use in-place when replacement is impractical and operating controls justify it. Separate persistent data into managed/external storage.
Launch-template discipline
Use launch templates, explicit numbered versions, approved AMI ID, instance profile, security groups, user data digest, block-device mappings/encryption, metadata options, monitoring, tags, and capacity settings. Do not deploy $Latest/$Default as immutable evidence.
AMI creation must record source commit/package, patch baseline, build image/tool, tests, scan/SBOM, account/Region, AMI/snapshot IDs, encryption, launch test, and deprecation. Copying across Regions/accounts creates new image/snapshot/KMS authorization and provenance links.
Instance refresh flow
create desired launch-template version
-> preflight quota/capacity/AZ/health/alarms
-> start refresh with desired configuration/preferences
-> replace batches + warm-up + health
-> checkpoints and external validation
-> all replacements -> bake time -> success
During a refresh with desired configuration, scale-out uses the desired configuration. Scaling policies continue, so warm-up must prevent stale metrics and over-scaling.
Healthy percentage and headroom
For desired capacity 40, minimum healthy 90%, and maximum healthy 120%, the deployment must retain roughly 36 healthy-equivalent capacity while allowing up to 48 total, subject to documented rounding/weights. The theoretical 8-instance surge may still fail due to EC2 quotas, Spot pools, subnet IPs, target limits, license/database connections, or budget.
Calculate per Availability Zone and weighted capacity for mixed-instance policies. Desired capacity must be larger than the largest weight where required. Test an AZ/instance-type capacity failure during refresh.
Warm-up, grace, draining, and lifecycle
Instance warm-up controls refresh progress and scaling metrics. Health-check grace prevents premature replacement; load-balancer health determines traffic readiness; deregistration delay drains connections. They are different timers.
Lifecycle hooks can pause launch/termination for configuration, registration, evidence, or draining. Send heartbeat/complete action within bounded timeout, make handlers idempotent, and define default result. A stuck hook can consume capacity and block refresh.
The instance must publish exact release/AMI identity and pass deep-but-bounded application validation before traffic.
Checkpoints and bake time
Checkpoints pause after replacement percentages and emit events for verification. Small fleets can skip requested percentages because replacing one instance jumps beyond them. Canceling does not restore already replaced instances, and restarting a partial refresh is not resume-from-checkpoint.
Define checkpoints from meaningful fleet counts and failure domains, not round percentages. Bake time waits after replacements before completion; choose it from delayed workload behavior. Default zero is not evidence of safety.
Skip matching
Skip matching replaces only instances whose modeled configuration differs. It can save time/cost, but it does not inspect code fetched by unchanged user data. If user data downloads mutable code without a launch-template/AMI change, disable skip matching or, preferably, make the desired artifact immutable and represented in configuration.
Understand limitations with attribute-based selection and other current features before enabling it.
Auto rollback
Auto rollback requires desired configuration and has restrictions: previous and desired launch-template versions must support rollback, numbered versions are needed instead of $Latest/$Default, and Parameter Store AMI aliases limit rollback behavior. A stable previous configuration must still launch.
CloudWatch alarms can fail/rollback a refresh when configured; alarms in ALARM or INSUFFICIENT_DATA can block start, and missing-data treatment affects progress. Up to the documented alarm limit may be specified. Rollback replaces already refreshed instances; it consumes time/capacity and does not reverse database/external writes.
Mixed instances and Spot
Mixed policies combine On-Demand/Spot, instance types/requirements, weights, and allocation strategies. During deployment, capacity rebalance/Spot interruptions and refresh replacement can overlap. Ensure baseline On-Demand or diversified capacity meets availability, and validate performance across every permitted type/architecture.
Do not bake architecture-specific binaries into one AMI and then allow incompatible types. Monitor weighted capacity, not only instance count.
Read-only inspection
aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names ASG_NAME --region ap-south-1
aws autoscaling describe-instance-refreshes --auto-scaling-group-name ASG_NAME --max-records 10 --region ap-south-1
aws ec2 describe-launch-template-versions --launch-template-id LT_ID --versions VERSION --region ap-south-1
aws elbv2 describe-target-health --target-group-arn TARGET_GROUP_ARN --region ap-south-1
Redact IDs, ARNs, IPs, AMIs, user data, roles, tags, network, and account details.
Failure game day
Inject 12 cases: AMI launch failure, wrong architecture, subnet IP exhaustion, one-AZ capacity loss, Spot interruption, target health path wrong, warm-up too short, lifecycle hook timeout, checkpoint skipped on small fleet, skip matching misses mutable code, alarm missing data, and rollback previous AMI unavailable.
For each state current/desired versions, fleet capacity, customer exposure, scaling interaction, stop/cancel/rollback/forward decision, data compatibility, and retest.
Acceptance
Design a 60-instance three-AZ release using immutable AMI refresh and compare in-place/blue-green alternatives. Submit launch-template/AMI provenance, per-AZ/weighted capacity math, preferences, checkpoints, timers, lifecycle handlers, alarms, rollback prerequisites, mixed-instance policy, telemetry, 12 failures, cost, and decommission.
Pass requires numbered immutable versions, measured headroom, exact runtime release identity, safe warm-up/drain, checkpoint semantics, verified rollback AMI/capacity, and no assumption that fleet rollback reverses state.