Lesson 366 · AWS Learning Path

AWS 366: EC2 and Auto Scaling deployment patterns

· Published · 5 min read

Labelled process diagram for AWS 366: Versioned intent to Automated validation to Controlled AWS change to Observed result and retained evidence, with decision, proof and rejection evidence.

Why this lesson matters

EC2 fleets combine image/revision delivery with Auto Scaling replacement, load-balancer registration, lifecycle hooks, warm-up, scaling policies, mixed instances, and state. A rollout can pass instance health while deploying the wrong AMI or exhausting one subnet.

Pattern choices

PatternMechanismStrengthMain risk
Golden AMI + instance refreshNew numbered launch-template version and rolling replacementImmutable/reproducible hostsImage pipeline/capacity time
CodeDeploy in-placeHooks update existing instancesLower temporary capacityDrift, mixed/partial host state
CodeDeploy blue-greenReplacement ASG/instances and traffic rerouteFast traffic return while old retainedDuplicate capacity and data compatibility
User-data pullInstance fetches code at bootSimple bootstrapMutable “latest”; skip-matching cannot see code change
Root-volume replacement refreshSupported mixed-instances configuration replaces root volumePreserves selected instance attachments/stateStrict requirements and no ELB support boundary

Prefer immutable images for stable fleet state; use in-place when replacement is impractical and operating controls justify it. Separate persistent data into managed/external storage.

Launch-template discipline

Use launch templates, explicit numbered versions, approved AMI ID, instance profile, security groups, user data digest, block-device mappings/encryption, metadata options, monitoring, tags, and capacity settings. Do not deploy $Latest/$Default as immutable evidence.

AMI creation must record source commit/package, patch baseline, build image/tool, tests, scan/SBOM, account/Region, AMI/snapshot IDs, encryption, launch test, and deprecation. Copying across Regions/accounts creates new image/snapshot/KMS authorization and provenance links.

Instance refresh flow

create desired launch-template version
 -> preflight quota/capacity/AZ/health/alarms
 -> start refresh with desired configuration/preferences
 -> replace batches + warm-up + health
 -> checkpoints and external validation
 -> all replacements -> bake time -> success

During a refresh with desired configuration, scale-out uses the desired configuration. Scaling policies continue, so warm-up must prevent stale metrics and over-scaling.

Healthy percentage and headroom

For desired capacity 40, minimum healthy 90%, and maximum healthy 120%, the deployment must retain roughly 36 healthy-equivalent capacity while allowing up to 48 total, subject to documented rounding/weights. The theoretical 8-instance surge may still fail due to EC2 quotas, Spot pools, subnet IPs, target limits, license/database connections, or budget.

Calculate per Availability Zone and weighted capacity for mixed-instance policies. Desired capacity must be larger than the largest weight where required. Test an AZ/instance-type capacity failure during refresh.

Warm-up, grace, draining, and lifecycle

Instance warm-up controls refresh progress and scaling metrics. Health-check grace prevents premature replacement; load-balancer health determines traffic readiness; deregistration delay drains connections. They are different timers.

Lifecycle hooks can pause launch/termination for configuration, registration, evidence, or draining. Send heartbeat/complete action within bounded timeout, make handlers idempotent, and define default result. A stuck hook can consume capacity and block refresh.

The instance must publish exact release/AMI identity and pass deep-but-bounded application validation before traffic.

Checkpoints and bake time

Checkpoints pause after replacement percentages and emit events for verification. Small fleets can skip requested percentages because replacing one instance jumps beyond them. Canceling does not restore already replaced instances, and restarting a partial refresh is not resume-from-checkpoint.

Define checkpoints from meaningful fleet counts and failure domains, not round percentages. Bake time waits after replacements before completion; choose it from delayed workload behavior. Default zero is not evidence of safety.

Skip matching

Skip matching replaces only instances whose modeled configuration differs. It can save time/cost, but it does not inspect code fetched by unchanged user data. If user data downloads mutable code without a launch-template/AMI change, disable skip matching or, preferably, make the desired artifact immutable and represented in configuration.

Understand limitations with attribute-based selection and other current features before enabling it.

Auto rollback

Auto rollback requires desired configuration and has restrictions: previous and desired launch-template versions must support rollback, numbered versions are needed instead of $Latest/$Default, and Parameter Store AMI aliases limit rollback behavior. A stable previous configuration must still launch.

CloudWatch alarms can fail/rollback a refresh when configured; alarms in ALARM or INSUFFICIENT_DATA can block start, and missing-data treatment affects progress. Up to the documented alarm limit may be specified. Rollback replaces already refreshed instances; it consumes time/capacity and does not reverse database/external writes.

Mixed instances and Spot

Mixed policies combine On-Demand/Spot, instance types/requirements, weights, and allocation strategies. During deployment, capacity rebalance/Spot interruptions and refresh replacement can overlap. Ensure baseline On-Demand or diversified capacity meets availability, and validate performance across every permitted type/architecture.

Do not bake architecture-specific binaries into one AMI and then allow incompatible types. Monitor weighted capacity, not only instance count.

Read-only inspection

aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names ASG_NAME --region ap-south-1
aws autoscaling describe-instance-refreshes --auto-scaling-group-name ASG_NAME --max-records 10 --region ap-south-1
aws ec2 describe-launch-template-versions --launch-template-id LT_ID --versions VERSION --region ap-south-1
aws elbv2 describe-target-health --target-group-arn TARGET_GROUP_ARN --region ap-south-1

Redact IDs, ARNs, IPs, AMIs, user data, roles, tags, network, and account details.

Failure game day

Inject 12 cases: AMI launch failure, wrong architecture, subnet IP exhaustion, one-AZ capacity loss, Spot interruption, target health path wrong, warm-up too short, lifecycle hook timeout, checkpoint skipped on small fleet, skip matching misses mutable code, alarm missing data, and rollback previous AMI unavailable.

For each state current/desired versions, fleet capacity, customer exposure, scaling interaction, stop/cancel/rollback/forward decision, data compatibility, and retest.

Acceptance

Design a 60-instance three-AZ release using immutable AMI refresh and compare in-place/blue-green alternatives. Submit launch-template/AMI provenance, per-AZ/weighted capacity math, preferences, checkpoints, timers, lifecycle handlers, alarms, rollback prerequisites, mixed-instance policy, telemetry, 12 failures, cost, and decommission.

Pass requires numbered immutable versions, measured headroom, exact runtime release identity, safe warm-up/drain, checkpoint semantics, verified rollback AMI/capacity, and no assumption that fleet rollback reverses state.

Official sources

Advertisement