AWS 364: Deployment configurations, health, automatic rollback, and alarms
Why this lesson matters
Deployment safety comes from the relationship between replacement rate, spare capacity, health semantics, traffic exposure, alarm delay, application validation, and reversible state. A named “automatic rollback” option cannot compensate for a wrong metric or irreversible write.
Platform-specific configurations
| Platform | Configuration controls | Important boundary |
|---|---|---|
| EC2/on-premises in-place | Minimum healthy hosts, optionally zonal settings where supported | Reduced capacity and mixed host state |
| EC2 blue-green | Replacement environment and traffic/reroute policy | Capacity, target selection, old-environment retention |
| Lambda | All-at-once, canary, or linear alias traffic | Version/alias, hook and low-volume signal |
| ECS | All-at-once, canary, or linear task-set traffic | LB/listener/target group; current NLB constraints |
A configuration defines progress rules, not business correctness. Custom canary/linear support varies by deployment method; CloudFormation-managed ECS blue-green has documented constraints.
Capacity math
For an EC2 group of 20 instances with minimum healthy hosts=90%, at least 18 must remain healthy, so at most two can be unavailable if rounding and target selection behave as expected. Verify service rounding/documentation and failure domains. If two instances are already unhealthy, deployment has no safe budget.
For per-zone resilience, calculate each Availability Zone independently. A fleet can satisfy a global minimum while one zone loses too much capacity. Include load-balancer deregistration delay, connection draining, startup, agent hooks, health-check intervals/thresholds, warm-up, and application cache population.
For a 10-percent canary receiving 200 requests/minute for five minutes, only 100 canary requests occur. If the critical workflow is 1 percent, expect one observation. That is not adequate evidence. Increase interval/cohort/synthetic test or choose a safer gate.
Layered health model
process ready -> target healthy -> API technically valid
-> dependency behavior -> business transaction correct
-> guardrails (security/data/cost) -> release accepted
Use multiple indicators:
- deployment/hook status;
- instance/task/function and target-group health;
- request rate, error rate, p95/p99 latency, saturation;
- queue age, database waits, external dependency failures;
- business success/failure and reconciliation;
- security/data integrity/residency/cost guardrails;
- exact version/digest distribution.
Health endpoints should test enough dependencies to reject unusable targets without creating cascading failure. Separate readiness from liveness. A deep health check on every request can overload dependencies; a shallow check can admit broken targets.
Alarm design
For every CloudWatch alarm define metric namespace/name, dimensions, statistic/percentile, period, evaluation periods, datapoints to alarm, threshold/comparison, missing-data behavior, low-sample behavior, action, owner, and test.
Absolute error counts mislead across changing traffic; rates need sound numerator/denominator and enough volume. Percentiles need representative samples. Composite alarms can combine guardrails but add dependency and evaluation latency.
Alarm evaluation plus metric ingestion must be faster than the traffic increment. If a canary advances after five minutes while its reliable business metric arrives after ten, automation is blind.
Test OK, ALARM, INSUFFICIENT_DATA, stale/wrong dimension, no traffic, and alarm-service permission failure. Decide fail-open/fail-closed per risk.
Stop, fail, rollback, and redeploy
- Stop/cancel prevents further progress where possible; in-flight provider work may continue.
- Failed marks the deployment unsuccessful; already changed targets/data remain.
- Automatic rollback starts a new deployment of a prior known-good revision for configured failure/alarm/stop events.
- Redeploy creates another release attempt, often after correction.
- Roll-forward deploys a new compatible fix when reverting is unsafe.
The prior revision, image, AppSpec, dependencies, keys, roles, and capacity must remain usable. Verify them periodically.
Data and schema boundary
Before traffic shift, classify changes:
| Change | Traffic rollback safety |
|---|---|
| Stateless code only | Usually safer if config/dependencies remain compatible |
| Additive nullable schema | Often compatible after tests |
| Destructive column/event change | Unsafe for mixed/old readers |
| External payment/message | Requires idempotency and reconciliation |
| Encryption/key format change | Requires old/new reader and key availability |
| Background migration | Requires checkpoint, dual-read/write or roll-forward plan |
Use expand/migrate/contract and preserve old-version compatibility through the rollback window. State the last safe rollback point and authority who closes it.
Bake time and observation
Observe long enough to encounter meaningful workload: request cycles, cache expiry, scheduled/background jobs, queue retries, database maintenance, and business settlement. A fixed 10-minute bake is not evidence for a daily process.
Define immediate automated gates plus longer post-deployment monitoring. Keep old capacity/revision until delayed risks pass or accept residual risk explicitly.
Read-only inspection
aws deploy list-deployment-configs --region ap-south-1
aws deploy get-deployment-config --deployment-config-name CONFIG_NAME --region ap-south-1
aws deploy get-deployment-group --application-name APPLICATION_NAME --deployment-group-name GROUP_NAME --region ap-south-1
aws cloudwatch describe-alarms --alarm-names ALARM_NAME --region ap-south-1
Redact names, roles, targets, alarm dimensions, ARNs, and account details.
Failure game day
Inject 12 cases: insufficient spare capacity, one-zone degradation, green passes shallow health but DB fails, canary too small, alarm wrong dimension, missing data, metric delay, composite dependency unavailable, hook false positive, automatic rollback revision unavailable, old code cannot read new writes, and rollback itself alarms.
For each record timeline, exposure, alarm state, deployment state, stop/rollback/roll-forward decision, data reconciliation, owner, communication, and changed test.
Acceptance
Submit platform/configuration matrix, capacity and canary sample calculations, layered health contract, alarm catalog, stop/fail/rollback state machine, prior-revision availability proof, schema/data compatibility, bake-time rationale, 12 game-day records, cost, and residual risks.
Pass requires measured capacity, statistically meaningful validation, tested missing-data behavior, version-bound rollback, no assumption that rollback reverses writes, and explicit human authority when automation is uncertain.