AWS 357: Rolling, blue-green and canary deployments
Why this lesson matters
Deployment strategies decide how much capacity changes at once, how many users see a defect, how long old and new versions coexist, and how quickly traffic can return. The strategy name alone says nothing about database compatibility, observability, session behavior, or recovery.
This lesson compares all-at-once, rolling, immutable, blue-green, canary, linear, and feature-flag release patterns with calculations and failure gates.
Separate the release dimensions
artifact deployment: where new binaries/configuration run
traffic release: which users/requests reach the new version
feature release: which behavior is enabled
data migration: which schema/data representation is authoritative
These can move at different times. A blue-green infrastructure switch may still expose a feature all at once. A canary traffic shift may fail because both versions share an incompatible database. A feature flag is not rollback if the new code already wrote irreversible data.
Strategy comparison
| Strategy | Capacity/change shape | Exposure | Reversal | Major weakness |
|---|---|---|---|---|
| All-at-once | Replace/shift everything | 100% immediately | Fast only if old state remains compatible | Largest blast radius |
| Rolling | Replace batches in same fleet | Grows by batch | Redeploy/roll batches | Mixed versions and reduced headroom |
| Immutable | Build replacement instances/tasks | None until replacement accepted | Keep prior fleet/image | Extra capacity/time |
| Blue-green | Two environments, traffic switch | Depends on shift | Route back while blue remains safe | Duplicate capacity and state compatibility |
| Canary | Small cohort then remainder | Bounded initially | Stop/route back | Weak signal at low traffic or wrong cohort |
| Linear | Repeated increments | Gradual | Stop/route back | Longer mixed-version window |
| Feature flag | Behavior controlled separately | Cohort/percentage | Disable flag | Flag debt and data side effects |
AWS implementations differ. CodeDeploy EC2/on-premises can use in-place or blue-green; Lambda and ECS CodeDeploy paths are blue-green with supported traffic shifting. ECS with NLB has current deployment-configuration limitations. ECS rolling updates and EKS strategies have different controllers and health semantics. Verify the exact target platform.
Requirements that choose the strategy
Collect:
- maximum user/error exposure and release window;
- required capacity/headroom and scaling speed;
- request volume needed for statistically useful validation;
- state/session affinity and long-lived connection behavior;
- API/event/schema compatibility across old and new versions;
- external side effects and idempotency;
- startup/warm-up/cache behavior;
- RTO/RPO and rollback deadline;
- compliance/approval constraints;
- observability latency and operator coverage;
- cost of duplicate capacity, transfer, logs, and tests.
Reject any strategy that violates a hard requirement before weighted scoring.
Capacity and exposure calculations
For a rolling fleet of 100 instances with 20-percent batches and minimum healthy capacity of 90, determine whether terminating 20 before launching replacements violates the minimum. A safer controller may surge replacements first, requiring at least 120 temporary capacity. Quotas, subnet IPs, target-group registration, warm-up, and downstream connection limits must support the surge.
For 12,000 requests/minute and a 5-percent canary lasting 10 minutes:
canary requests = 12,000 * 0.05 * 10 = 6,000
If a critical operation occurs in only 0.1 percent of requests, expect about six such operations. That may be insufficient to validate its error rate. Segment by journey/tenant/Region and run synthetic checks without confusing synthetic traffic with customer proof.
Define an error-budget gate. If baseline has 0.2-percent errors and canary reaches 1.0 percent over 6,000 requests, compare absolute errors, confidence/volume, latency percentiles, business failures, and guardrails. Do not promote based on averages alone.
Database and event compatibility
Use expand/migrate/contract:
- Expand schema/API/event so old and new code can coexist.
- Deploy compatible readers/writers.
- Backfill with throttling, checkpoints, and reconciliation.
- Switch authoritative behavior after evidence.
- Observe through rollback window.
- Contract only after no old consumer remains.
Examples: add nullable column before writing it; dual-read during migration; publish additive event fields; version breaking contracts; use transactional outbox/idempotency for side effects. Avoid destructive rename/drop in the same release as the new reader.
Rollback after new writes may require roll-forward, compensating transaction, or data reconciliation. Routing alone is not enough.
Traffic and cohort design
Traffic weighting can occur through load balancers, aliases, service mesh, DNS, or application flags. Each has different stickiness, cache, connection, retry, and propagation behavior.
Choose cohort intentionally: random request, stable user/tenant hash, internal users, one Region/AZ, or synthetic traffic. Random per-request can send one session to both versions; one-AZ canary may confound version with zonal conditions; internal-only traffic may not represent production workload.
Document what happens to long-lived TCP/WebSocket connections, cookies, in-flight jobs, queues, scheduled work, and background consumers when traffic weights change.
Health gates and observability
Use layers:
- platform: process/task/instance readiness and target health;
- service: request rate, errors, p95/p99 latency, saturation;
- dependency: database, queue age, external failures, throttles;
- business: successful order/payment/document outcomes;
- guardrail: security, data integrity, cost, residency;
- deployment: version/digest distribution and hook outcomes.
Every gate has query, threshold, evaluation window, missing-data behavior, owner, stop action, and recovery. Health check success is not business correctness. Alarm delay must fit traffic-shift intervals.
Strategy runbooks
Rolling
Verify headroom and compatibility, drain targets, replace bounded batch, wait readiness/warm-up, validate, continue. Stop when minimum healthy capacity, errors, latency, queue age, or replacement failures breach limits.
Blue-green
Create green from immutable revision, validate privately, synchronize compatible state, run smoke/load/security tests, shift traffic, observe, then retain blue through explicit rollback window. Prevent both environments from running singleton jobs or consuming the same queue incorrectly.
Canary/linear
Validate version and cohort, shift first increment, wait a statistically and operationally meaningful interval, compare baseline/canary, stop automatically on guardrail, continue increments, then verify 100-percent state. Do not delete old capacity immediately.
Failure game day
Work through ten injects:
- New tasks are healthy but p99 latency doubles after warm-up.
- Canary receives too little traffic for a decision.
- Old code cannot read a schema written by new code.
- Sticky sessions hide failing users from aggregate metrics.
- Queue consumers process the same message with different semantics.
- Alarm has wrong dimensions and stays
INSUFFICIENT_DATA. - Rollback restores traffic but target version wrote external payments.
- Blue and green both run a singleton scheduler.
- Capacity quota prevents green from reaching required size.
- One Availability Zone fails during a rolling update.
For each state detection, automatic/manual decision, data authority, traffic action, capacity action, reconciliation, communication, and retest.
Decision exercise and acceptance
Choose strategies for: stateless web API, payment service with writes, Lambda event processor, ECS service with NLB, 100-instance EC2 fleet, and database schema release. Provide hard requirements, rejected strategies, capacity/exposure math, cohort, gates, data compatibility, rollback deadline, costs, and game-day results.
Pass requires no strategy chosen by name alone, explicit mixed-version compatibility, enough canary sample evidence, bounded automatic stop, immutable version correlation, no claim that traffic rollback reverses data, and platform-specific AWS limitations.