AWS 368: ECS blue-green deployment and traffic shifting
Why this lesson matters
ECS blue-green deployment runs a replacement service revision beside the original, validates it, shifts traffic, observes it, and retains the original briefly for recovery. Safety depends on task-definition immutability, target groups, listeners, health checks, capacity, hooks, alarms, and data compatibility working together.
Deployment choices
ECS supports rolling replacement and newer native deployment strategies, while existing systems may use the CodeDeploy blue-green controller. Do not mix their configuration models. Record which controller owns traffic and rollback before changing a service.
| Choice | Best fit | Key tradeoff |
|---|---|---|
| ECS rolling | Ordinary stateless update | Lower duplicate capacity, mixed revisions during rollout |
| ECS native blue-green | ECS-managed parallel revisions and traffic | Extra capacity; service configuration owns lifecycle |
| CodeDeploy blue-green | Existing CodeDeploy/AppSpec workflow | More resources and a separate deployment control plane |
| Canary or linear shift | Enough traffic and measurable risk | Longer overlap and observation cost |
| All-at-once shift | Non-production or externally controlled traffic | Fast, but minimal production sampling |
For CodeDeploy blue-green, the ECS service uses the CODE_DEPLOY controller. A deployment group names the cluster/service, two target groups, production listener, optional test listener, traffic configuration, alarms, rollback, and original-task-set termination wait. The AppSpec points to a task-definition revision, container name/port, and optional Lambda hooks.
Native ECS blue-green instead uses the ECS controller with a blue-green strategy and ECS service deployment configuration. It can use lifecycle hooks, CloudWatch alarms, a circuit breaker, bake time, and load-balancer advanced configuration. Confirm feature and load-balancer restrictions in current documentation for the selected mode.
End-to-end traffic flow
immutable image digest -> task-definition revision -> green tasks
-> green target health -> optional test listener/hook
-> production listener shifts traffic -> alarm/bake observation
-> terminate blue, or route back and investigate
Two target groups prevent blue and green registrations from being confused. For awsvpc networking, target type is normally ip. Health-check path, port, matcher, interval, threshold, startup time, security-group rules, subnet routes, container port, and listener rule must agree. A shallow /health that ignores critical dependencies can promote a broken release; a probe that fails whenever an optional dependency fails can cause needless rollback.
CodeDeploy ECS lifecycle order includes BeforeInstall, replacement installation, AfterInstall, test traffic, AfterAllowTestTraffic, BeforeAllowTraffic, production traffic, and AfterAllowTraffic. Hooks should verify exact image/release identity and user outcomes through the intended listener, not directly call a task IP.
Capacity, scaling, and state
Blue-green can require roughly twice the task capacity during overlap, plus deployment and failure headroom. Check Fargate quotas or EC2 cluster resources, subnet IPs, ENIs, target limits, NAT paths, service quotas, database connections, licenses, and budget. Capacity-provider scaling and service auto scaling can react during deployment, so distinguish deployment growth from demand growth.
Task definitions are immutable revisions, but an image tag can move. Pin the container image digest and retain scan/SBOM/signature evidence. Give the task execution role only image/log/secret retrieval permissions and the task role only application permissions. The CodeDeploy or ECS infrastructure role needs scoped authority to manage service and load-balancer resources; hooks need separate roles.
Rollback changes traffic and desired service revision. It does not reverse database migrations, queue messages, cache changes, or calls to another system. Use expand-contract schema changes: add backward-compatible structures, deploy compatible code, migrate data, then remove old structures only after the rollback window closes. Make consumers tolerate both message versions.
Read-only inspection
aws ecs describe-services --cluster CLUSTER --services SERVICE --region ap-south-1
aws ecs describe-task-definition --task-definition FAMILY:REVISION --region ap-south-1
aws deploy get-deployment-group --application-name APP --deployment-group-name GROUP --region ap-south-1
aws deploy list-deployments --application-name APP --deployment-group-name GROUP --region ap-south-1
aws elbv2 describe-rules --listener-arn LISTENER_ARN --region ap-south-1
aws elbv2 describe-target-health --target-group-arn TARGET_GROUP_ARN --region ap-south-1
Correlate service controller/strategy, task-definition ARN, deployment/task-set status, image digest, desired/running/pending counts, listener weights, target health reasons, alarms, hook events, and termination timing. A PRIMARY deployment or healthy target alone does not prove the user transaction.
Workshop and failure game day
Design a three-AZ service with desired count 12. Calculate tasks and subnet IPs during full overlap plus one failed replacement batch. Define test traffic, a five-step linear or two-step canary shift, bake time, termination wait, and abort thresholds. Explain whether session affinity, long-lived connections, DNS, and background workers follow the load-balancer shift.
Test at least twelve failures: missing image permission, wrong architecture, container crash, insufficient CPU, subnet IP exhaustion, security-group mismatch, wrong container port, bad health matcher, test listener hitting blue, hook timeout, alarm missing data, auto-scaling interference, incompatible migration, and rollback with unhealthy blue tasks. Preserve ECS service events, stopped-task reason, target-health reason, listener rule, hook logs, alarm history, and user probe.
Cost and cleanup
Price duplicate Fargate or EC2 capacity, ALB/NLB hours and capacity units, NAT/data processing, logs, metrics, traces, ECR storage/scanning, hooks, and retained old tasks. Longer bake and termination windows buy evidence at a real cost.
This lesson creates nothing. In an approved lab, stop test executions, restore listener ownership, delete only owned services/task sets, deployment groups, target groups, listeners, images, logs, roles, and networking resources, then verify no target or ENI remains. Retain audit evidence according to policy.
Acceptance
Submit a controller decision, task-definition and digest provenance, listener/target diagram, health contract, capacity and subnet-IP math, IAM role split, hook assertions, alarm/bake/rollback policy, schema compatibility plan, fourteen failure results, cost, and cleanup. Pass requires exact traffic evidence, capacity for overlap, user-level validation, and recovery that has been tested rather than assumed.