Lesson 368 · AWS Learning Path

AWS 368: ECS blue-green deployment and traffic shifting

· Published · 5 min read

Labelled process diagram for AWS 368: Versioned intent to Automated validation to Controlled AWS change to Observed result and retained evidence, with decision, proof and rejection evidence.

Why this lesson matters

ECS blue-green deployment runs a replacement service revision beside the original, validates it, shifts traffic, observes it, and retains the original briefly for recovery. Safety depends on task-definition immutability, target groups, listeners, health checks, capacity, hooks, alarms, and data compatibility working together.

Deployment choices

ECS supports rolling replacement and newer native deployment strategies, while existing systems may use the CodeDeploy blue-green controller. Do not mix their configuration models. Record which controller owns traffic and rollback before changing a service.

ChoiceBest fitKey tradeoff
ECS rollingOrdinary stateless updateLower duplicate capacity, mixed revisions during rollout
ECS native blue-greenECS-managed parallel revisions and trafficExtra capacity; service configuration owns lifecycle
CodeDeploy blue-greenExisting CodeDeploy/AppSpec workflowMore resources and a separate deployment control plane
Canary or linear shiftEnough traffic and measurable riskLonger overlap and observation cost
All-at-once shiftNon-production or externally controlled trafficFast, but minimal production sampling

For CodeDeploy blue-green, the ECS service uses the CODE_DEPLOY controller. A deployment group names the cluster/service, two target groups, production listener, optional test listener, traffic configuration, alarms, rollback, and original-task-set termination wait. The AppSpec points to a task-definition revision, container name/port, and optional Lambda hooks.

Native ECS blue-green instead uses the ECS controller with a blue-green strategy and ECS service deployment configuration. It can use lifecycle hooks, CloudWatch alarms, a circuit breaker, bake time, and load-balancer advanced configuration. Confirm feature and load-balancer restrictions in current documentation for the selected mode.

End-to-end traffic flow

immutable image digest -> task-definition revision -> green tasks
 -> green target health -> optional test listener/hook
 -> production listener shifts traffic -> alarm/bake observation
 -> terminate blue, or route back and investigate

Two target groups prevent blue and green registrations from being confused. For awsvpc networking, target type is normally ip. Health-check path, port, matcher, interval, threshold, startup time, security-group rules, subnet routes, container port, and listener rule must agree. A shallow /health that ignores critical dependencies can promote a broken release; a probe that fails whenever an optional dependency fails can cause needless rollback.

CodeDeploy ECS lifecycle order includes BeforeInstall, replacement installation, AfterInstall, test traffic, AfterAllowTestTraffic, BeforeAllowTraffic, production traffic, and AfterAllowTraffic. Hooks should verify exact image/release identity and user outcomes through the intended listener, not directly call a task IP.

Capacity, scaling, and state

Blue-green can require roughly twice the task capacity during overlap, plus deployment and failure headroom. Check Fargate quotas or EC2 cluster resources, subnet IPs, ENIs, target limits, NAT paths, service quotas, database connections, licenses, and budget. Capacity-provider scaling and service auto scaling can react during deployment, so distinguish deployment growth from demand growth.

Task definitions are immutable revisions, but an image tag can move. Pin the container image digest and retain scan/SBOM/signature evidence. Give the task execution role only image/log/secret retrieval permissions and the task role only application permissions. The CodeDeploy or ECS infrastructure role needs scoped authority to manage service and load-balancer resources; hooks need separate roles.

Rollback changes traffic and desired service revision. It does not reverse database migrations, queue messages, cache changes, or calls to another system. Use expand-contract schema changes: add backward-compatible structures, deploy compatible code, migrate data, then remove old structures only after the rollback window closes. Make consumers tolerate both message versions.

Read-only inspection

aws ecs describe-services --cluster CLUSTER --services SERVICE --region ap-south-1
aws ecs describe-task-definition --task-definition FAMILY:REVISION --region ap-south-1
aws deploy get-deployment-group --application-name APP --deployment-group-name GROUP --region ap-south-1
aws deploy list-deployments --application-name APP --deployment-group-name GROUP --region ap-south-1
aws elbv2 describe-rules --listener-arn LISTENER_ARN --region ap-south-1
aws elbv2 describe-target-health --target-group-arn TARGET_GROUP_ARN --region ap-south-1

Correlate service controller/strategy, task-definition ARN, deployment/task-set status, image digest, desired/running/pending counts, listener weights, target health reasons, alarms, hook events, and termination timing. A PRIMARY deployment or healthy target alone does not prove the user transaction.

Workshop and failure game day

Design a three-AZ service with desired count 12. Calculate tasks and subnet IPs during full overlap plus one failed replacement batch. Define test traffic, a five-step linear or two-step canary shift, bake time, termination wait, and abort thresholds. Explain whether session affinity, long-lived connections, DNS, and background workers follow the load-balancer shift.

Test at least twelve failures: missing image permission, wrong architecture, container crash, insufficient CPU, subnet IP exhaustion, security-group mismatch, wrong container port, bad health matcher, test listener hitting blue, hook timeout, alarm missing data, auto-scaling interference, incompatible migration, and rollback with unhealthy blue tasks. Preserve ECS service events, stopped-task reason, target-health reason, listener rule, hook logs, alarm history, and user probe.

Cost and cleanup

Price duplicate Fargate or EC2 capacity, ALB/NLB hours and capacity units, NAT/data processing, logs, metrics, traces, ECR storage/scanning, hooks, and retained old tasks. Longer bake and termination windows buy evidence at a real cost.

This lesson creates nothing. In an approved lab, stop test executions, restore listener ownership, delete only owned services/task sets, deployment groups, target groups, listeners, images, logs, roles, and networking resources, then verify no target or ENI remains. Retain audit evidence according to policy.

Acceptance

Submit a controller decision, task-definition and digest provenance, listener/target diagram, health contract, capacity and subnet-IP math, IAM role split, hook assertions, alarm/bake/rollback policy, schema compatibility plan, fourteen failure results, cost, and cleanup. Pass requires exact traffic evidence, capacity for overlap, user-level validation, and recovery that has been tested rather than assumed.

Official sources

Advertisement