AWS 162: Architecture: choose EC2, Lambda, ECS, Fargate, EKS, Beanstalk, ECS Express Mode, or Batch
Why this lesson matters
Choose EC2, Lambda, ECS, Fargate, EKS, Beanstalk, ECS Express Mode, or Batch from workload and team constraints.
Compute selection assigns operational work. The best platform is the simplest one that satisfies execution duration, protocol, state, scaling, isolation, hardware, portability and team constraints with tested recovery. “Serverless,” “containers,” and “managed” do not remove architecture responsibilities; they move them.
What you will be able to do
By the end, you can:
- explain architecture: choose ec2, lambda, ecs, fargate, eks, beanstalk, ecs express mode, or batch in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Choose EC2, Lambda, ECS, Fargate, EKS, Beanstalk, ECS Express Mode, or Batch from workload and team constraints. |
| Scope and boundary | The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Compute platform architecture decision. |
| Evidence of success | Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Compute platform architecture decision. |
| Cost model | Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design. |
| Safe rejection rule | Avoid platform fashion, hidden generated resources, or teaching App Runner to new customers after access closure. |
How the request flows
+------------------------+
| Workload constraints |
+------------------------+
|
v
+------------------------------+
| Candidate execution models |
+------------------------------+
|
v
+------------------------------------+
| Selected platform and guardrails |
+------------------------------------+
|
v
+---------------------------------------------------+
| Deployment, health, recovery, and cost evidence |
+---------------------------------------------------+
For Compute platform architecture decision, the important boundary is this: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Compute platform architecture decision. Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Compute platform architecture decision. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use the simplest platform that meets the measurable requirement and can be safely operated by the team. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid platform fashion, hidden generated resources, or teaching App Runner to new customers after access closure. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
Workload-first selector
| Platform | Strong starting fit | Reject or qualify when |
|---|---|---|
| EC2 + ASG | OS/kernel control, legacy agents, special networking/storage/hardware, steady fleets | team cannot own AMI patching/capacity/host security; burst/scale-to-zero dominates |
| default Lambda | short event/request handlers, bursty traffic, scale-to-zero economics | standard invocation duration/state/server/port/GPU or sustained-cost requirements fail; also evaluate durable/Managed Instances explicitly |
| ECS on EC2/ECS Managed Instances | container orchestration with AWS-native API and fleet economics/control | Kubernetes API/ecosystem is mandatory or team cannot own selected capacity boundary |
| ECS on Fargate | stateless/finite containers without node operations | unsupported host/GPU/privileged/storage need or steady reservations cost more than justified alternatives |
| ECS Express Mode | fast stateless HTTPS container deployment with standard Fargate/ALB topology | custom topology/governance/shared-ALB or generated-resource behavior fails requirements |
| EKS managed nodes/Auto Mode/Fargate | Kubernetes API/ecosystem/portability and a capable platform team | Kubernetes complexity has no measurable product/organization value |
| Elastic Beanstalk | supported language/Docker web or worker app with guided EC2 environment | platform lifecycle/customization or container-orchestration needs exceed model |
| AWS Batch | finite queued jobs, arrays/dependencies, fair-share and specialized/Spot compute | always-on request service or custom workflow is required |
Also compare Step Functions/Lambda durable functions for orchestration, and App Runner only for existing-customer migration because onboarding is closed to new customers. A product can combine platforms: API on ECS, event handler on Lambda and nightly jobs on Batch - but every boundary adds identity, delivery, observability and failure coupling.
Score mandatory dimensions
Document runtime/duration, inbound protocol/connection, average/peak/burst, CPU-memory-GPU/architecture, local/durable state, startup/latency, tenancy/isolation, network placement, availability/RPO/RTO, deployment frequency, compliance, current quotas/Regions, portability and team on-call skill. Make mandatory constraints pass/fail before weighted scoring.
Calculate monthly compute at average and peak, idle floor, ALB/NAT/public IPv4, logs/metrics, storage, image scanning, data transfer, control-plane/management fees and engineering labor. Compare Savings Plans/RI/Spot only after capacity model. A cheaper runtime that needs an unsupported operator team is not lower TCO.
For migration, define image/package and configuration contract, externalize sessions/data, add health/telemetry, shadow/canary traffic, reconcile state, cut over DNS/routing and retain reversible old capacity until acceptance. Avoid simultaneous writes without a conflict/idempotency plan.
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Use the Console service search and open EC2 Instances, Lambda Functions, ECS Services, EKS Clusters, Elastic Beanstalk Environments, and Batch Job queues; confirm the account and Region before reading the page.
- Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
- Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
- Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws ec2 describe-instances --query 'Reservations[].Instances[?State.Name!=`terminated`].InstanceId' --output text
aws lambda list-functions --query 'Functions[].FunctionName' --output text
aws ecs list-clusters --output text
aws eks list-clusters --output text
aws elasticbeanstalk describe-environments --query 'Environments[].EnvironmentName' --output text
aws batch describe-job-queues --query 'jobQueues[].jobQueueName' --output text
Expected interpretation
Inventory cannot choose a platform. The decision needs duration, protocol, state, packaging, scaling, startup, hardware, portability, control, skills, reliability, security, and total cost.
Practical work
Write p08-platform-adr.md for a synchronous API, event handler, scheduled batch, stateful legacy service, Kubernetes product, and quick internal web app. Select one platform for each and reject two alternatives.
Include numeric demand/cost, identity/network/data path, deployment/rollback, scaling/downstream guard, AZ/Region recovery, observability and operator ownership. Run table-top failures: bad release, one-AZ loss, registry outage, exhausted concurrency/nodes/IPs, stateful restart, downstream throttle and platform/runtime retirement. State the evidence that would reverse every selection.
Diagnose this topic from its own evidence
A bad selection reveals itself as permanent adapters: a 15-minute Lambda repeatedly timing out, Kubernetes operated only through vendor tickets, Fargate requiring prohibited host access, Batch emulating a web server, or EC2 scaling to zero with unsafe boot time. Return to the failed requirement and change platform/architecture; do not hide mismatch with retries or broad privileges. Validate costs against observed usage and team incident performance after launch.
Cost and cleanup
Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Choose EC2, Lambda, ECS, Fargate, EKS, Beanstalk, ECS Express Mode, or Batch from workload and team constraints.
- Which scope or ownership boundary must be proved first?
Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Compute platform architecture decision.
- What evidence is strong enough to accept the result?
Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Compute platform architecture decision.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid platform fashion, hidden generated resources, or teaching App Runner to new customers after access closure.
- Which cost dimensions and retained resources need an owner?
Expected direction: Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Lesson acceptance
Pass when all workload decisions trace to numeric mandatory requirements, assign shared responsibility, reject two credible alternatives, and include security, scaling, failure, migration/rollback and full TCO. Fail if keywords choose platforms, Kubernetes is selected for résumé value, serverless means no operations, or current App Runner/Beanstalk/platform lifecycle is ignored.