AWS 291: Modernization path from virtual machines to managed, container, or serverless services
Why this lesson matters
Moving an application from a virtual machine into a container changes its package, but does not automatically improve its architecture. Replacing a process with a Lambda function changes the runtime contract, failure behavior, scaling model, and cost unit. Replacing a database with a managed service transfers some operations to AWS, but the application team still owns schema, queries, access, recovery objectives, and data correctness.
The architect must choose the least complex operating model that satisfies the workload. That can be EC2. It can also be ECS, EKS, Lambda, Batch, or a purpose-built managed service. The correct portfolio will usually contain several models.
This lesson turns workload evidence into an incremental modernization path. AWS288 covered application decomposition and safe extraction patterns. AWS290 covered the financial case. Here you decide where each component should run, what responsibility moves, what remains, and how to migrate without coupling every change into one irreversible release.
What you will be able to do
By the end, you can:
- distinguish hosting change, replatforming, and application refactoring;
- characterize web, API, worker, scheduled, stream, and stateful components;
- choose among EC2, ECS, EKS, Fargate, Lambda, Batch, and managed services;
- explain control plane, data plane, scaling unit, isolation, and failure boundaries;
- reject container or serverless designs that violate runtime requirements;
- separate compute, data, identity, network, deployment, and observability transitions;
- design a staged migration with contracts, traffic control, rollback, and decommission gates;
- identify platform-team skills and retained operational duties;
- model complete steady-state and transition costs; and
- produce evidence strong enough for an architecture review.
Before you start
- Use supplied evidence unless the account owner authorizes read-only inspection.
- Do not create clusters, services, functions, VPC endpoints, databases, images, repositories, or commitments in this lesson.
- Never copy secrets, environment values, image credentials, customer data, account IDs, or private resource names into coursework.
- Use a non-root federated identity. Listing a resource does not authorize changing it.
- Treat service availability, quotas, runtime versions, Regions, and prices as dated evidence. Verify them for the target Region before approval.
- AWS App Runner is closed to new customers. Existing customers can continue to use it, but AWS recommends ECS Express Mode as the migration destination. Do not propose App Runner for a new customer.
1. Begin with the component, not a favorite service
Decompose the application only far enough to expose different runtime and ownership needs. A useful component record includes:
| Evidence | Questions |
|---|---|
| Entry and exit | HTTP, queue, schedule, stream, file, database call, or administrator session? |
| Runtime | Language, version, process model, startup time, duration, architecture, OS APIs, privileged operations? |
| Demand | Constant, scheduled, bursty, unpredictable, latency-sensitive, or throughput-oriented? |
| State | Local files, sessions, locks, caches, database transactions, durable messages? |
| Resources | CPU, memory, GPU, ephemeral disk, network, IOPS, ports, protocol, maximum execution time? |
| Dependencies | Libraries, sidecars, agents, licenses, host paths, kernel modules, shared files, fixed IP allowlists? |
| Reliability | Availability, RTO, RPO, retry safety, ordering, duplicate tolerance, back-pressure? |
| Security | Trust boundary, data class, internet reachability, secrets, keys, tenant isolation, audit needs? |
| Delivery | Release frequency, rollback time, test automation, artifact provenance, schema compatibility? |
| Team | Linux, containers, Kubernetes, event-driven design, and managed-service operating skills? |
A single VM can contain five candidates with different destinations: an HTTP process, a nightly report, a queue worker, a local database, and a cron cleanup script. Moving the whole VM into one large container preserves hidden coupling. First map processes, ports, files, schedules, service accounts, outbound destinations, and startup order.
2. Understand the operating models
EC2: retain server control
Choose EC2 when the workload needs operating-system control, a special driver or agent, a long-lived licensed installation, unusual networking, direct host access, or a migration step with minimal code change. EC2 also provides a useful containment destination while teams remove technical blockers.
AWS manages the physical facility and virtualization. You still own the guest OS, hardening, patching, capacity, process supervision, scaling, deployment, backups, telemetry, and application. Auto Scaling and Systems Manager reduce work; they do not remove server ownership.
Keeping a component on EC2 is not failure. It is a valid decision when requirements and skills justify it. Record the blocker and revisit date instead of forcing risk into the current wave.
Containers: standardize the deployable unit
A container image packages application files and dependencies into immutable layers. At runtime it is still an isolated process sharing a kernel, not a miniature VM. Store durable data outside the task or pod. Send logs to an external destination. Supply configuration and secrets at deployment or runtime, not in image layers.
Containerization is most useful when the process can:
- start from an image without interactive setup;
- expose explicit health and readiness behavior;
- tolerate replacement of an individual replica;
- keep durable state in an external service;
- shut down gracefully within a defined drain period; and
- run without privileged host assumptions.
Before choosing orchestration, prove the image build, software bill of materials, vulnerability policy, signature or provenance, registry lifecycle, runtime user, read-only paths, signal handling, resource limits, and architecture compatibility.
ECS: AWS-native orchestration
Amazon ECS manages task placement and service lifecycle. A task definition declares images, CPU and memory, ports, environment references, IAM roles, logging, health checks, and storage. A service maintains desired count and coordinates deployments. A task is the scheduling unit; containers in one task share its lifecycle and, with awsvpc, its network interface.
Use ECS when the team wants managed orchestration, AWS integrations, strong network and security control, and no Kubernetes API requirement. Choose its capacity separately:
- Fargate removes worker-instance provisioning. Price and size each task; observe supported configurations and platform constraints.
- EC2 capacity gives instance choice, packing control, host-level capabilities, and potential cost efficiency at stable scale, but the team owns instance lifecycle and capacity availability.
- ECS Managed Instances, where available and suitable, can transfer more instance management while retaining broader EC2 capabilities. Verify current Regional support and constraints.
ECS Express Mode is the simplified path for a new internet-facing container service. With a small set of inputs it provisions an ECS service on Fargate plus load balancing, scaling, and networking. Simplicity does not remove review of the generated architecture, IAM, exposure, logs, scaling limits, and underlying resource cost.
EKS: choose the Kubernetes interface deliberately
Amazon EKS manages the Kubernetes control plane. The organization still owns Kubernetes versions and add-ons, workload manifests, policies, namespaces, ingress, autoscaling, observability, network behavior, storage, workload security, and usually worker capacity. EKS can run pods on managed node groups, self-managed nodes, Auto Mode, or Fargate, subject to each model's constraints.
Choose EKS when Kubernetes is a strategic platform interface, required software depends on its APIs or ecosystem, or a capable platform team can operate it across enough workloads to justify the complexity. Portability is not free: cloud load balancers, IAM, storage classes, DNS, secrets, and databases remain platform-specific integrations.
Do not select EKS merely because the application is containerized. A small team with one simple API may gain cost and operational burden without receiving useful Kubernetes value.
Lambda: event-driven functions
AWS Lambda runs function code in managed execution environments in response to synchronous or asynchronous invocation and event-source polling. The function scales in concurrent execution units, not servers. It fits short, event-driven, stateless processing with supported runtimes or container images and clearly bounded resource and duration needs.
Validate current quotas and configuration ranges, including timeout, memory, ephemeral storage, deployment package, concurrency, payloads, and event-source behavior. A container image does not turn Lambda into an unlimited container runtime; the Lambda execution contract still applies.
Design for environment reuse but never depend on it. Initialize reusable clients outside the handler where appropriate, keep durable state externally, use timeouts, and make retries safe. Understand invocation semantics:
- a synchronous caller normally owns retries and response handling;
- asynchronous invocation can retry and route failed events to configured destinations;
- stream and queue event-source mappings poll and deliver batches with service-specific retry, ordering, and concurrency behavior.
Reserved concurrency can protect a downstream dependency or reserve capacity. Provisioned concurrency addresses startup latency for selected functions but adds cost. Neither fixes slow dependencies, unsafe code, or missing back-pressure.
Reject Lambda when execution can exceed its contract, the process requires a persistent server, host customization, unsupported protocol handling, extremely predictable continuously high capacity better served elsewhere, or a dependency cannot tolerate burst concurrency.
Batch and purpose-built services
AWS Batch is designed for queued batch jobs and schedules compute based on job requirements. It is usually clearer than inventing a permanent web service to execute offline work. EventBridge Scheduler can initiate timed work without retaining cron on a VM.
Managed services can remove more undifferentiated operation than changing compute alone. Evaluate RDS or Aurora for relational data, DynamoDB for suitable key-value/document access, ElastiCache for caching, S3 or EFS for file needs, SQS or SNS for decoupling, EventBridge for events, Step Functions for visible workflow state, and managed streaming or search services where requirements fit.
Managed does not mean unowned. The customer still owns data models, authorization, network access, encryption choices, retention, quotas, performance, recovery validation, and cost.
3. A decision sequence that resists hype
Apply these questions in order:
- Can a managed capability replace custom operation? Prefer a fit-for-purpose service when its contract, portability, compliance, and cost are acceptable.
- Does the component need OS or host control? If yes, retain EC2 or use container-on-EC2 only when container packaging still adds value.
- Is it naturally event-driven and bounded? Evaluate Lambda, including duration, payload, concurrency, state, latency, and downstream capacity.
- Is it asynchronous finite work? Evaluate Batch, ECS tasks, or Step Functions orchestration rather than a permanent service.
- Does it need a long-running container? Evaluate ECS first unless Kubernetes requirements are real.
- Is Kubernetes itself required and supportable? If yes, evaluate EKS and its platform operating model.
- What is the simplest model the team can operate during failure at 03:00? Include deployment, diagnosis, recovery, security, and cost, not just creation.
| Workload | Likely starting option | Reject or reconsider when |
|---|---|---|
| Legacy vendor server with kernel driver | EC2 | A supported managed or container target becomes available |
| Stateless HTTP API with several replicas | ECS on Fargate or ECS Express Mode | Host control, unusual protocol, or continuous economics require another model |
| Platform with Kubernetes-native controllers | EKS | Kubernetes is only a resume skill or no platform team exists |
| Short object-created transformation | Lambda | Duration, local resource, burst, or payload constraints do not fit |
| Hours-long queued scientific jobs | AWS Batch | Interactive low-latency service behavior is required |
| Relational database with routine administration burden | RDS or Aurora candidate | Engine features, licensing, latency, topology, or control requirements do not fit |
4. Separate the transitions
Do not change six planes in one release unless the business risk demands it:
- Artifact: manual server build to versioned image or function package.
- Compute: VM process to task, pod, function, or managed job.
- Data: local or self-managed store to external or managed data service.
- Integration: direct calls to stable APIs, queues, events, or workflows.
- Traffic: DNS, load balancer, API gateway, queue mapping, or consumer weight.
- Operations: identity, secrets, telemetry, deployment, scaling, recovery, and support ownership.
A reversible path might externalize sessions, move file data, build an immutable image, deploy it to ECS beside the VM, mirror safe traffic, shift a small percentage, observe business and technical metrics, then retire the VM. Database conversion can remain a separately governed change.
Use branch by abstraction when both old and new implementations must sit behind a stable interface. Use a strangler route when a proxy or event boundary can send one capability to the new implementation. Use expand-and-contract for schema and event evolution: add backward-compatible fields, migrate readers and writers, verify, then remove the old shape.
5. State, messages, and correctness
Local state is the most common hidden blocker. Classify every write:
- cache that can be lost;
- temporary data needed only during one attempt;
- session state needed by later requests;
- durable business record;
- shared file contract;
- coordination lock; or
- audit evidence.
Give each class an explicit owner, service, durability, retention, encryption, backup, and recovery test. Do not mount shared storage merely to preserve an accidental local-file design without checking concurrency and consistency.
Scaling creates duplicate and out-of-order work. Assign idempotency keys, deduplication windows, transaction boundaries, retry limits, dead-letter handling, poison-message quarantine, and replay procedures. A queue decouples availability only when producers, consumers, retention, visibility timeout, and backlog alarms are designed together.
For a database move, define source of truth and write ownership at every phase. Dual writes without an outbox, reconciliation, and failure policy commonly diverge. Rollback after new writes requires reverse replication, replay, compensation, or a business-approved loss boundary. “Point DNS back” is not a data rollback plan.
6. Identity, network, and security ownership
Replace long-lived server credentials with workload identities:
- EC2 instance profiles for instances;
- ECS task roles for application containers, separate from task execution roles;
- EKS workload identity mechanisms for pods, separate from node roles; and
- Lambda execution roles for functions.
Scope permissions to actions and resources, use resource policies where the target service requires them, and test both allowed and denied behavior. Secrets belong in an approved secret store with rotation and controlled retrieval, not images, task definitions, manifests, function variables, or logs.
Draw ingress and egress separately. Include public or private endpoints, subnets, security groups, network ACLs, load balancers, API Gateway, service discovery, DNS, NAT or endpoints, inspection, inter-AZ paths, and downstream allowlists. Fargate removes instance management, not VPC design. Lambda can use service networking without your VPC, or connect to VPC resources through configured subnets and security groups; that decision affects reachability, IP capacity, and egress.
Scan source dependencies and artifacts. Generate provenance and an SBOM, scan images and packages, pin immutable versions, establish critical-vulnerability gates and exceptions, use non-root execution where supported, minimize writable paths, and rebuild rather than patching a running container.
7. Scaling and failure behavior
Name the scaling signal and bottleneck. CPU may work for compute-bound services; request count, queue depth per worker, concurrency, latency, or custom business throughput can be more meaningful. Define minimum, maximum, target, cooldown, startup time, drain behavior, and downstream capacity.
Scaling one tier can overload the next. One thousand function invocations can exhaust a database connection limit. More consumers can increase lock contention or violate an external API quota. Use admission control, reserved concurrency, queue buffering, connection pooling, rate limits, and load shedding as appropriate.
Map failure at each layer:
client -> DNS/edge -> entry service -> compute replica -> dependency -> data store
| | |
v v v
auth/routing retry/scale quota/recovery
Test loss of one replica, one Availability Zone where the design claims tolerance, dependency timeout, bad deployment, exhausted quota, expired secret, queue backlog, poison event, and partial data migration. Verify detection, impact, automated response, operator action, rollback, and recovery time.
8. Delivery and observability contract
Use one immutable artifact through environments. A deployment must define readiness, health, graceful shutdown, drain time, rollout policy, rollback trigger, and database compatibility. Blue/green and canary reduce exposure only when metrics can distinguish versions and rollback remains data-safe.
Correlate a request or event across old and new systems. Capture deployment version, trace or correlation ID, latency, errors, saturation, restarts, throttles, concurrency, queue age, failed records, dependency health, and business outcomes. Central logging needs redaction, access control, retention, query cost, and outage behavior.
Do not call a migration successful because tasks are RUNNING or a function returned HTTP 200. Acceptance includes output correctness, authorization, service-level objectives, recovery, operational handoff, cost, and source decommissioning.
9. Cost and operating model
Compare equivalent quality of service and include transition cost.
| Model | Important cost drivers |
|---|---|
| EC2 | Instance time, OS/software, EBS, load balancing, scaling headroom, operations |
| ECS on EC2 | EC2 capacity and unused packing space, EBS, load balancing, registry, orchestration operations |
| ECS/EKS on Fargate | Requested task or pod CPU, memory, ephemeral storage, running duration, public IPv4 where used |
| EKS | Cluster charge and any version-related tier, worker capacity, add-ons, load balancers, logs, platform team |
| Lambda | Requests, duration and configured resources, provisioned concurrency, transfer, logs, connected services |
| Managed services | Capacity or requests, storage, I/O, backup, transfer, replicas, support, observability |
Add NAT processing, endpoints, inter-AZ transfer, load-balancer units, image storage/scanning, logs/metrics/traces, backup, security services, non-production, support, training, platform engineering, dual run, and decommissioning. Serverless can be economical for variable or intermittent demand; continuously high demand may favor a different model after measurement. Never compare an unprotected single VM with a Multi-AZ managed design as if service quality were equal.
10. Read-only evidence, when authorized
These commands inventory names only. They do not prove configuration or suitability and can reveal identifiers, so redact shared output.
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws ec2 describe-instances \
--filters Name=instance-state-name,Values=pending,running,stopping,stopped \
--query 'Reservations[].Instances[].{Id:InstanceId,Type:InstanceType,State:State.Name}'
aws ecs list-clusters
aws eks list-clusters
aws lambda list-functions \
--query 'Functions[].{Name:FunctionName,Runtime:Runtime,Memory:MemorySize,Timeout:Timeout}'
aws rds describe-db-instances \
--query 'DBInstances[].{Id:DBInstanceIdentifier,Engine:Engine,MultiAZ:MultiAZ}'
Continue only for resources the owner identifies. Capture configuration, tags, metrics, deployments, events, IAM relationships, networking, logs, backups, costs, and timestamps. A list call is discovery, not end-to-end evidence.
11. Guided modernization-path workshop
Use the supplied application with these components:
- a Linux/Nginx web tier with local sessions and uploaded files;
- a Java order API with a fixed pool of database connections;
- a queue worker whose messages can be delivered more than once;
- a four-hour nightly report needing 12 GB memory and 20 GB scratch space;
- a one-minute cleanup cron job;
- PostgreSQL on a VM with a reporting replica;
- a third-party licensing daemon requiring a host-bound license; and
- a stream processor that maintains local checkpoints.
Demand triples during quarterly close. The target requires Multi-AZ service, RTO 60 minutes, RPO 15 minutes, no public database, auditable deployment, and rollback without losing accepted orders.
Produce these artifacts:
- A process, port, state, dependency, identity, and data-flow inventory.
- A current responsibility matrix covering OS through application and data.
- At least three portfolio options: conservative EC2/replatform, ECS-centered, and mixed managed/serverless.
- A component decision table with evidence, rejected alternatives, confidence, and revisit trigger.
- An explicit decision on ECS versus EKS and EC2 versus Fargate capacity.
- A Lambda fit test for the cleanup job and reasons the four-hour report cannot simply become one function.
- A managed PostgreSQL assessment that separates engine compatibility, extensions, connection behavior, availability, recovery, and migration.
- A state-externalization and idempotency plan.
- Identity and bidirectional network-flow diagrams.
- A six-stage transition with contracts, traffic percentages, data ownership, gates, and rollback.
- Scaling math for quarterly close plus database connection protection.
- Failure experiments and observable acceptance signals.
- A 12-month cost model with transition and retained-source cost.
- An operating model naming platform, application, security, database, network, and finance owners.
A strong answer will probably retain the licensing daemon on EC2 initially. It may use ECS for the web/API/worker, Batch or an ECS task for reporting, Lambda for bounded cleanup, and a managed database only after compatibility proof. This is not the only valid answer. Evidence must drive the choice.
12. Diagnose weak modernization plans
| Symptom | Likely mistake | Repair |
|---|---|---|
| One container per old VM | Packaging changed but component boundaries did not | Inventory processes and state, then separate only where valuable |
| EKS chosen for portability | Kubernetes integration and team cost were omitted | Document required APIs, platform ownership, and cloud dependencies |
| Lambda times out or floods a database | Runtime or concurrency contract was ignored | Redesign work units, buffer, limit concurrency, or choose another compute model |
| Scaling creates duplicate orders | Delivery retries and idempotency were absent | Add stable keys, transactional handling, reconciliation, and replay tests |
| Rollback loses new data | Traffic rollback was mistaken for data rollback | Define write ownership, reverse flow, compensation, or stop-write gate |
| Fargate service cannot reach the internet or dependency | Route, NAT, endpoint, DNS, or security group is missing | Trace both directions and verify the exact destination |
| Container works only as root | Host assumptions were packaged unchanged | Remove privilege, externalize state, and define writable paths |
| Cost rises after modernization | HA, logs, transfer, idle minimums, or transition were omitted | Normalize service level and model all connected resources |
| Team cannot restore service | Responsibility transfer was assumed, not designed | Create runbooks, access, alarms, game days, and handoff gates |
Knowledge check
- Why is containerizing a VM not automatically modernization?
It changes packaging but can preserve coupling, state, patching assumptions, and weak delivery practices.
- When is EC2 the responsible choice?
When host control or compatibility is required, or as a governed intermediate step with a revisit trigger.
- What separates ECS from Fargate?
ECS orchestrates tasks and services; Fargate is one compute-capacity option for running them.
- Why should EKS require an affirmative reason?
Kubernetes provides a powerful platform contract but introduces lifecycle, policy, add-on, skills, and cost responsibilities.
- Why can Lambda scaling damage a healthy database?
Invocation concurrency can grow faster than downstream connection or transaction capacity.
- What makes a queue consumer safe to retry?
Defined idempotency, transaction boundaries, visibility timing, retry limits, failed-message handling, and reconciliation.
- Why is DNS reversal not enough for rollback?
New durable writes may exist only in the target and require replay, replication, compensation, or an explicit loss decision.
- What proves modernization success?
Correct behavior, security, SLOs, recovery, operability, economics, accountable ownership, and retired source obligations.
Lesson acceptance
You may continue only when your submission contains:
- complete workload-characteristic and current-responsibility evidence;
- at least three viable portfolio options with rejected alternatives;
- component decisions spanning VM, container, function, batch, and managed-service boundaries;
- explicit ECS/EKS, EC2/Fargate, and event-driven fit reasoning;
- state, identity, network, scaling, deployment, and observability designs;
- a staged transition with separate compute and data rollback;
- failure experiments and measurable gates;
- complete steady-state and transition cost drivers;
- operating owners and skills gaps; and
- a decision record whose assumptions and revisit triggers are traceable.