AWS 288: Replatforming, refactoring and application modernization
Why this lesson matters
Modernization is not “replace EC2 with containers” or “split a monolith into microservices.” It changes how software is owned, deployed, scaled, secured, observed and recovered. A managed database can remove undifferentiated operations but introduce feature limits. Containers can standardize deployment but add image, orchestration and platform responsibilities. Serverless can match event-driven demand but impose runtime, concurrency and integration constraints. Microservices can enable team autonomy only when service and data boundaries are genuinely independent.
The architect's job is to connect each change to a measured business outcome and deliver it in reversible stages. A well-structured modular monolith may be safer and cheaper than a distributed system. Rehost now and modernize later can be the correct strategy when a data-center deadline conflicts with the time needed to redesign and prove behavior.
Outcomes
By the end, you can:
- distinguish rehost, replatform, refactor/re-architect and ordinary optimization;
- assess strategic, functional, technical, financial and digital readiness;
- select managed database, container, serverless and event/integration patterns from constraints;
- identify business capabilities, subdomains, transactions and team-aligned boundaries;
- explain why code boundaries without data ownership create a distributed monolith;
- use strangler fig, branch by abstraction and anti-corruption layers incrementally;
- design saga, outbox, idempotency and reconciliation for distributed workflows;
- define observability, deployment, security, resilience and operating-model changes;
- measure value and total cost without counting service adoption as success; and
- produce a staged modernization blueprint for a supplied monolith.
Strategy boundaries
| Strategy | Amount of architectural change | Example |
|---|---|---|
| Rehost | minimal workload change | VM to EC2 using MGN |
| Replatform | targeted platform substitution without redesigning the core business architecture | self-managed PostgreSQL to compatible RDS/Aurora; app to managed containers |
| Refactor/re-architect | substantial code, data and interaction redesign for cloud-native outcomes | monolith capability extracted into independently owned service |
| Optimize | improve an already selected design | right-size, tune cache, change autoscaling policy |
One application can combine strategies. The web tier may replatform to ECS, the database replatform to Aurora, a batch job refactor to events/Lambda, and a proprietary reporting module remain temporarily on EC2. Record component-level decisions and transition architecture.
Modernization is optional, not automatically superior
Modernize when evidence links current constraints to desired outcomes such as faster safe releases, independent scaling, improved recovery, reduced licensing, better product experimentation or lower operational toil.
Reject or defer when:
- the application is near retirement or replacement;
- no business outcome or owner funds the change;
- source code, tests or domain expertise are absent;
- a fixed migration deadline leaves no time to redesign and validate;
- vendor support or compliance forbids the target;
- traffic/operations do not justify added distributed-system complexity; or
- the organization lacks platform, security, SRE and product ownership capacity.
AWS large-migration guidance notes that refactor is the most complex migration strategy and often recommends rehost/replatform first, then modernize. Treat that as a tradeoff, not a universal rule.
Five readiness lenses
AWS modernization assessment guidance uses five useful lenses:
- Strategic/business fit: differentiation, revenue, customer pain, roadmap and urgency.
- Functional adequacy: missing capabilities, usability, workflow and product fit.
- Technical adequacy: architecture, quality, testability, dependencies, security and operability.
- Financial fit: current run/change cost, licenses, one-time investment and expected benefit.
- Digital readiness: teams, product model, DevSecOps, platform, data, governance and skills.
Assess more than code. A technically extractable service can fail if no team owns its API, on-call, budget and data. Record evidence, confidence, prerequisites and gaps. Produce a roadmap, one or two detailed blueprints/MVPs and an action plan, not a portfolio of vague aspirations.
Baseline before choosing technology
Capture:
- business transactions, users, peak/seasonal demand and SLOs;
- deployment frequency, lead time, change-failure rate and recovery time;
- architecture/components, code languages/frameworks and support dates;
- inbound/outbound dependencies, protocols, latency and failure behavior;
- database tables, ownership, transactions, reports and batch jobs;
- state/session/file/cache and consistency requirements;
- security boundaries, identities, secrets and compliance evidence;
- build/test/deploy automation and test coverage;
- incidents, toil, bottlenecks, scaling constraints and technical debt;
- full infrastructure, platform, license and labor cost; and
- team topology, domain knowledge and operating responsibilities.
Without a baseline, “30 percent faster” or “more resilient” cannot be accepted later.
Replatform options and hidden work
Managed database
Potential value: automated infrastructure provisioning, backups, patching, monitoring, Multi-AZ and supported scaling. Hidden work: engine/version/extension compatibility, superuser restrictions, storage/IO behavior, parameters, connections, certificates, authentication, maintenance, backup/restore and failure tests. Heterogeneous engines also require schema/code/data conversion from AWS284.
Managed containers
Potential value: immutable packaging, scheduler placement, health-based replacement and consistent pipelines. Hidden work: image supply chain, registry lifecycle, task/pod identity, secrets, networking, service discovery, load balancing, logging, autoscaling, storage, draining, deployment rollback, capacity and platform ownership.
ECS reduces Kubernetes-specific operations; EKS supports Kubernetes ecosystems and portability but requires cluster/add-on/governance skill. Fargate removes worker-node management but not application, IAM, network, image, observability or cost design. App Runner can simplify suitable web/service workloads but is not a universal container platform.
Managed messaging/integration
Replacing self-managed brokers or schedulers can reduce operations, but semantics matter: ordering, duplicates, acknowledgement, retry, poison messages, transactions, routing, protocol compatibility and retention. Service substitution is complete only when producers/consumers and operations pass tests.
Replatforming can be valuable even if no microservice is created. Do not rename “containerized monolith” to “microservices.”
Serverless fit
Lambda and managed event services fit short-lived/event-driven work with variable demand and independent operations. Evaluate:
- runtime and package/image support;
- duration, memory, CPU, ephemeral storage and concurrency;
- event source semantics, batching, ordering and duplicate delivery;
- cold-start/latency and provisioned capacity needs;
- VPC connectivity, ENI behavior, downstream connection limits;
- retries, destinations/dead-letter queues and idempotency;
- state externalization and workflow duration;
- observability/correlation and local/integration testing; and
- per-request plus downstream cost at steady and burst load.
Do not split a long synchronous request into many functions solely to use serverless. Step Functions can make workflows explicit, but state-machine transitions, compensation and execution history become design/operational concerns.
Containers versus serverless versus modular monolith
| Requirement | Likely direction |
|---|---|
| existing long-running portable process, custom runtime or steady compute | managed containers |
| bursty event handler with bounded execution and managed integrations | Lambda/event services |
| unclear domain boundaries, small team, coordinated transactions | modular monolith on a managed platform |
| independent products/teams and scaling/failure needs with clear ownership | carefully bounded services |
This is not a service-selection shortcut. Validate latency, throughput, state, security, compliance, skills, lifecycle and total cost.
Find service boundaries from the domain
Useful decomposition views:
- business capability: what the organization does, such as catalog, ordering or billing;
- subdomain/bounded context: model and language that remain internally consistent;
- transaction: operations that must succeed atomically or compensate;
- change coupling: code/data that changes and deploys together;
- scaling/failure: components with materially different load or isolation needs; and
- service per team: stable ownership aligned to cognitive capacity.
A service should be independently deployable and own its behavior/data contract. Too small creates chatty calls, operational overhead and cross-service transactions. Too large preserves bottlenecks. Boundary discovery is iterative; start with seams where value and evidence are strongest.
Data ownership is the hard boundary
If several “microservices” directly update the same schema, releases and failures remain coupled. Establish one authoritative owner for a data domain and expose versioned APIs/events. Other services may maintain derived read models, not bypass ownership.
Database decomposition requires:
- table and column ownership map;
- cross-schema joins, stored code, triggers and reports;
- transaction/consistency requirements;
- historical/backfill and reconciliation design;
- access control and data classification per domain;
- analytics/reporting path that avoids operational ownership bypass;
- backup/restore and disaster consistency across stores; and
- retirement of old reads/writes.
Polyglot persistence is optional. Select different databases only when access patterns and team skills justify them; otherwise it multiplies operations.
Distributed transaction patterns
Local ACID transactions no longer span independently owned stores. Choose explicitly:
- Saga choreography: services react to events; simple for few participants but dependency/observability complexity grows.
- Saga orchestration: a coordinator directs steps; clearer complex workflow but the orchestrator needs resilience and ownership.
- Transactional outbox: business data and an outgoing event are committed together locally, then a publisher delivers the event.
- Idempotent consumer: duplicate messages produce one intended business effect.
- Compensation: semantically reverses prior actions; it may not restore time or erase external side effects.
Define business invariants, timeout, retry/backoff, ordering, deduplication key, poison-message handling, reconciliation and human repair. “Eventual consistency” is not permission to display impossible balances indefinitely.
Incremental modernization patterns
Strangler fig
Place a routing/proxy layer before the monolith. Initially send all traffic to legacy. Extract one capability, route its eligible traffic to the new service, compare behavior, expand, then remove old code/data only after no callers remain. It requires code/request interception and clear domain understanding.
Branch by abstraction
Create an abstraction around an internal dependency, implement old and new paths behind it, switch gradually and remove the legacy implementation after acceptance. Useful where external routing cannot intercept internal calls.
Anti-corruption layer
Translate between the legacy model/protocol and the new domain so old concepts do not leak into every new service. Own mappings, failures, versioning and eventual removal.
Use feature flags/canaries for controlled exposure, but maintain flag ownership and cleanup. A permanent compatibility layer becomes technical debt.
A safe strangler stage
For extracting Customer Preferences from an order monolith:
- baseline existing preference behavior, data and SLO;
- identify API/DB/batch/report callers and owner;
- define new service contract, data ownership and authorization;
- build service, IaC, pipeline, tests, observability, backup and on-call;
- backfill data with checksums and reconciliation;
- capture ongoing changes using outbox/CDC or controlled dual-write;
- initially shadow reads without affecting responses;
- route a deterministic canary cohort to new reads;
- enable writes with idempotency and rollback data strategy;
- expand only while business/technical metrics match;
- stop legacy writes, prove no callers, retain rollback window; and
- remove old code/tables/routes/flags and verify cost/security inventories.
Rollback differs before and after new authoritative writes. State which system owns truth at every stage.
API and event contracts
Modernization can shift coupling from code to network contracts. Define:
- schema and semantic versioning;
- backward/forward compatibility and deprecation window;
- authentication/authorization per caller and tenant;
- timeout, retry, circuit breaker and rate limit;
- pagination, idempotency and correlation IDs;
- event ordering, timestamp/source, duplicate behavior and retention;
- consumer contract/integration tests; and
- ownership, SLO and change communication.
Avoid long synchronous chains. One request calling six services serially compounds latency and failure. Use cached/materialized data or asynchronous workflows where business requirements permit.
Security changes with architecture
For every new component define workload identity, least-privilege permissions, network reachability, secrets/certificate lifecycle, encryption/KMS ownership, supply-chain controls, vulnerability patching, data classification, tenant isolation and audit evidence.
Containers need signed/scanned images, nonroot/minimal runtime, immutable deployment and controlled registry promotion. Functions need scoped execution roles, dependency scanning and concurrency protection. Events need resource policies and data minimization. Managed databases still need authentication, grants, parameters, backups and monitoring.
More services create more identities and policy relationships. Measure authorization paths; do not copy one broad role to every component.
Resilience and observability
Define service SLOs from the end-to-end user journey. Instrument metrics, structured logs and traces with correlation IDs across old/new paths. Observe deployment/configuration versions, queues, retries, circuit state, concurrency, dependency latency, saturation and business outcomes.
Test:
- instance/task/function replacement;
- AZ/dependency/identity/KMS/DNS failure;
- queue backlog and poison events;
- partial saga and compensation;
- database failover/restore;
- throttling and downstream exhaustion;
- rollback/roll-forward of code and schema; and
- loss of the legacy compatibility layer.
Adding retries without deadlines/backoff/idempotency can amplify an outage.
Deployment and database evolution
Independent deployment requires backward-compatible change:
- expand schema/API/event contract;
- deploy producers/consumers that support old and new;
- backfill and verify;
- switch reads/writes gradually;
- observe through rollback window; and
- contract/remove old fields only after all consumers prove migration.
Use blue/green, canary or feature flags according to data compatibility. Application rollback cannot undo destructive schema changes. Store deployment, migration and contract versions together in evidence.
Operating model and team readiness
Service autonomy requires “you build it, you run it” capabilities or an explicit alternative. Define product owner, development, platform, security, data and SRE responsibilities. Provide paved-road templates for repository, IaC, pipeline, identity, secrets, logging, alarms, SLOs, backup and cost tags.
Measure cognitive load. EKS, service mesh, multiple databases and event platforms may exceed a small team's capacity. A platform team should reduce complexity through supported products, not centralize every application change.
Economics and outcome measures
Include build/migration labor, parallel run, platform engineering, pipelines, observability, data migration, training, licenses, support, per-service resources, transfer, logs and increased nonproduction environments.
Measure outcomes such as:
- feature lead time and deployment frequency;
- change-failure rate and mean recovery time;
- availability/latency and scaling efficiency;
- security remediation and patch latency;
- operational tickets/toil and on-call burden;
- infrastructure/license/unit-transaction cost; and
- business conversion, retention or processing time.
Count services only as inventory, not success. Set target and review date. Stop or change direction if value does not justify complexity.
Troubleshooting modernization
| Symptom | Likely cause | Evidence-led response |
|---|---|---|
| More deploys but more incidents | weak tests/contracts/observability or team overload | inspect change failures and dependency paths; strengthen gates/platform |
| “Microservices” release together | shared schema, synchronous coupling, central ownership | map change/data coupling; merge or establish real boundaries |
| Event duplicates business action | no idempotency/deduplication | use business key, durable outcome and replay tests |
| Saga never completes | missing timeout/compensation/observability | inspect state/event timeline; add explicit recovery/human repair |
| Serverless throttles database | unbounded concurrency/connections | cap concurrency, pool/proxy/batch and protect downstream |
| Container cost exceeds VM | overprovisioning, fragmentation, platform overhead | calculate unit cost and right-size/bin-pack; reconsider platform |
| Strangler never retires legacy | no caller inventory/exit gate or compatibility layer became permanent | track routes/data/owners and fund deletion explicitly |
Hands-on workshop: modernize one monolith
Build a blueprint containing:
- five-lens readiness assessment and measurable outcomes;
- current component/dependency/data/transaction map;
- rehost, two replatform and three refactor options;
- weighted decision plus rejected choices;
- candidate business capabilities/bounded contexts and team owners;
- database/table ownership and distributed transaction treatment;
- container/serverless/managed-service fit by component;
- strangler route, anti-corruption layer and 12-stage extraction;
- API/event contracts, outbox/idempotency/saga design;
- security, SLO, telemetry, deployment and restore gates;
- target operating model, skills and full cost;
- rollback/write-authority and legacy decommission plan; and
- 30/90/180-day outcome review.
Inject failures: unclear domain boundary, two services write one table, duplicate payment event, downstream database connection exhaustion, canary latency regression, and an unowned compatibility layer. Decide whether to split, merge, defer, replatform or fix.
Knowledge check
- Is containerizing a monolith automatically refactoring?
No. It is commonly a replatform unless architecture/code/data ownership materially changes.
- Why might a modular monolith be correct?
Unclear domains, small teams or strong transaction coupling may benefit from one deployment while internal modules improve.
- Why is shared database write access dangerous?
It preserves schema/release/failure coupling and bypasses service ownership.
- What problem does a transactional outbox solve?
It atomically records business state and the intent to publish an event in one local transaction.
- What is the point of strangler fig?
Replace functionality incrementally through controlled routing, reducing big-bang risk.
- Why is compensation not simple rollback?
It performs a new business action and cannot always erase external side effects or elapsed time.
- How is modernization success measured?
Against agreed business, delivery, reliability, security, operations and unit-cost outcomes, not service count.
Lesson acceptance
Pass only when the blueprint links each code/data/platform/operating-model change to evidence and measurable value, with incremental stages and safe exit. Reject it if it assumes microservices are mandatory, calls containerization a completed refactor, allows shared-table writes, omits duplicate/partial workflow handling, uses a big-bang rewrite, lacks team/on-call ownership, treats managed services as operation-free, or cannot retire the legacy path.