AWS 328: Microservices and component decoupling
Why this lesson matters
Microservices are independently owned and deployable business capabilities communicating across a network. They can improve team autonomy, scaling, release speed, and failure isolation, but replace in-process calls and transactions with network latency, partial failure, duplicated messages, versioned contracts, distributed data, and harder operations.
Decoupling is the goal only when it improves a measured outcome. A well-structured modular monolith is often safer than premature microservices. Service boundaries should follow business cohesion and ownership, not tables, classes, or fashionable AWS icons.
Learning outcomes
By the end, you can:
- decide between modular monolith and microservices;
- find bounded contexts and service ownership boundaries;
- choose synchronous APIs, queues, events, or workflows by behavior;
- design contracts, versioning, idempotency, and backpressure;
- preserve consistency through local transactions, sagas, and reconciliation;
- apply strangler, branch-by-abstraction, and anti-corruption patterns;
- design observability, security, delivery, and rollback across services;
- produce a staged decomposition of an order monolith.
1. When to split and when to stay together
Use evidence such as independent change cadence, distinct scaling, failure isolation, domain ownership, security boundary, and team autonomy. Do not split merely because the codebase is large.
| Signal | Keep or improve modular monolith | Consider service boundary |
|---|---|---|
| Domain | Rules are tightly cohesive and evolving together | Stable bounded capability with clear language |
| Team | One small team owns and releases it | Durable team can own build and operations |
| Scale | Components scale similarly | One capability has materially different demand |
| Failure | In-process transaction is valuable | Isolation/degradation has business value |
| Delivery | Coordinated release is acceptable | Independent release solves measured delay |
| Data | Strong cross-component transactions dominate | Data ownership and eventual consistency are acceptable |
Microservices without independent ownership become a distributed monolith: many deployables that still require coordinated release.
2. Find boundaries
Map user journeys, business capabilities, domain language, commands, events, data ownership, policies, and change history. High cohesion means related rules change together; low coupling means a boundary needs little knowledge of others.
Useful decomposition perspectives include business capability, subdomain, transaction, and service per team. Avoid one service per database table or CRUD entity. A service should own behavior and its data, not expose its schema as a remote database.
Start with seams: stable interfaces, painful release boundaries, independently scaling jobs, or capabilities with clear owners. Record reasons to keep other areas together.
3. Interaction choices
Synchronous API
Use when the caller needs an immediate answer. Define protocol, authentication, authorization, timeout, retry eligibility, idempotency, version, quota, and error contract. Long call chains compound latency and availability. Use degradation, caching, or asynchronous continuation where business rules allow.
Queue
Use for durable work buffering, load leveling, and consumer decoupling. With SQS define Standard versus FIFO need, visibility timeout, retention, receive count, dead-letter queue, redrive, ordering scope, deduplication, and backpressure. A DLQ is not an archive; it needs alarms, diagnosis, replay safety, and ownership.
Event
An event states a fact that occurred. EventBridge or SNS can fan out to interested consumers. Producers should not depend on each consumer's implementation. Define source, type, version, subject, occurrence time, correlation, idempotency key, data classification, retention/replay approach, and schema compatibility.
Workflow
Use Step Functions or another explicit orchestrator when sequence, branching, waiting, timeout, retries, compensation, and audit must be visible. Choreography can be simple for few participants, but becomes hard to understand as hidden dependencies grow.
4. Contract design
Keep contracts consumer-focused and backward compatible. Add optional fields before removing or changing semantics. Do not reuse a field for a different meaning. Maintain schema ownership, examples, validation, compatibility policy, deprecation notice, consumer inventory, and sunset evidence.
Events should be self-contained enough for the intended consumer without carrying large sensitive payloads. Use a claim-check pattern: store a large payload securely and send an authorized reference with integrity and lifecycle controls.
Contract tests complement, not replace, end-to-end tests. Test old consumer/new producer and new consumer/old producer during rolling change.
5. Delivery semantics and idempotency
Assume messages can be delivered more than once. Exactly-once business behavior is achieved with idempotency and state, even when transport offers deduplication features.
An idempotency record includes operation key, request fingerprint, status, result/reference, creation, and expiry. Claim the key atomically before side effects. Decide behavior for same key/different payload, in-progress work, expired records, and partial failure. Retention must exceed realistic retry/replay duration.
Retries need bounded attempts, exponential backoff, jitter, and retryable-error classification. Retrying validation errors or permanent authorization failures wastes capacity. Ensure visibility timeout exceeds normal processing with safe extensions for long work.
6. Data ownership and consistency
Each service owns writes to its datastore. Other services use contracts or replicated views. Shared physical infrastructure can be transitional, but shared writable schemas couple deployment, authorization, and failure.
Cross-service workflows use local transactions and consistency patterns:
- transactional outbox: write business state and outgoing event in one local transaction, then relay it;
- saga choreography: participants react to events and publish outcomes;
- saga orchestration: coordinator commands steps and tracks state;
- compensation: perform a business action that addresses a prior committed action;
- reconciliation: compare authoritative records and repair divergence.
Compensation is not database rollback. Refunding a captured payment creates another auditable transaction. Some effects, such as email or shipment, cannot be undone; design prevention or remediation.
Avoid distributed two-phase commit across services unless the platform and requirements genuinely justify its availability and coupling trade-offs.
7. Failure isolation and backpressure
Set timeout shorter than the caller's remaining budget. Use circuit breakers for repeated dependency failure, bulkheads for tenant/function isolation, queues for buffering, concurrency limits to protect downstream systems, and load shedding for noncritical work.
Trace a retry storm: caller retries, gateway retries, service retries, and SDK retries can multiply one failure into many requests. Assign retry ownership to one suitable layer and measure attempts per successful business outcome.
Design poison-message handling, partial batch response where supported, replay rate, quarantine, and manual approval for side-effecting replay.
8. AWS service mapping
API Gateway or load balancers can expose APIs; Lambda, ECS, EKS, or EC2 can run components; SQS buffers work; SNS fans out messages; EventBridge routes events; Step Functions orchestrates; DynamoDB, Aurora/RDS, and other stores serve owned access patterns. Choose from requirements and team capability, not from a predetermined “microservices stack.”
Read-only inventory:
export AWS_DEFAULT_REGION="ap-south-1"
aws apigatewayv2 get-apis --output table
aws events list-event-buses --output table
aws sqs list-queues --output table
aws stepfunctions list-state-machines --output table
aws ecs list-clusters --output table
Inventory does not prove contracts, ownership, traceability, or safe replay.
9. Security and observability
Authenticate service identity and authorize least privilege at each boundary. Protect resource policies, execution roles, queue/topic/event-bus policies, secrets, KMS keys, egress, and tenant context. Validate messages; an internal event is not automatically trustworthy.
Propagate correlation and trace context without leaking sensitive data. Observe request success/latency, queue age/depth, event delivery, workflow failure, retries, throttles, DLQ, idempotency conflict, compensation, and business completion. Build a service map from telemetry, not only documentation.
Each service needs owner, SLO, on-call route, dashboard, alarms, runbook, dependency list, quota model, cost allocation, deployment process, and recovery plan.
10. Incremental decomposition
The strangler fig pattern routes selected capabilities to a new implementation while others remain in the monolith. Branch by abstraction introduces an internal interface so implementation can change behind it. An anti-corruption layer translates between old and new domain models.
Stages: characterize current behavior; add observability; establish seam; build new owned data and contract; shadow or test; route a small cohort; reconcile; increase traffic; stop old writes; validate; remove old code/data after retention and rollback decisions.
Dual writes create divergence unless one atomic source plus outbox/CDC and reconciliation is designed. Keep a declared source of truth at every stage.
11. Guided workshop: order monolith
A monolith contains catalog, cart, order, payment, inventory, shipping, notifications, and reporting in one relational database. Promotions overload inventory queries; payment changes require whole-system releases; shipping provider outages block checkout; one team owns everything.
Produce:
- user journeys and current coupling map;
- domain/capability model;
- team and operational ownership proposal;
- modular-monolith versus service decision for each boundary;
- one justified first extraction and one explicit keep-together choice;
- synchronous API contracts and timeout budgets;
- event catalog and compatibility rules;
- queue, retry, backpressure, DLQ, and replay design;
- idempotency state machine for payment;
- data ownership and replicated-view map;
- outbox design;
- order saga with compensation and reconciliation;
- trust and least-privilege map;
- service-level telemetry and tracing model;
- deployment/contract-test strategy;
- strangler stages, traffic shift, rollback, and old-path removal;
- failure-injection plan;
- cost, quota, risk, and acceptance record.
12. Troubleshooting
| Symptom | Investigate |
|---|---|
| Duplicate charge | Idempotency claim, key scope/expiry, retries, side-effect ordering |
| Queue age rises | Producer rate, consumer concurrency, duration, downstream throttle |
| DLQ fills after release | Contract compatibility, poison item, IAM/KMS, timeout |
| Saga stuck | Missing event, correlation, state transition, compensation failure |
| Latency varies | Synchronous chain, cold path, connection, dependency percentile |
| Services release together | Shared schema/library, contract break, team ownership |
Cost and cleanup
This no-create lesson has no AWS cleanup. Model API requests, messages, workflow transitions, compute, state, logs/traces, NAT/transfer, duplicate work, environments, and team/on-call cost. Microservices often increase platform and observability cost.
Knowledge check
- Is one service per table a good boundary? No; boundaries follow cohesive business capability and ownership.
- Why is idempotency still needed with retries? The same work can be delivered or attempted more than once.
- What does an outbox solve? Atomic business-state and event intent recording in one local transaction.
- Is compensation rollback? No; it is a new business action addressing an earlier committed action.
- When is a modular monolith preferable? When distributed complexity exceeds measured autonomy, scaling, or isolation benefit.
Lesson acceptance
Submit all 18 artifacts. Boundaries must be justified by domain and ownership evidence, every interaction must define failure and delivery behavior, data writes must have one owner, retries must be idempotent and bounded, and extraction must preserve rollback/reconciliation and observable business outcomes.