AWS 323: Architecture trade-offs and service selection
Why this lesson matters
Architecture is the practice of choosing among acceptable compromises. Two designs can both work while differing in failure behavior, security ownership, delivery speed, cost, portability, skills, and reversibility. A service is not “best” in isolation; it is suitable for a specific workload and organization under explicit constraints.
The SAP-C02 exam tests this reasoning through scenarios. The correct answer normally satisfies hard requirements with the least unnecessary operational burden. In production, the same choice must survive review, implementation, incidents, cost scrutiny, and later change.
Learning outcomes
By the end, you can:
- turn approved requirements into knockout criteria and weighted criteria;
- create genuinely viable options at equivalent levels of detail;
- compare AWS services by responsibility, behavior, limits, and total cost;
- expose security, reliability, performance, sustainability, and operational trade-offs;
- use experiments to resolve uncertainty;
- test a decision against weight and assumption changes;
- document selection, rejected alternatives, risk, and reversal conditions;
- recognize common service-selection traps in exams and real reviews.
1. Requirements before scoring
Start from AWS322's accepted requirements. Separate:
- knockout criteria: legal, safety, security, residency, protocol, recovery, or service-availability conditions an option must satisfy;
- weighted criteria: factors where better and worse can be compared;
- assumptions: statements that require validation;
- preferences: negotiable team or vendor choices;
- future possibilities: useful context that must not dominate today's decision.
Never allow a high weighted score to compensate for a failed mandatory constraint. If a payment workload must keep regulated records in an approved geography, a lower-cost option outside it is not viable and should not enter the weighted ranking.
Express each criterion with a measurement method. “Low cost” becomes three-year total cost at baseline, expected, and peak demand. “Easy to operate” becomes named on-call skills, patching ownership, deployment steps, restore procedure, and incident evidence.
2. Build comparable options
Include at least three paths when realistic:
- retain or incrementally improve the current design;
- use a more managed AWS design;
- use an alternative architecture shape.
Do not create a favored option in deployment detail while describing alternatives in one sentence. For every option cover the same views:
| View | Questions |
|---|---|
| Request and data flow | Where do identity, traffic, state, and events travel? |
| Responsibility | What does AWS manage, and what does the team own? |
| Failure | Which component, AZ, Region, dependency, or operator action can fail? |
| Security | Where are trust, authorization, encryption, keys, and audit enforced? |
| Performance | What are latency, throughput, concurrency, scaling, and quota boundaries? |
| Operations | Who deploys, patches, observes, restores, and supports it? |
| Cost | What drives steady, variable, transfer, support, and people costs? |
| Change | How is it migrated, reversed, extended, or retired? |
A managed service transfers some undifferentiated work to AWS, but not accountability for data, configuration, identity, capacity limits, recovery objectives, or application correctness.
3. Service-selection ladder
Use this order instead of matching one keyword to one service:
- State the user and business outcome.
- Apply hard constraints.
- Identify workload shape: synchronous, asynchronous, streaming, batch, transactional, analytical, file, object, graph, time series, or mixed.
- Identify consistency, ordering, latency, throughput, durability, RTO, and RPO needs.
- Identify identity, network, residency, encryption, audit, and isolation boundaries.
- Determine operating capacity, skills, deployment model, and support hours.
- Compare viable service families and architecture shapes.
- Verify Regional availability, feature support, quotas, and integration behavior.
- Model full cost and growth.
- Resolve material unknowns through documentation, prototype, benchmark, or failure test.
Examples:
- Choose SQS when durable asynchronous buffering and consumer decoupling fit; choose EventBridge for event routing and integration patterns; choose Kinesis Data Streams when ordered high-throughput stream processing and shard/on-demand behavior fit. They are not interchangeable “messaging services.”
- Choose RDS or Aurora when relational semantics and managed operations fit; DynamoDB when access patterns support key-value/document behavior and predictable scale; Redshift for analytical warehousing. “Needs a database” is insufficient.
- Choose Lambda for event-driven bounded execution where its runtime, duration, concurrency, package, and networking model fit; containers or EC2 when workload/runtime/control requirements make them preferable.
4. Decision matrix without false precision
Create a matrix only after narrative analysis. Suggested criteria include security, reliability, performance, cost, sustainability, operability, delivery time, organizational skills, quota risk, portability, and reversibility.
Use a documented scoring scale, for example:
| Score | Meaning |
|---|---|
| 1 | Requirement is technically possible only with major unresolved risk |
| 2 | Significant compensating controls or organizational change required |
| 3 | Meets requirement with understood normal work |
| 4 | Meets requirement well with favorable evidence |
| 5 | Strongly exceeds the criterion with verified evidence |
For each score cite evidence and confidence. A total such as 3.74 is not scientific truth. The matrix makes assumptions and priorities visible; the decision still needs engineering judgment.
Sensitivity analysis
Recalculate when:
- the top criterion's weight increases or decreases;
- an uncertain score moves to its plausible minimum or maximum;
- demand is baseline, expected, and stress-case;
- a commitment discount is removed;
- one critical dependency or Region is unavailable;
- staffing or delivery-date assumptions change.
If tiny weight changes reverse the winner, the decision is fragile. Resolve important uncertainty or prefer the more reversible option.
5. Trade-offs across the six pillars
Security can add control boundaries and review time. Reliability often adds redundancy, testing, and cost. Performance can increase capacity or caching complexity. Cost optimization can reduce headroom. Sustainability can favor higher utilization and less retained data. Operational Excellence can favor managed services and automation. No pillar is a universal trump card; hard requirements and business impact determine acceptable balance.
Ask for second-order effects. A cache improves latency but needs correctness and invalidation. Cross-Region replication improves recovery but changes cost, consistency, key, residency, and failback. Event-driven decoupling absorbs spikes but introduces asynchronous user experience, retries, idempotency, observability, and eventual consistency.
6. Current evidence and safe queries
This is a no-create workshop. Read-only discovery can support, but cannot complete, a decision:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws ec2 describe-regions --query 'Regions[].RegionName' --output text
aws service-quotas list-services --max-results 30 --output table
aws wellarchitected list-workloads --output table
aws pricing describe-services --region us-east-1 --max-results 30 --output table
The Pricing API is called in us-east-1, returns large structured catalogs, and does not itself calculate architecture cost. Use AWS Pricing Calculator or governed cost models with exact Region, usage, data transfer, tiering, commitments, support, and growth assumptions. Verify current AWS documentation rather than assuming feature parity across Regions.
7. Experiment design
Run a proof of concept only for a decision-changing uncertainty. State hypothesis, representative workload, environment, variables, controls, success threshold, duration, evidence, owner, cost cap, and cleanup.
Useful experiments include:
- p95/p99 latency and error behavior at expected and burst load;
- failover, retry, duplicate, and recovery behavior;
- restore time and data-loss measurement;
- compatibility with required protocol, library, or data shape;
- operator deployment and incident-response time;
- quota and scale behavior;
- cost per business transaction.
A toy “hello world” proves API access, not production suitability.
8. Architecture decision record
The ADR should include status, context, requirements, options, knockout results, matrix, sensitivity, evidence, selected option, disadvantages, compensating controls, risks, assumptions, experiment results, migration, rollback, cost owner, and review triggers.
Classify decisions as one-way or two-way doors. Irreversible or expensive-to-reverse choices need stronger evidence. Reversible choices should still have guardrails. Triggers can include demand thresholds, feature availability, regulatory change, sustained cost variance, repeated incidents, or skill loss.
9. Guided workshop: order platform
Compare three architectures for a retailer processing orders in India and Singapore. Requirements: p95 API latency below 500 ms in-region, 2,000 normal and 8,000 burst requests/second, no duplicate charge, 99.95% monthly availability, one-hour RTO, five-minute RPO, payment-data controls, six-year order retention, delivery in six months, and an approved cost range. The team knows Linux, relational databases, and containers but has limited event-driven experience.
Produce:
- requirement and evidence register;
- knockout criteria;
- three equivalent option diagrams;
- responsibility matrix;
- request, event, and data flows;
- identity and trust analysis;
- failure and recovery analysis;
- consistency/idempotency decision;
- capacity and quota model;
- three-scenario cost model;
- operational ownership and skill-gap plan;
- weighted decision matrix with evidence;
- sensitivity analysis;
- two decision-changing experiments;
- risk and compensating-control register;
- complete ADR with reversal triggers.
10. Diagnosis and exam reasoning
Reject options that violate an explicit constraint before optimizing. Prefer managed services when they satisfy requirements and reduce unwanted operations, but do not assume managed means no operations. Watch for answers that add multi-Region, custom software, self-managed clusters, or data copies without a stated need. On the exam, “most operationally efficient” and “least cost” change the ranking only after mandatory requirements pass.
Common review failures are manipulated weights, missing retain/current option, ignored transfer cost, average instead of percentile performance, untested recovery, undocumented disadvantages, and a matrix filled with unsupported numbers.
Cost and cleanup
The lesson creates no resources. An approved future experiment must have a budget, tags, resource manifest, expiration, and dependency-ordered cleanup. Retain evidence and confirm that billing stops.
Knowledge check
- Can a weighted score rescue an option that violates residency? No; knockout criteria are mandatory.
- Why use sensitivity analysis? To reveal whether small weight or assumption changes reverse the decision.
- When is a POC useful? When it resolves a material uncertainty that can change selection.
- Why compare operational responsibility? Services transfer different work to AWS while leaving customer accountability.
- What makes an ADR reviewable? Explicit requirements, equivalent options, evidence, disadvantages, risks, and revisit triggers.
Lesson acceptance
Submit all 16 artifacts. Every viable option must be described at equal depth, hard constraints must be applied before scoring, every material score needs evidence and confidence, sensitivity must be tested, selected disadvantages must be explicit, and the ADR must include experiments, risks, rollback, owner, and review triggers.