AWS 313: Architecture: select analytics, streaming, AI, IoT, end-user, and specialized services without overengineering
Why this lesson matters
Architecture exams and real projects often present several services that could technically work. The best answer is not the service with the most features. It is the smallest supportable design that meets measurable requirements, protects data, survives expected failures, can be operated by the team, and has an exit path.
Overengineering appears as streaming for a daily report, a data lake for one clean table, a digital twin for a dashboard, blockchain for a trusted database, generative AI for deterministic lookup, Kubernetes for one batch job, or five overlapping observability systems. Underengineering is equally dangerous: one unpartitioned stream, a shared device certificate, an always-on endpoint with no demand, a desktop with uncontrolled export, or an AI agent with administrator tools.
This capstone integrates AWS300 through AWS312. You will select among analytics, streaming, search, ML/AI, IoT, end-user computing, messaging, media, and blockchain patterns while explicitly rejecting unnecessary components.
What you will be able to do
By the end, you can:
- translate a vague request into functional and quality requirements;
- separate facts, assumptions, constraints, preferences, and unknowns;
- choose batch, micro-batch, streaming, request/response, edge, desktop, or specialized processing;
- map service capability, operating responsibility, failure mode, quota, and cost;
- compare at least three viable options without feature-count bias;
- detect hidden complexity and irreversible commitments;
- design evidence, rollback, migration, and decommissioning before approval;
- reject AI, IoT, blockchain, real-time, or managed-specialty services when unjustified;
- defend a minimum viable architecture and a measured evolution path; and
- produce a complete P18 service-selection decision record.
Before you start
- Use the detailed lessons AWS300-AWS312 as technical references. This lesson does not replace their build and diagnostic work.
- Do not create resources. All scenarios use fictional organizations and synthetic data.
- Current service lifecycle status is a requirement: reject ended/unavailable paths identified in earlier lessons.
- Record uncertainty instead of inventing requirements. An assumption needs an owner and validation date.
- Avoid diagrams that show only service icons. Every arrow must name data, protocol, direction, volume, identity, and failure behavior.
1. Convert the request into measurable requirements
Start with an outcome statement:
For [user], provide [outcome] from [approved source]
within [latency/freshness], at [volume/peak],
with [availability/RTO/RPO], under [security/legal constraints],
measured by [acceptance evidence], owned by [team].
Classify every statement:
| Type | Example | Treatment |
|---|---|---|
| Fact | 10,000 devices currently send one 2 KB event/minute | Verify source and date. |
| Assumption | Traffic may double in 18 months | Model and assign validation owner. |
| Constraint | Data must remain in approved geography | Treat as gate, not preference. |
| Requirement | Dashboard reflects accepted events in 15 minutes | Test end-to-end. |
| Preference | Team likes Kafka | Score only after requirements. |
| Unknown | Maximum reconnect burst | Measure or design a bounded experiment. |
Capture data volume and shape, producers/consumers, latency/freshness, ordering, replay, delivery semantics, retention, consistency, query patterns, concurrency, languages/protocols, user locations, data class, identities, failure impact, RTO/RPO, skills, support model, quota, and budget.
“Real time,” “serverless,” “AI-powered,” and “enterprise scale” are not requirements. Replace them with numbers and observable behavior.
2. Choose the processing tempo
| Tempo | Use when | Typical direction |
|---|---|---|
| Request/response | User waits for one bounded result | API plus managed service/database |
| Scheduled batch | Inputs arrive together and delay is acceptable | S3, Glue/EMR, Athena, Batch Transform |
| Micro-batch | Minutes of freshness with efficient grouping | Scheduled/event batch over incremental data |
| Continuous stream | Each event must be processed with seconds/subseconds latency, state, or replay | Kinesis/MSK plus Flink/Lambda consumers |
| Edge/local | WAN loss or physical latency requires local behavior | Greengrass/device logic plus cloud sync |
| Human workflow | Decision requires qualified review | Queue/case UI and explicit approval |
Do not use a stream solely because data arrives continuously. If consumers need a daily aggregate, durable object ingestion plus batch may be simpler. Do not use batch when safety alerts need seconds and late handling.
3. Use a service-selection ladder
Evaluate in order:
- Can deterministic application logic solve it?
- Can an existing managed feature solve it?
- Does a purpose-built service fit task, data, Region, scale, and lifecycle?
- Is a configurable managed platform needed?
- Is custom infrastructure justified by a measured gap?
Move down only with evidence. The lowest layer transfers more patching, scaling, security, upgrade, and on-call responsibility to the team.
Analytics and data
- Athena plus Glue Catalog for serverless SQL over governed S3 data.
- Lake Formation when fine-grained lake permissions and cross-account governance are required.
- EMR when distributed-framework/runtime control or large custom Spark workloads justify it.
- Quick Sight for governed BI; Quick Suite adds broader research/agentic capabilities with new control needs.
- Data Exchange for licensed external datasets, not data truth.
Streaming, search, and events
- Kinesis Data Streams for AWS-managed ordered shard/on-demand streams.
- Amazon Data Firehose for managed buffered delivery/transformation.
- Managed Service for Apache Flink for stateful event-time stream processing.
- MSK when Kafka protocol/ecosystem/portability is an actual requirement.
- OpenSearch for search/log/vector analytics patterns, not primary transactional truth.
- SQS/SNS/EventBridge for queue, pub/sub, and event-routing semantics before selecting a streaming platform.
AI and ML
- deterministic rule or search before probabilistic generation;
- purpose-built AI API for a supported standard task;
- Bedrock for managed foundation-model applications;
- SageMaker AI for custom build/train/deploy/lifecycle control;
- human decision and deterministic authorization for high-impact effects.
IoT
- IoT Core for device identity/messaging/shadows/rules;
- Device Management/Defender for fleet operations/security;
- Greengrass V2 for supported edge workloads;
- SiteWise for industrial data modeling;
- TwinMaker for contextual operational twins only when that experience is needed.
End-user and specialist
- WorkSpaces Personal for persistent assigned desktops;
- WorkSpaces Applications for centrally streamed apps/non-persistent sessions;
- Amplify for appropriate web build/hosting/backend workflows;
- End User Messaging for transactional channels and Connect for campaigns/contact center;
- Media services only for actual codec, transport, packaging, origin, DRM, or ad workflows;
- Managed Blockchain only for genuine multi-party trust/ledger requirements.
4. Expose hidden operational ownership
For every candidate complete this table:
| Dimension | Required evidence |
|---|---|
| Identity | Human, workload, device, service roles and resource policy |
| Network | Every path, DNS, TLS, ports, endpoint, ingress/egress |
| Data | Source of truth, schema, quality, retention, deletion, residency |
| Runtime | Capacity, scaling, version, deployment, quotas |
| Failure | Detection, degraded mode, retry/replay, recovery, rollback |
| Security | Threats, least privilege, encryption, secrets, audit |
| Operations | Metrics/logs/traces, runbook, on-call, support boundary |
| Cost | Unit driver, normal/peak/failure volume, idle/commitment cost |
| Lifecycle | Service/model/software status, upgrades, exit/decommission |
A managed service removes selected infrastructure tasks. It does not remove data contracts, authorization, testing, quotas, observability, incident ownership, or cost governance.
5. Compare options fairly
Create at least three viable options. Establish knockout gates before weighted scoring. A design that violates residency, RTO, safety, or licensing cannot win by being cheap.
Example weighted score:
| Criterion | Weight | Measurement |
|---|---|---|
| Functional fit | 20 | Required behavior without unsupported workaround |
| Security/compliance | 20 | All mandatory controls and evidence |
| Reliability/recovery | 15 | Tested objectives and failure isolation |
| Operability/skills | 15 | Team can deploy, diagnose, upgrade, and support |
| Performance/scale | 10 | Normal, peak, backlog, and quota proof |
| Cost/TCO | 10 | Three-year demand scenarios and people cost |
| Change/exit | 10 | Migration, portability, rollback, decommission |
Score 0-5 with evidence and confidence. Sensitivity-test uncertain weights and volumes. Do not assign every preferred option 5.
6. Apply anti-overengineering tests
Reject or simplify a component when:
- no requirement maps uniquely to it;
- removing it does not break an acceptance test;
- it duplicates another system of record;
- the team cannot state its failure mode or owner;
- it exists only for hypothetical scale without a trigger;
- it adds a second synchronization or consistency boundary;
- a managed feature meets the need;
- its data must later be copied elsewhere merely to be usable;
- it is unavailable, closing, or unsupported for new work; or
- its rollback/decommission path is unknown.
Prefer reversible decisions. A queue can often be added before a consumer when measured bursts appear. A blockchain data model, proprietary campaign estate, on-chain personal record, or organization-wide identity choice can be expensive to reverse.
Define evolution triggers, not speculative components:
Start: S3 + Glue Catalog + Athena
Trigger: p95 query exceeds 20 seconds for 30 days at approved optimization
Then evaluate: materialized tables, Redshift, or another workload-specific engine
7. Prove the end-to-end path
For each critical flow, write:
producer identity
-> protocol/endpoint
-> authorization
-> buffer/storage
-> transformation
-> consumer/query
-> user/action
-> evidence
Then reverse it for return traffic, acknowledgement, and failure. Name partition keys, ordering scope, duplicate handling, schema version, late events, retries, poison data, backpressure, retention, encryption key, logs, and owner.
Capacity calculations must include:
- events/requests per second and payload units;
- partition/shard/broker/hot-key distribution;
- daily/monthly storage and compaction;
- consumer processing rate and backlog catch-up;
- reconnect/retry/failure burst;
- query concurrency and scan;
- model tokens or endpoint capacity;
- desktop/session peak and login burst; and
- quota headroom and increase lead time.
8. Diagnose poor architecture decisions
| Symptom | Likely design error | Correction |
|---|---|---|
| Daily report owns a 24x7 Kafka platform | Tempo and operational fit ignored | Reassess batch/object/serverless SQL. |
| Same data has three “sources of truth” | Services selected by feature, not ownership | Assign authority and derive read models. |
| AI answers are reviewed but still wrong | Human control lacks source/evidence/time | Improve task, retrieval, reviewer UI and evaluation. |
| IoT cloud outage stops local safety | Cloud service crossed physical safety boundary | Restore local interlock/safe state. |
| Desktop bill grows with no users | Running-mode/capacity owner absent | Use schedules/AutoStop/fleet scaling and lifecycle. |
| Search index is used as transaction database | Query feature confused with source of truth | Put authoritative writes in suitable database. |
| “Serverless” workload hits quota at launch | Capacity planning omitted | Load test, request quotas, throttle and degrade. |
| Blockchain cannot delete personal data | Technology rejected privacy requirement too late | Keep governed data off-chain or reject blockchain. |
9. Complete the P18 scenario
Fictional Northstar Manufacturing requests one “real-time AI platform” for:
- 10,000 refrigerator controllers sending minute telemetry and urgent alarms;
- plant OPC-UA data and OEE dashboards;
- daily executive sales/inventory BI;
- partner Kafka events at 5,000/s with seven-day replay;
- support document search and answer drafting;
- 300 contractors needing two managed Windows applications;
- order and OTP SMS;
- live quarterly video events; and
- three logistics partners requesting shared custody history.
Produce one portfolio architecture, but decide each workload independently. A defensible starting direction might include:
- IoT Core plus local safety and possibly Greengrass; SiteWise only for the industrial model;
- S3/Glue/Athena and Quick Sight for daily BI;
- MSK because the partner Kafka contract is explicit, with measured partition/replay design;
- authorized retrieval with Bedrock and mandatory support-agent review;
- WorkSpaces Applications for contractor apps;
- End User Messaging with consent/idempotency/spend controls;
- IVS or a media workflow selected from latency/production requirements, not a permanent broadcast stack by default; and
- database plus signed audit records unless consortium trust analysis proves Fabric necessary.
This is not an answer key. Requirements and evidence can produce another choice.
10. Required decision dossier
Submit:
- one-page business outcomes and actors;
- facts/assumptions/constraints/unknowns register;
- workload-by-workload measurable requirements;
- data classification, ownership, retention, and geography;
- processing-tempo decisions;
- three options per major workload;
- knockout gates and weighted matrix;
- explicit rejected services with reasons;
- end-to-end flow proofs;
- capacity/quota calculations;
- security and blast-radius map;
- failure/degraded/recovery/rollback table;
- operating model and RACI;
- normal/peak/failure/TCO cost;
- evolution triggers and decommission plan; and
- architecture review presentation and challenge log.
Run game days for eight cases: IoT reconnect storm, poisoned stream partition, stale lake permission, AI source injection, desktop capacity exhaustion, duplicate OTP trigger, video-origin failure, and one partner leaving the custody consortium.
Read-only evidence commands
aws sts get-caller-identity --query Arn --output text
aws service-quotas list-services --max-results 20 --output table
aws cloudwatch describe-alarms --max-records 20 --output table
aws resourcegroupstaggingapi get-resources --resources-per-page 20 --output table
aws ce get-cost-and-usage \
--time-period Start=2026-09-01,End=2026-09-02 \
--granularity DAILY --metrics UnblendedCost
These inventories prove limited control-plane state, not architecture fitness. Redact account and cost data.
Cost and cleanup
Build cost from service units and workload volume, then add logs, transfer, NAT, KMS, support, licensing, human review, idle minimums, commitments, and engineering/on-call effort. Include failure scenarios such as replay, retraining, reconnect bursts, and temporary dual running during migration.
T0 creates nothing. Preserve the decision record, evidence timestamps, source links, assumptions, approvals, and rejected options. Give every approved experiment a tagged owner, budget, expiry, cleanup command, and readback proof.
Knowledge check
- Why start with tempo? It eliminates architectures whose latency/processing model does not fit.
- What is a knockout gate? A mandatory requirement that scoring cannot compensate for.
- Why compare three viable options? To expose tradeoffs and avoid validating a predetermined favorite.
- Managed service versus no operations? It transfers selected tasks but leaves data, access, testing, quotas, incidents, and cost.
- When add speculative scale? Only after a measured trigger and reevaluation.
- Why prove every arrow? Service icons hide identity, data, protocol, direction, and failure.
- What makes rejection valuable? It prevents cost, coupling, risk, and unsupported lifecycle dependencies.
- What proves architecture success? User/business acceptance plus security, failure, operations, cost, and lifecycle evidence.
Lesson acceptance
Pass when every selected component maps to a requirement and every requirement maps to end-to-end acceptance evidence, with viable alternatives, quantified tradeoffs, owners, and exit paths. Fail if the design begins with service icons, calls every flow real-time, uses preference as fact, omits a simpler option, ignores current service lifecycle, cannot remove any component, or lacks rollback and cost.