Lesson 313 · AWS Learning Path

AWS 313: Architecture: select analytics, streaming, AI, IoT, end-user, and specialized services without overengineering

· Published · 10 min read

Labelled process diagram for AWS 313: Business and data requirements to Smallest suitable current service set to Integrated workload path to Outcome, operations, cost, and reversal evidence, with decision, proof and...

Why this lesson matters

Architecture exams and real projects often present several services that could technically work. The best answer is not the service with the most features. It is the smallest supportable design that meets measurable requirements, protects data, survives expected failures, can be operated by the team, and has an exit path.

Overengineering appears as streaming for a daily report, a data lake for one clean table, a digital twin for a dashboard, blockchain for a trusted database, generative AI for deterministic lookup, Kubernetes for one batch job, or five overlapping observability systems. Underengineering is equally dangerous: one unpartitioned stream, a shared device certificate, an always-on endpoint with no demand, a desktop with uncontrolled export, or an AI agent with administrator tools.

This capstone integrates AWS300 through AWS312. You will select among analytics, streaming, search, ML/AI, IoT, end-user computing, messaging, media, and blockchain patterns while explicitly rejecting unnecessary components.

What you will be able to do

By the end, you can:

  • translate a vague request into functional and quality requirements;
  • separate facts, assumptions, constraints, preferences, and unknowns;
  • choose batch, micro-batch, streaming, request/response, edge, desktop, or specialized processing;
  • map service capability, operating responsibility, failure mode, quota, and cost;
  • compare at least three viable options without feature-count bias;
  • detect hidden complexity and irreversible commitments;
  • design evidence, rollback, migration, and decommissioning before approval;
  • reject AI, IoT, blockchain, real-time, or managed-specialty services when unjustified;
  • defend a minimum viable architecture and a measured evolution path; and
  • produce a complete P18 service-selection decision record.

Before you start

  • Use the detailed lessons AWS300-AWS312 as technical references. This lesson does not replace their build and diagnostic work.
  • Do not create resources. All scenarios use fictional organizations and synthetic data.
  • Current service lifecycle status is a requirement: reject ended/unavailable paths identified in earlier lessons.
  • Record uncertainty instead of inventing requirements. An assumption needs an owner and validation date.
  • Avoid diagrams that show only service icons. Every arrow must name data, protocol, direction, volume, identity, and failure behavior.

1. Convert the request into measurable requirements

Start with an outcome statement:

For [user], provide [outcome] from [approved source]
within [latency/freshness], at [volume/peak],
with [availability/RTO/RPO], under [security/legal constraints],
measured by [acceptance evidence], owned by [team].

Classify every statement:

TypeExampleTreatment
Fact10,000 devices currently send one 2 KB event/minuteVerify source and date.
AssumptionTraffic may double in 18 monthsModel and assign validation owner.
ConstraintData must remain in approved geographyTreat as gate, not preference.
RequirementDashboard reflects accepted events in 15 minutesTest end-to-end.
PreferenceTeam likes KafkaScore only after requirements.
UnknownMaximum reconnect burstMeasure or design a bounded experiment.

Capture data volume and shape, producers/consumers, latency/freshness, ordering, replay, delivery semantics, retention, consistency, query patterns, concurrency, languages/protocols, user locations, data class, identities, failure impact, RTO/RPO, skills, support model, quota, and budget.

“Real time,” “serverless,” “AI-powered,” and “enterprise scale” are not requirements. Replace them with numbers and observable behavior.

2. Choose the processing tempo

TempoUse whenTypical direction
Request/responseUser waits for one bounded resultAPI plus managed service/database
Scheduled batchInputs arrive together and delay is acceptableS3, Glue/EMR, Athena, Batch Transform
Micro-batchMinutes of freshness with efficient groupingScheduled/event batch over incremental data
Continuous streamEach event must be processed with seconds/subseconds latency, state, or replayKinesis/MSK plus Flink/Lambda consumers
Edge/localWAN loss or physical latency requires local behaviorGreengrass/device logic plus cloud sync
Human workflowDecision requires qualified reviewQueue/case UI and explicit approval

Do not use a stream solely because data arrives continuously. If consumers need a daily aggregate, durable object ingestion plus batch may be simpler. Do not use batch when safety alerts need seconds and late handling.

3. Use a service-selection ladder

Evaluate in order:

  1. Can deterministic application logic solve it?
  2. Can an existing managed feature solve it?
  3. Does a purpose-built service fit task, data, Region, scale, and lifecycle?
  4. Is a configurable managed platform needed?
  5. Is custom infrastructure justified by a measured gap?

Move down only with evidence. The lowest layer transfers more patching, scaling, security, upgrade, and on-call responsibility to the team.

Analytics and data

  • Athena plus Glue Catalog for serverless SQL over governed S3 data.
  • Lake Formation when fine-grained lake permissions and cross-account governance are required.
  • EMR when distributed-framework/runtime control or large custom Spark workloads justify it.
  • Quick Sight for governed BI; Quick Suite adds broader research/agentic capabilities with new control needs.
  • Data Exchange for licensed external datasets, not data truth.

Streaming, search, and events

  • Kinesis Data Streams for AWS-managed ordered shard/on-demand streams.
  • Amazon Data Firehose for managed buffered delivery/transformation.
  • Managed Service for Apache Flink for stateful event-time stream processing.
  • MSK when Kafka protocol/ecosystem/portability is an actual requirement.
  • OpenSearch for search/log/vector analytics patterns, not primary transactional truth.
  • SQS/SNS/EventBridge for queue, pub/sub, and event-routing semantics before selecting a streaming platform.

AI and ML

  • deterministic rule or search before probabilistic generation;
  • purpose-built AI API for a supported standard task;
  • Bedrock for managed foundation-model applications;
  • SageMaker AI for custom build/train/deploy/lifecycle control;
  • human decision and deterministic authorization for high-impact effects.

IoT

  • IoT Core for device identity/messaging/shadows/rules;
  • Device Management/Defender for fleet operations/security;
  • Greengrass V2 for supported edge workloads;
  • SiteWise for industrial data modeling;
  • TwinMaker for contextual operational twins only when that experience is needed.

End-user and specialist

  • WorkSpaces Personal for persistent assigned desktops;
  • WorkSpaces Applications for centrally streamed apps/non-persistent sessions;
  • Amplify for appropriate web build/hosting/backend workflows;
  • End User Messaging for transactional channels and Connect for campaigns/contact center;
  • Media services only for actual codec, transport, packaging, origin, DRM, or ad workflows;
  • Managed Blockchain only for genuine multi-party trust/ledger requirements.

4. Expose hidden operational ownership

For every candidate complete this table:

DimensionRequired evidence
IdentityHuman, workload, device, service roles and resource policy
NetworkEvery path, DNS, TLS, ports, endpoint, ingress/egress
DataSource of truth, schema, quality, retention, deletion, residency
RuntimeCapacity, scaling, version, deployment, quotas
FailureDetection, degraded mode, retry/replay, recovery, rollback
SecurityThreats, least privilege, encryption, secrets, audit
OperationsMetrics/logs/traces, runbook, on-call, support boundary
CostUnit driver, normal/peak/failure volume, idle/commitment cost
LifecycleService/model/software status, upgrades, exit/decommission

A managed service removes selected infrastructure tasks. It does not remove data contracts, authorization, testing, quotas, observability, incident ownership, or cost governance.

5. Compare options fairly

Create at least three viable options. Establish knockout gates before weighted scoring. A design that violates residency, RTO, safety, or licensing cannot win by being cheap.

Example weighted score:

CriterionWeightMeasurement
Functional fit20Required behavior without unsupported workaround
Security/compliance20All mandatory controls and evidence
Reliability/recovery15Tested objectives and failure isolation
Operability/skills15Team can deploy, diagnose, upgrade, and support
Performance/scale10Normal, peak, backlog, and quota proof
Cost/TCO10Three-year demand scenarios and people cost
Change/exit10Migration, portability, rollback, decommission

Score 0-5 with evidence and confidence. Sensitivity-test uncertain weights and volumes. Do not assign every preferred option 5.

6. Apply anti-overengineering tests

Reject or simplify a component when:

  • no requirement maps uniquely to it;
  • removing it does not break an acceptance test;
  • it duplicates another system of record;
  • the team cannot state its failure mode or owner;
  • it exists only for hypothetical scale without a trigger;
  • it adds a second synchronization or consistency boundary;
  • a managed feature meets the need;
  • its data must later be copied elsewhere merely to be usable;
  • it is unavailable, closing, or unsupported for new work; or
  • its rollback/decommission path is unknown.

Prefer reversible decisions. A queue can often be added before a consumer when measured bursts appear. A blockchain data model, proprietary campaign estate, on-chain personal record, or organization-wide identity choice can be expensive to reverse.

Define evolution triggers, not speculative components:

Start: S3 + Glue Catalog + Athena
Trigger: p95 query exceeds 20 seconds for 30 days at approved optimization
Then evaluate: materialized tables, Redshift, or another workload-specific engine

7. Prove the end-to-end path

For each critical flow, write:

producer identity
 -> protocol/endpoint
 -> authorization
 -> buffer/storage
 -> transformation
 -> consumer/query
 -> user/action
 -> evidence

Then reverse it for return traffic, acknowledgement, and failure. Name partition keys, ordering scope, duplicate handling, schema version, late events, retries, poison data, backpressure, retention, encryption key, logs, and owner.

Capacity calculations must include:

  • events/requests per second and payload units;
  • partition/shard/broker/hot-key distribution;
  • daily/monthly storage and compaction;
  • consumer processing rate and backlog catch-up;
  • reconnect/retry/failure burst;
  • query concurrency and scan;
  • model tokens or endpoint capacity;
  • desktop/session peak and login burst; and
  • quota headroom and increase lead time.

8. Diagnose poor architecture decisions

SymptomLikely design errorCorrection
Daily report owns a 24x7 Kafka platformTempo and operational fit ignoredReassess batch/object/serverless SQL.
Same data has three “sources of truth”Services selected by feature, not ownershipAssign authority and derive read models.
AI answers are reviewed but still wrongHuman control lacks source/evidence/timeImprove task, retrieval, reviewer UI and evaluation.
IoT cloud outage stops local safetyCloud service crossed physical safety boundaryRestore local interlock/safe state.
Desktop bill grows with no usersRunning-mode/capacity owner absentUse schedules/AutoStop/fleet scaling and lifecycle.
Search index is used as transaction databaseQuery feature confused with source of truthPut authoritative writes in suitable database.
“Serverless” workload hits quota at launchCapacity planning omittedLoad test, request quotas, throttle and degrade.
Blockchain cannot delete personal dataTechnology rejected privacy requirement too lateKeep governed data off-chain or reject blockchain.

9. Complete the P18 scenario

Fictional Northstar Manufacturing requests one “real-time AI platform” for:

  • 10,000 refrigerator controllers sending minute telemetry and urgent alarms;
  • plant OPC-UA data and OEE dashboards;
  • daily executive sales/inventory BI;
  • partner Kafka events at 5,000/s with seven-day replay;
  • support document search and answer drafting;
  • 300 contractors needing two managed Windows applications;
  • order and OTP SMS;
  • live quarterly video events; and
  • three logistics partners requesting shared custody history.

Produce one portfolio architecture, but decide each workload independently. A defensible starting direction might include:

  • IoT Core plus local safety and possibly Greengrass; SiteWise only for the industrial model;
  • S3/Glue/Athena and Quick Sight for daily BI;
  • MSK because the partner Kafka contract is explicit, with measured partition/replay design;
  • authorized retrieval with Bedrock and mandatory support-agent review;
  • WorkSpaces Applications for contractor apps;
  • End User Messaging with consent/idempotency/spend controls;
  • IVS or a media workflow selected from latency/production requirements, not a permanent broadcast stack by default; and
  • database plus signed audit records unless consortium trust analysis proves Fabric necessary.

This is not an answer key. Requirements and evidence can produce another choice.

10. Required decision dossier

Submit:

  1. one-page business outcomes and actors;
  2. facts/assumptions/constraints/unknowns register;
  3. workload-by-workload measurable requirements;
  4. data classification, ownership, retention, and geography;
  5. processing-tempo decisions;
  6. three options per major workload;
  7. knockout gates and weighted matrix;
  8. explicit rejected services with reasons;
  9. end-to-end flow proofs;
  10. capacity/quota calculations;
  11. security and blast-radius map;
  12. failure/degraded/recovery/rollback table;
  13. operating model and RACI;
  14. normal/peak/failure/TCO cost;
  15. evolution triggers and decommission plan; and
  16. architecture review presentation and challenge log.

Run game days for eight cases: IoT reconnect storm, poisoned stream partition, stale lake permission, AI source injection, desktop capacity exhaustion, duplicate OTP trigger, video-origin failure, and one partner leaving the custody consortium.

Read-only evidence commands

aws sts get-caller-identity --query Arn --output text
aws service-quotas list-services --max-results 20 --output table
aws cloudwatch describe-alarms --max-records 20 --output table
aws resourcegroupstaggingapi get-resources --resources-per-page 20 --output table
aws ce get-cost-and-usage \
  --time-period Start=2026-09-01,End=2026-09-02 \
  --granularity DAILY --metrics UnblendedCost

These inventories prove limited control-plane state, not architecture fitness. Redact account and cost data.

Cost and cleanup

Build cost from service units and workload volume, then add logs, transfer, NAT, KMS, support, licensing, human review, idle minimums, commitments, and engineering/on-call effort. Include failure scenarios such as replay, retraining, reconnect bursts, and temporary dual running during migration.

T0 creates nothing. Preserve the decision record, evidence timestamps, source links, assumptions, approvals, and rejected options. Give every approved experiment a tagged owner, budget, expiry, cleanup command, and readback proof.

Knowledge check

  1. Why start with tempo? It eliminates architectures whose latency/processing model does not fit.
  2. What is a knockout gate? A mandatory requirement that scoring cannot compensate for.
  3. Why compare three viable options? To expose tradeoffs and avoid validating a predetermined favorite.
  4. Managed service versus no operations? It transfers selected tasks but leaves data, access, testing, quotas, incidents, and cost.
  5. When add speculative scale? Only after a measured trigger and reevaluation.
  6. Why prove every arrow? Service icons hide identity, data, protocol, direction, and failure.
  7. What makes rejection valuable? It prevents cost, coupling, risk, and unsupported lifecycle dependencies.
  8. What proves architecture success? User/business acceptance plus security, failure, operations, cost, and lifecycle evidence.

Lesson acceptance

Pass when every selected component maps to a requirement and every requirement maps to end-to-end acceptance evidence, with viable alternatives, quantified tradeoffs, owners, and exit paths. Fail if the design begins with service icons, calls every flow real-time, uses preference as fact, omits a simpler option, ignores current service lifecycle, cannot remove any component, or lacks rollback and cost.

Official sources

Advertisement