Lesson 138 · AWS Learning Path

AWS 138: Architecture: select relational, key-value, document, graph, time-series, cache, and warehouse services

· Published · 11 min read

Labelled process diagram for AWS 138: Access patterns and data ownership to Candidate data models to Minimum justified database set to Failure, recovery, and cost review, with decision, proof and rejection evidence.

Why this lesson matters

Make a defensible database portfolio decision across relational, key-value, document, graph, time-series, cache, and warehouse needs.

Database selection begins with the workload, not an AWS product logo. The same application may need transactions, millisecond key lookups, flexible documents, relationship traversal, ephemeral cache entries and analytical scans - but assigning every need to a different service can create more failure modes than business value. This lesson turns requirements into a defensible minimum portfolio.

What you will be able to do

By the end, you can:

  • explain architecture: select relational, key-value, document, graph, time-series, cache, and warehouse services in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeMake a defensible database portfolio decision across relational, key-value, document, graph, time-series, cache, and warehouse needs.
Scope and boundaryOne application can use several data stores, but each adds identity, networking, consistency, backup, monitoring, skills, migration, and cost responsibilities.
Evidence of successThe final ADR maps each access pattern to one source of truth, defines integrations, rejects unnecessary stores, and includes failure, recovery, security, and cost evidence.
Cost modelThe cost model includes idle minimums, capacity or requests, storage, I/O, backups, replicas, transfer, caches, analytics scans, and operational effort.
Safe rejection ruleAvoid a service-per-feature design that duplicates truth, hides consistency, increases recovery coupling, and cannot be operated by the team.

How the request flows

+--------------------------------------+
|  Access patterns and data ownership  |
+--------------------------------------+
                   |
                   v
+-------------------------+
|  Candidate data models  |
+-------------------------+
            |
            v
+----------------------------------+
|  Minimum justified database set  |
+----------------------------------+
                 |
                 v
+--------------------------------------+
|  Failure, recovery, and cost review  |
+--------------------------------------+

For Database architecture decision, the important boundary is this: One application can use several data stores, but each adds identity, networking, consistency, backup, monitoring, skills, migration, and cost responsibilities. The final ADR maps each access pattern to one source of truth, defines integrations, rejects unnecessary stores, and includes failure, recovery, security, and cost evidence. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.

Architecture decision table

SituationDirectionReason
Requirement matchesUse polyglot persistence only where a specialized store produces measurable value and has an accountable owner.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid a service-per-feature design that duplicates truth, hides consistency, increases recovery coupling, and cannot be operated by the team.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

Start with access patterns and guarantees

For every data set, write the operation before naming a service: “fetch an order by ID,” “atomically reserve inventory,” “find paths up to three hops,” or “aggregate five years of events by month.” Record request rate, item/row size, growth, latency percentile, read/write ratio, query flexibility, transaction scope, consistency, retention, RPO, RTO, Regions and compliance. A noun such as “customer data” is not an access pattern.

Separate the system of record from derived views. Search indexes, caches, materialized reports and graph projections may be rebuilt from authoritative data. If two writable systems both claim to own the same field, define conflict resolution and failure handling or simplify the design.

Workload signalStrong starting candidateReject or qualify when
relational joins, constraints, SQL and multi-row transactionsRDS or Aurora with the compatible engineunbounded horizontal write scale or key-only access dominates; engine compatibility is not exact
known partition-key access at large scale and low latencyDynamoDBad-hoc joins/queries dominate, keys are unknown, or hot partitions cannot be redesigned
JSON document aggregate with document-oriented APIDocumentDB when its supported MongoDB compatibility is sufficientthe application depends on unsupported MongoDB behavior; validate compatibility, indexes and limits
relationship traversal and path questionsNeptune Database; Neptune Analytics for analytical graph workordinary foreign-key joins are enough or graph projection freshness is undefined
Apache Cassandra/CQL-compatible wide-column patternsKeyspacesCassandra feature compatibility or partition design does not fit; verify supported CQL behavior
high-resolution InfluxDB time-series workloadTimestream for InfluxDBoperational ownership, cardinality, retention or current service availability does not fit
ephemeral acceleration, sessions, counters or locksElastiCache for Valkey/Redis OSS/Memcachedit would become the only durable copy or stale/evicted data breaks correctness
columnar BI, large scans and warehouse SQLRedshift provisioned or Serverlessrequest-by-request OLTP or tiny operational lookups dominate
search, relevance, log exploration and aggregationsOpenSearch Serviceit is being treated as the transactional source of truth

No table chooses automatically. For example, a JSON column in PostgreSQL may be simpler than adding a document database; a relational adjacency table may handle a small graph; DynamoDB Streams may feed a Redshift analytical copy. The architect must show measurable reasons for specialization.

Make consistency and integration explicit

Within one database, use the strongest transaction boundary required by the invariant. Across databases, assume there is no universal ACID transaction. A common pattern is:

API transaction
  -> write authoritative business row/item + outbox record atomically
  -> stream/change-data-capture publisher
  -> durable event destination
  -> idempotent consumers update cache/search/graph/warehouse projections
  -> reconciliation finds missing, duplicate or divergent projections

Delivery can be at least once, out of order or delayed. Give every event an immutable ID, entity ID, version and occurred-at time. Consumers must detect duplicates and stale versions. Define what users see while a projection lags, how poison events reach a dead-letter path, how replay works, and how deletion/privacy requests propagate. “Eventual consistency” is not a design until the acceptable staleness and repair path are numeric.

Evaluate architecture dimensions

Score candidates from 1 (poor) to 5 (strong), but attach evidence to every score:

  • correctness: transactions, consistency, constraints, concurrency and conflict rules;
  • access fit: exact read/write/query patterns and indexes without scans or hot keys;
  • scale/performance: throughput, latency percentiles, item/row limits, partitions, connections and concurrency;
  • availability/recovery: Multi-AZ behavior, replicas, backup/PITR, cross-Region replication, tested RPO/RTO and failback;
  • security/governance: IAM/control plane, data-plane authentication, private networking, encryption, secrets, audit, residency and deletion;
  • operations: patching, schema/index changes, monitoring, capacity, troubleshooting, upgrades and team skill;
  • migration/reversibility: export format, change capture, cutover, rollback and vendor/API compatibility;
  • cost: idle floor plus requests/compute, storage, I/O, backup, replica, transfer, licensing and engineering effort.

Weight correctness and mandatory compliance as pass/fail gates rather than allowing cheap cost to cancel a failed requirement. Document assumptions such as growth and traffic and add a date to volatile pricing, quotas and feature support.

Worked portfolio example

An online shop has orders, catalog browsing, sessions, recommendations, telemetry and finance reporting. A defensible first iteration might be:

NeedDecisionOwnership and synchronization
orders, payments and inventory reservationPostgreSQL-compatible RDS/Aurorarelational system of record; one transaction protects business invariants
product catalogbegin in the relational store; consider DynamoDB only after measured key-access scale justifies itorder lines retain immutable product facts needed for history
sessions and hot product fragmentsElastiCache with bounded TTL; durable login/account truth remains outside cachecache miss reloads from the owner; eviction must not lose truth
recommendationsdo not add Neptune merely because relationships exist; adopt graph only if path/traversal queries show valuederived projection fed idempotently from durable events
telemetrystreaming/object-storage pipeline; choose a current time-series engine only for required interactive time-window queriesretention classes and deletion policy are explicit
finance and BIRedshift after volume/concurrency warrants a warehouse; start simpler when it does notCDC/batch projection reconciled to the order system

This design intentionally rejects several services at launch. Add a specialist store only when a benchmark, operational requirement or cost model shows why the existing owner cannot meet the need.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open the AWS database category and the Pricing Calculator estimate prepared in AWS122.
  2. Review RDS, DynamoDB, ElastiCache, Redshift, and purpose-built inventories without changing anything.
  3. Use each service's current documentation and limits page to validate the final decision assumptions.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws rds describe-db-instances --query 'DBInstances[].Engine' --output text
aws dynamodb list-tables --output text
aws elasticache describe-cache-clusters --query 'CacheClusters[].Engine' --output text
aws redshift describe-clusters --query 'Clusters[].ClusterIdentifier' --output text

Expected interpretation

The commands provide an ownership inventory, not an architecture answer. The ADR must begin with workload evidence and may correctly decide that no database resource should be added.

Practical work

Submit p06-database-adr.md for orders, catalog, sessions, recommendations, telemetry, cache and reporting. Include:

  1. workload assumptions and an access-pattern/guarantee table;
  2. one authoritative owner for every field and a diagram of derived copies;
  3. at least two candidates per workload, weighted scoring and explicit rejection reasons;
  4. normal data flow plus timeout, throttling, duplicate-event, stale-projection, Region-failure and restore flows;
  5. identity, network, encryption, secret, audit and data-residency boundaries;
  6. backup/PITR, RPO/RTO, restore validation, failover and failback ownership;
  7. capacity and five-year cost model with dates and sensitivity at 10x traffic/data;
  8. migration phases, dual-write avoidance, CDC/reconciliation, cutover gates and rollback;
  9. observability: service metrics, application SLIs, alarms, logs, traces and business reconciliation;
  10. a final minimum architecture and conditions that would trigger reconsideration.

Run three tabletop tests: loss of the cache, six-hour projection lag and loss of the primary database Region. For each, state user-visible behavior, preserved invariants, alarms, operator steps, recovery evidence and acceptable data loss. Also test a negative dependency: stop the projection consumer and prove the source-of-truth write remains correct while freshness alarms trigger.

Diagnose this topic from its own evidence

Evidence patternArchitectural defectCorrection
decision begins with “we want to use service X”technology-first selectionrestate measurable access patterns and score alternatives
same field is writable in two storesambiguous ownership and conflict risknominate one owner; make other copies derived or specify tested conflict semantics
synchronous request calls three databasescoupled latency and availabilityshrink the correctness path; move noncritical projections to durable asynchronous flow
scans, throttles or connection exhaustion at forecast loadmodel/capacity mismatchredesign keys/indexes/pooling, benchmark, or select a better-fit store
restore succeeds but application reconciliation failsbackup was tested below the business layervalidate schema, rows, permissions, downstream projections and application queries
cross-Region replica is called “zero data loss” without evidenceasynchronous RPO hiddenmeasure lag, define loss window and choose stronger controls if required
cache outage causes data losscache accidentally became system of recordpersist truth durably and design miss/rebuild behavior
cost estimate includes compute onlyincomplete TCOadd storage, I/O, backup, transfer, replicas, minimum capacity and operations

Cost and cleanup

The cost model includes idle minimums, capacity or requests, storage, I/O, backups, replicas, transfer, caches, analytics scans, and operational effort.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Make a defensible database portfolio decision across relational, key-value, document, graph, time-series, cache, and warehouse needs.

  1. Which scope or ownership boundary must be proved first?

Expected direction: One application can use several data stores, but each adds identity, networking, consistency, backup, monitoring, skills, migration, and cost responsibilities.

  1. What evidence is strong enough to accept the result?

Expected direction: The final ADR maps each access pattern to one source of truth, defines integrations, rejects unnecessary stores, and includes failure, recovery, security, and cost evidence.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid a service-per-feature design that duplicates truth, hides consistency, increases recovery coupling, and cannot be operated by the team.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: The cost model includes idle minimums, capacity or requests, storage, I/O, backups, replicas, transfer, caches, analytics scans, and operational effort.

Lesson acceptance

Pass only when another architect can trace every requirement to evidence and reproduce the decision. The ADR must identify authoritative data owners, transaction and consistency boundaries, integrations and replay, security, availability, tested recovery, observability, migration/rollback and full cost dimensions. It must reject at least two plausible services with technical reasons, include the three tabletop failures, and avoid unsupported claims such as “managed means no operations” or “multi-Region means no data loss.”

Fail the lesson if it is merely a service comparison table, duplicates write ownership, omits degraded behavior, assumes backups equal recovery, cites undated pricing/features, or introduces a database without an accountable operator and deletion/retention policy.

Official sources

Advertisement