AWS 138: Architecture: select relational, key-value, document, graph, time-series, cache, and warehouse services
Why this lesson matters
Make a defensible database portfolio decision across relational, key-value, document, graph, time-series, cache, and warehouse needs.
Database selection begins with the workload, not an AWS product logo. The same application may need transactions, millisecond key lookups, flexible documents, relationship traversal, ephemeral cache entries and analytical scans - but assigning every need to a different service can create more failure modes than business value. This lesson turns requirements into a defensible minimum portfolio.
What you will be able to do
By the end, you can:
- explain architecture: select relational, key-value, document, graph, time-series, cache, and warehouse services in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Make a defensible database portfolio decision across relational, key-value, document, graph, time-series, cache, and warehouse needs. |
| Scope and boundary | One application can use several data stores, but each adds identity, networking, consistency, backup, monitoring, skills, migration, and cost responsibilities. |
| Evidence of success | The final ADR maps each access pattern to one source of truth, defines integrations, rejects unnecessary stores, and includes failure, recovery, security, and cost evidence. |
| Cost model | The cost model includes idle minimums, capacity or requests, storage, I/O, backups, replicas, transfer, caches, analytics scans, and operational effort. |
| Safe rejection rule | Avoid a service-per-feature design that duplicates truth, hides consistency, increases recovery coupling, and cannot be operated by the team. |
How the request flows
+--------------------------------------+
| Access patterns and data ownership |
+--------------------------------------+
|
v
+-------------------------+
| Candidate data models |
+-------------------------+
|
v
+----------------------------------+
| Minimum justified database set |
+----------------------------------+
|
v
+--------------------------------------+
| Failure, recovery, and cost review |
+--------------------------------------+
For Database architecture decision, the important boundary is this: One application can use several data stores, but each adds identity, networking, consistency, backup, monitoring, skills, migration, and cost responsibilities. The final ADR maps each access pattern to one source of truth, defines integrations, rejects unnecessary stores, and includes failure, recovery, security, and cost evidence. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use polyglot persistence only where a specialized store produces measurable value and has an accountable owner. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid a service-per-feature design that duplicates truth, hides consistency, increases recovery coupling, and cannot be operated by the team. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
Start with access patterns and guarantees
For every data set, write the operation before naming a service: “fetch an order by ID,” “atomically reserve inventory,” “find paths up to three hops,” or “aggregate five years of events by month.” Record request rate, item/row size, growth, latency percentile, read/write ratio, query flexibility, transaction scope, consistency, retention, RPO, RTO, Regions and compliance. A noun such as “customer data” is not an access pattern.
Separate the system of record from derived views. Search indexes, caches, materialized reports and graph projections may be rebuilt from authoritative data. If two writable systems both claim to own the same field, define conflict resolution and failure handling or simplify the design.
| Workload signal | Strong starting candidate | Reject or qualify when |
|---|---|---|
| relational joins, constraints, SQL and multi-row transactions | RDS or Aurora with the compatible engine | unbounded horizontal write scale or key-only access dominates; engine compatibility is not exact |
| known partition-key access at large scale and low latency | DynamoDB | ad-hoc joins/queries dominate, keys are unknown, or hot partitions cannot be redesigned |
| JSON document aggregate with document-oriented API | DocumentDB when its supported MongoDB compatibility is sufficient | the application depends on unsupported MongoDB behavior; validate compatibility, indexes and limits |
| relationship traversal and path questions | Neptune Database; Neptune Analytics for analytical graph work | ordinary foreign-key joins are enough or graph projection freshness is undefined |
| Apache Cassandra/CQL-compatible wide-column patterns | Keyspaces | Cassandra feature compatibility or partition design does not fit; verify supported CQL behavior |
| high-resolution InfluxDB time-series workload | Timestream for InfluxDB | operational ownership, cardinality, retention or current service availability does not fit |
| ephemeral acceleration, sessions, counters or locks | ElastiCache for Valkey/Redis OSS/Memcached | it would become the only durable copy or stale/evicted data breaks correctness |
| columnar BI, large scans and warehouse SQL | Redshift provisioned or Serverless | request-by-request OLTP or tiny operational lookups dominate |
| search, relevance, log exploration and aggregations | OpenSearch Service | it is being treated as the transactional source of truth |
No table chooses automatically. For example, a JSON column in PostgreSQL may be simpler than adding a document database; a relational adjacency table may handle a small graph; DynamoDB Streams may feed a Redshift analytical copy. The architect must show measurable reasons for specialization.
Make consistency and integration explicit
Within one database, use the strongest transaction boundary required by the invariant. Across databases, assume there is no universal ACID transaction. A common pattern is:
API transaction
-> write authoritative business row/item + outbox record atomically
-> stream/change-data-capture publisher
-> durable event destination
-> idempotent consumers update cache/search/graph/warehouse projections
-> reconciliation finds missing, duplicate or divergent projections
Delivery can be at least once, out of order or delayed. Give every event an immutable ID, entity ID, version and occurred-at time. Consumers must detect duplicates and stale versions. Define what users see while a projection lags, how poison events reach a dead-letter path, how replay works, and how deletion/privacy requests propagate. “Eventual consistency” is not a design until the acceptable staleness and repair path are numeric.
Evaluate architecture dimensions
Score candidates from 1 (poor) to 5 (strong), but attach evidence to every score:
- correctness: transactions, consistency, constraints, concurrency and conflict rules;
- access fit: exact read/write/query patterns and indexes without scans or hot keys;
- scale/performance: throughput, latency percentiles, item/row limits, partitions, connections and concurrency;
- availability/recovery: Multi-AZ behavior, replicas, backup/PITR, cross-Region replication, tested RPO/RTO and failback;
- security/governance: IAM/control plane, data-plane authentication, private networking, encryption, secrets, audit, residency and deletion;
- operations: patching, schema/index changes, monitoring, capacity, troubleshooting, upgrades and team skill;
- migration/reversibility: export format, change capture, cutover, rollback and vendor/API compatibility;
- cost: idle floor plus requests/compute, storage, I/O, backup, replica, transfer, licensing and engineering effort.
Weight correctness and mandatory compliance as pass/fail gates rather than allowing cheap cost to cancel a failed requirement. Document assumptions such as growth and traffic and add a date to volatile pricing, quotas and feature support.
Worked portfolio example
An online shop has orders, catalog browsing, sessions, recommendations, telemetry and finance reporting. A defensible first iteration might be:
| Need | Decision | Ownership and synchronization |
|---|---|---|
| orders, payments and inventory reservation | PostgreSQL-compatible RDS/Aurora | relational system of record; one transaction protects business invariants |
| product catalog | begin in the relational store; consider DynamoDB only after measured key-access scale justifies it | order lines retain immutable product facts needed for history |
| sessions and hot product fragments | ElastiCache with bounded TTL; durable login/account truth remains outside cache | cache miss reloads from the owner; eviction must not lose truth |
| recommendations | do not add Neptune merely because relationships exist; adopt graph only if path/traversal queries show value | derived projection fed idempotently from durable events |
| telemetry | streaming/object-storage pipeline; choose a current time-series engine only for required interactive time-window queries | retention classes and deletion policy are explicit |
| finance and BI | Redshift after volume/concurrency warrants a warehouse; start simpler when it does not | CDC/batch projection reconciled to the order system |
This design intentionally rejects several services at launch. Add a specialist store only when a benchmark, operational requirement or cost model shows why the existing owner cannot meet the need.
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Open the AWS database category and the Pricing Calculator estimate prepared in AWS122.
- Review RDS, DynamoDB, ElastiCache, Redshift, and purpose-built inventories without changing anything.
- Use each service's current documentation and limits page to validate the final decision assumptions.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws rds describe-db-instances --query 'DBInstances[].Engine' --output text
aws dynamodb list-tables --output text
aws elasticache describe-cache-clusters --query 'CacheClusters[].Engine' --output text
aws redshift describe-clusters --query 'Clusters[].ClusterIdentifier' --output text
Expected interpretation
The commands provide an ownership inventory, not an architecture answer. The ADR must begin with workload evidence and may correctly decide that no database resource should be added.
Practical work
Submit p06-database-adr.md for orders, catalog, sessions, recommendations, telemetry, cache and reporting. Include:
- workload assumptions and an access-pattern/guarantee table;
- one authoritative owner for every field and a diagram of derived copies;
- at least two candidates per workload, weighted scoring and explicit rejection reasons;
- normal data flow plus timeout, throttling, duplicate-event, stale-projection, Region-failure and restore flows;
- identity, network, encryption, secret, audit and data-residency boundaries;
- backup/PITR, RPO/RTO, restore validation, failover and failback ownership;
- capacity and five-year cost model with dates and sensitivity at 10x traffic/data;
- migration phases, dual-write avoidance, CDC/reconciliation, cutover gates and rollback;
- observability: service metrics, application SLIs, alarms, logs, traces and business reconciliation;
- a final minimum architecture and conditions that would trigger reconsideration.
Run three tabletop tests: loss of the cache, six-hour projection lag and loss of the primary database Region. For each, state user-visible behavior, preserved invariants, alarms, operator steps, recovery evidence and acceptable data loss. Also test a negative dependency: stop the projection consumer and prove the source-of-truth write remains correct while freshness alarms trigger.
Diagnose this topic from its own evidence
| Evidence pattern | Architectural defect | Correction |
|---|---|---|
| decision begins with “we want to use service X” | technology-first selection | restate measurable access patterns and score alternatives |
| same field is writable in two stores | ambiguous ownership and conflict risk | nominate one owner; make other copies derived or specify tested conflict semantics |
| synchronous request calls three databases | coupled latency and availability | shrink the correctness path; move noncritical projections to durable asynchronous flow |
| scans, throttles or connection exhaustion at forecast load | model/capacity mismatch | redesign keys/indexes/pooling, benchmark, or select a better-fit store |
| restore succeeds but application reconciliation fails | backup was tested below the business layer | validate schema, rows, permissions, downstream projections and application queries |
| cross-Region replica is called “zero data loss” without evidence | asynchronous RPO hidden | measure lag, define loss window and choose stronger controls if required |
| cache outage causes data loss | cache accidentally became system of record | persist truth durably and design miss/rebuild behavior |
| cost estimate includes compute only | incomplete TCO | add storage, I/O, backup, transfer, replicas, minimum capacity and operations |
Cost and cleanup
The cost model includes idle minimums, capacity or requests, storage, I/O, backups, replicas, transfer, caches, analytics scans, and operational effort.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Make a defensible database portfolio decision across relational, key-value, document, graph, time-series, cache, and warehouse needs.
- Which scope or ownership boundary must be proved first?
Expected direction: One application can use several data stores, but each adds identity, networking, consistency, backup, monitoring, skills, migration, and cost responsibilities.
- What evidence is strong enough to accept the result?
Expected direction: The final ADR maps each access pattern to one source of truth, defines integrations, rejects unnecessary stores, and includes failure, recovery, security, and cost evidence.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid a service-per-feature design that duplicates truth, hides consistency, increases recovery coupling, and cannot be operated by the team.
- Which cost dimensions and retained resources need an owner?
Expected direction: The cost model includes idle minimums, capacity or requests, storage, I/O, backups, replicas, transfer, caches, analytics scans, and operational effort.
Lesson acceptance
Pass only when another architect can trace every requirement to evidence and reproduce the decision. The ADR must identify authoritative data owners, transaction and consistency boundaries, integrations and replay, security, availability, tested recovery, observability, migration/rollback and full cost dimensions. It must reject at least two plausible services with technical reasons, include the three tabletop failures, and avoid unsupported claims such as “managed means no operations” or “multi-Region means no data loss.”
Fail the lesson if it is merely a service comparison table, duplicates write ownership, omits degraded behavior, assumes backups equal recovery, cites undated pricing/features, or introduces a database without an accountable operator and deletion/retention policy.