AWS 122: Choosing the correct database
Why this lesson matters
Choose a database by data model and access pattern before choosing an engine name.
A database is selected from the shape of data, required operations, consistency, scale, resilience, and operating model - not from familiarity or an exam keyword. This lesson builds that decision method before introducing individual AWS engines.
What you will be able to do
By the end, you can:
- explain choosing the correct database in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Choose a database by data model and access pattern before choosing an engine name. |
| Scope and boundary | The decision spans relational, key-value, document, graph, wide-column, time-series, cache, search, ledger, and warehouse workloads. Managed responsibility differs by service. |
| Evidence of success | The chosen service supports the required queries, consistency, scale, availability, recovery, security, operations, and migration path. |
| Cost model | Instances or capacity units, storage, I/O, backup, replication, transfer, requests, and reserved commitments differ by database family. |
| Safe rejection rule | Do not force a familiar relational engine onto every workload or select a database from a single exam keyword. |
How the request flows
+----------------------+
| Business question |
+----------------------+
|
v
+---------------------------------+
| Data model and access pattern |
+---------------------------------+
|
v
+------------------------------------+
| Consistency, scale, and recovery |
+------------------------------------+
|
v
+-----------------------------+
| Managed database decision |
+-----------------------------+
For AWS database portfolio, the important boundary is this: The decision spans relational, key-value, document, graph, wide-column, time-series, cache, search, ledger, and warehouse workloads. Managed responsibility differs by service. The chosen service supports the required queries, consistency, scale, availability, recovery, security, operations, and migration path. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use the portfolio when requirements include stable transactions, flexible key access, graph traversal, time-series ingestion, caching, or analytics. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Do not force a familiar relational engine onto every workload or select a database from a single exam keyword. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
Start with access patterns, not products
Write the operations the application must perform, with units and guarantees:
operation: GetOrderById
key/filter: order_id exact equality
result size: one order plus up to 100 lines
rate: 2,000 average / 12,000 peak requests per second
write/read ratio: 1:5
consistency: read-your-write for checkout; stale up to 30 s for analytics
transaction: order + payment reservation must commit atomically
latency: p99 under 80 ms in ap-south-1
growth: 3 TB now, 20 TB in 36 months
availability: multi-AZ; RPO < 1 minute; RTO < 15 minutes
retention/security: seven years, encryption, tenant isolation, deletion exceptions
“Store customer data” is not an access pattern. Include creates, updates, deletes, point reads, range scans, joins, aggregations, traversals, full-text/vector search, streams/change capture, TTL, bulk loads, backups, and administrative reports. Identify which operations are synchronous business paths and which can be asynchronous projections.
Database families and AWS directions
| Data/operation model | AWS services to evaluate | Reject when |
|---|---|---|
| Relational transactions, SQL, joins, constraints | Amazon RDS engines, Amazon Aurora; self-managed EC2 for requirements RDS cannot meet | Schema/joins are not actually required and horizontal key access dominates, or unsupported engine/OS control is mandatory |
| Distributed SQL with active-active regional design | Aurora DSQL where current Region/feature/semantic requirements fit | Required PostgreSQL compatibility/features, transaction model, latency, or operational constraints are unsupported |
| Key-value/document at massive scale | DynamoDB | Ad hoc joins, arbitrary relational queries, or access patterns cannot be modeled as keys/indexes |
| JSON document with MongoDB compatibility requirement | Amazon DocumentDB after compatibility testing | Exact MongoDB version/feature/operator behavior is required but unsupported |
| Graph traversal and relationships | Amazon Neptune | Workload is simple key lookup/relational join or required graph language/feature is unsupported |
| Wide-column Cassandra-compatible | Amazon Keyspaces | Cassandra compatibility/consistency/feature limits or access patterns do not fit |
| Time-series measurements | Amazon Timestream services/current offerings, including Timestream for InfluxDB where appropriate | Workload is primarily transactions, documents, or arbitrary joins |
| In-memory cache/session/leaderboard | ElastiCache or MemoryDB according to durability/compatibility needs | Cache is being mistaken for the system of record without accepted loss/recovery semantics |
| Analytics warehouse/columnar SQL | Amazon Redshift | Millisecond OLTP row transactions dominate or a data lake/query service is a better operational fit |
| Search, log analytics, vector retrieval | Amazon OpenSearch Service or purpose-built vector options | Strong relational transactions are required or search index cannot be rebuilt/governed as designed |
| Immutable cryptographically verifiable journal | Amazon QLDB only for existing supported lifecycle; verify current service availability and migration guidance | New adoption conflicts with current service status or a normal audit table/event log meets the need |
One application can use several databases deliberately. This “polyglot persistence” reduces compromise only when ownership, synchronization, failure handling, reconciliation, backup, deletion, and cost are explicit. Otherwise it creates distributed inconsistency.
Consistency and transaction questions
- Strong read: returns the latest successful write according to the service operation's guarantee.
- Eventually consistent read: can return an older value for a period; often cheaper/faster/more available at scale.
- Read-your-write: one user/workflow sees its own commit; this may require routing or tokens beyond generic eventual reads.
- Atomicity: all operations in a transaction commit or none do.
- Isolation: concurrent transactions do not create prohibited intermediate outcomes; isolation levels differ.
- Durability: committed data survives defined failures; this is not the same as availability.
CAP-style slogans do not select a database. Name the actual network/failure condition and the business behavior. Also distinguish database replication from backup: replicas can copy corruption or deletes, while backups provide historical recovery but not necessarily immediate serving.
Managed responsibility boundary
AWS can manage physical hosts, service software installation, portions of patching, storage, and failover depending on the service. The customer still owns data model, indexes, query plans, users/roles, credentials, network access, encryption choices/keys, parameter changes, version scheduling, capacity limits, monitoring, backups/retention, restore testing, schema migration, application retries, and cost.
Serverless/on-demand pricing removes some provisioning decisions, not responsibility. Provisioned engines can be more predictable or economical for steady usage. Model requests/read-write units/ACUs/nodes/instances, storage, I/O, backup, replication, transfer, streams, indexes, monitoring, support, and migration/engineering cost.
Decision workflow
- Inventory entities, relationships, volume, growth, item/row/document sizes, and data lifecycle.
- Write every access pattern with rate, latency, result size, consistency, and transaction requirement.
- Define failure scope, availability, RPO, RTO, Region, residency, and restore validation.
- Define security: identities, network path, encryption/key control, tenant isolation, audit, retention/deletion.
- Shortlist data models first, then AWS services/engines and exact versions.
- Test the hardest queries, peak writes, hot-key/skew behavior, failover, backup/restore, and schema evolution using production-shaped data.
- Calculate normal, peak, replication, recovery, and retained-backup cost.
- Record the selected option, rejected alternatives, reversal conditions, risks, and owner in an architecture decision record.
Worked selections
Orders and payments
Multi-row constraints, transactions, SQL reporting, and established relational skills point to RDS/Aurora. DynamoDB can fit only after redesigning transactions/access patterns and accepting denormalized projections; it is not selected merely for scale.
Shopping-cart sessions
Known key lookup by customer, TTL, high scale, and limited transactions point to DynamoDB. Add DAX only if measured read latency/load justifies a cache. An RDS table remains possible at smaller scale, but adds connection and vertical/sharding considerations.
Fraud relationship traversal
Queries such as “accounts connected within four transfers to known fraud” point to Neptune graph traversal. A relational recursive query may work at modest scale; benchmark both and evaluate graph language/team skills.
Operational dashboard from transactional data
Do not run unbounded analytics on the writer by default. Use a read replica for safe operational reports, or CDC/ETL into Redshift/OpenSearch/data lake depending on query. Define lag and reconciliation.
Product search
Keep authoritative product/price transactions in a database and build a searchable OpenSearch projection. Search results can lag; checkout revalidates price/stock against the source of truth.
Required decision evidence
- A complete access-pattern table and growth model.
- At least three candidate services with feature/version/Region evidence.
- A small benchmark proving the highest-risk query/write path and skew behavior.
- Positive transaction/read, negative authorization/constraint, dependency outage, failover, and restore tests.
- Data migration, CDC/reconciliation, schema/version change, rollback, monitoring, and cleanup plans.
- A monthly normal/peak/recovery cost model and a decision reversal condition.
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Open the AWS database service category and identify RDS, DynamoDB, DocumentDB, Neptune, Keyspaces, ElastiCache, Redshift, and Timestream for InfluxDB.
- Open RDS Databases and DynamoDB Tables in ap-south-1 and observe that an empty list is a valid starting state.
- Open AWS Pricing Calculator and locate database estimates without saving or purchasing anything.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws rds describe-db-instances --query 'DBInstances[].{Id:DBInstanceIdentifier,Engine:Engine,AZ:AvailabilityZone,Status:DBInstanceStatus}' --output table
aws dynamodb list-tables --output table
aws redshift describe-clusters --query 'Clusters[].{Id:ClusterIdentifier,State:ClusterStatus}' --output table
Expected interpretation
The commands inventory three database families. Empty output says nothing about which data model fits the workload, and it may also reflect Region, account, or permission scope.
Practical work
Build a requirement matrix for orders, sessions, product documents, social relationships, device telemetry, cached catalog reads, and business reporting. Select and reject services with measurable reasons.
Diagnose this topic from its own evidence
If no candidate satisfies the matrix, revisit the access patterns and determine whether one authoritative store plus asynchronous projections is safer than forcing one engine. Diagnose bad selections by the first violated requirement: unsupported query/transaction, hot partition, lag, unavailable Region/feature, recovery failure, security boundary, operating skill, or total cost. Preserve benchmark and failure evidence; do not change the requirement to defend a preferred product.
Cost and cleanup
Instances or capacity units, storage, I/O, backup, replication, transfer, requests, and reserved commitments differ by database family.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Choose a database by data model and access pattern before choosing an engine name.
- Which scope or ownership boundary must be proved first?
Expected direction: The decision spans relational, key-value, document, graph, wide-column, time-series, cache, search, ledger, and warehouse workloads. Managed responsibility differs by service.
- What evidence is strong enough to accept the result?
Expected direction: The chosen service supports the required queries, consistency, scale, availability, recovery, security, operations, and migration path.
- Which tempting design or shortcut must be rejected?
Expected direction: Do not force a familiar relational engine onto every workload or select a database from a single exam keyword.
- Which cost dimensions and retained resources need an owner?
Expected direction: Instances or capacity units, storage, I/O, backup, replication, transfer, requests, and reserved commitments differ by database family.
Lesson acceptance
Pass only when every access pattern has units and consistency/transaction semantics, at least three credible candidates are compared, the hardest path is benchmarked, failure and restore are tested, all data copies/flows are governed, and the final ADR explains rejection and reversal conditions.