Lesson 136 · AWS Learning Path

AWS 136: DocumentDB, Neptune, Keyspaces and Timestream for InfluxDB

· Published · 11 min read

Labelled process diagram for AWS 136: Native data shape to Compatible API and key or graph model to Purpose-built database to Query, backup, and cost evidence, with decision, proof and rejection evidence.

Why this lesson matters

Compare document, graph, wide-column, and current managed InfluxDB time-series workloads without directing new accounts to closed Timestream for LiveAnalytics.

These services are not interchangeable “NoSQL databases.” Each has a native data model, query/API compatibility boundary, topology, scaling unit, availability, backup path, and price model. New customers should follow current AWS guidance toward Timestream for InfluxDB rather than attempting to create Timestream for LiveAnalytics.

What you will be able to do

By the end, you can:

  • explain documentdb, neptune, keyspaces and timestream for influxdb in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeCompare document, graph, wide-column, and current managed InfluxDB time-series workloads without directing new accounts to closed Timestream for LiveAnalytics.
Scope and boundaryDocumentDB offers MongoDB-compatible document workloads, Neptune supports graph models, Keyspaces supports Apache Cassandra-compatible wide-column access, and Timestream for InfluxDB provides managed InfluxDB instances.
Evidence of successThe data model, query language, compatibility boundary, network placement, availability, backup, encryption, monitoring, migration, and client behavior are tested.
Cost modelInstances or capacity, storage, I/O or requests, backups, data transfer, monitoring, and multi-AZ or replica choices vary by service.
Safe rejection ruleAvoid compatibility assumptions, using a specialized database only because its label matches a noun, or selecting Timestream for LiveAnalytics for a new account.

How the request flows

+----------------------+
|  Native data shape   |
+----------------------+
           |
           v
+-----------------------------------------+
|  Compatible API and key or graph model  |
+-----------------------------------------+
                    |
                    v
+--------------------------+
|  Purpose-built database  |
+--------------------------+
             |
             v
+------------------------------------+
|  Query, backup, and cost evidence  |
+------------------------------------+

For Purpose-built databases, the important boundary is this: DocumentDB offers MongoDB-compatible document workloads, Neptune supports graph models, Keyspaces supports Apache Cassandra-compatible wide-column access, and Timestream for InfluxDB provides managed InfluxDB instances. The data model, query language, compatibility boundary, network placement, availability, backup, encryption, monitoring, migration, and client behavior are tested. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.

Architecture decision table

SituationDirectionReason
Requirement matchesUse the service whose native data model and query pattern remove real application or operational complexity.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid compatibility assumptions, using a specialized database only because its label matches a noun, or selecting Timestream for LiveAnalytics for a new account.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

Amazon DocumentDB

DocumentDB stores JSON-like documents and implements MongoDB-compatible APIs for supported versions/features. Compatibility does not mean it runs MongoDB server code or supports every operator, index, driver, tool, transaction, change-stream, collation, or administration command. Run the AWS compatibility tool/current matrix and application test suite before migration.

Instance-based clusters separate writer/readers from a shared cluster volume replicated across AZs, with cluster/reader/instance endpoints, backups/PITR, snapshots, KMS, parameter groups, and failover. Current engine versions and storage limits/configurations evolve; DocumentDB 5.0+ can offer standard versus I/O-Optimized economics, and newer versions increase storage limits. Elastic clusters shard for much larger scale but have a different feature/API/backup/operations matrix.

Model document boundaries to keep common reads atomic and under size limits. Avoid unbounded arrays and joins hidden in application loops. Index every measured query but account for write/storage amplification. Explain query plans/profiler, connections, cursor behavior, replica lag, cache, CPU, I/O, and change-stream retention.

Security uses private VPC access, SGs, TLS with current CA, Secrets Manager, database users/roles, IAM control plane, KMS, audit/profiler logs, and snapshot policies. Test writer/reader routing, unsupported MongoDB feature, failover reconnect, PITR restore, and migration reconciliation.

Amazon Neptune Database and Neptune Analytics

Neptune Database is a managed graph database. Choose one graph model:

  • property graph queried with Gremlin or openCypher;
  • RDF triples queried with SPARQL.

Property-graph data cannot be queried through SPARQL and RDF data cannot be queried through Gremlin/openCypher. Language compatibility has service-specific limits. Graph benefit appears when queries traverse relationships (fraud rings, knowledge graphs, recommendations, network dependencies), not merely because rows have foreign keys.

Neptune clusters have writer/read replicas, shared storage, endpoints, subnet/SG/TLS/authentication, parameter groups, backups, streams, logs, and failover concepts. Serverless capacity and Global Database are separate choices with version/Region constraints. Bulk loader uses S3 plus IAM role and must validate failed records, duplicates, edge/vertex IDs, and rollback.

Design vertex/edge labels, properties, cardinality, traversal direction, supernodes, indexes/statistics, query timeouts, and result bounds. Use explain/profile/query status and slow-query logs. A high-degree supernode or unbounded traversal can consume memory and time despite a small result.

Neptune Analytics is a separate graph analytics engine for algorithms/vector/large analytical workloads and does not replace an always-on transactional Neptune Database automatically. Define data load/snapshot/export, freshness, security, and cost.

Required tests: bounded two-hop traversal, unauthorized graph access, malformed loader row, supernode/timeout diagnosis, writer failover, backup restore, and result reconciliation.

Amazon Keyspaces for Apache Cassandra

Keyspaces provides a serverless, Cassandra Query Language compatible wide-column service. It presents regional endpoints and a Cassandra-compatible data model without customer-managed Cassandra nodes/rings/repair. Compatibility is bounded: supported CQL commands, data types, functions, drivers, consistency levels, compaction/storage behavior, and limits must be checked.

Partition key distributes rows; clustering columns order rows within a partition. Start from exact CQL queries. Low-cardinality/hot partitions, unbounded partition growth, tombstones, large rows, and unsupported ALLOW FILTERING-style thinking produce failures. Time-bucket high-volume series and model query-first tables.

Writes use LOCAL_QUORUM. Supported reads include ONE, LOCAL_ONE, and LOCAL_QUORUM, with different consistency/capacity. Other familiar Cassandra levels are unsupported. On-demand or provisioned/auto-scaling capacity is measured in request/capacity units with item-size rounding; TTL, PITR, client-side timestamps, CDC integrations, and networking/IAM authorization add cost and behavior.

Multi-Region keyspaces are active-active and replicate asynchronously. Last-writer-wins uses client-side timestamps; clock/application conflicts can discard a business update. Every replica must be sized for local plus replicated writes. Route 53/application routing handles regional traffic - there is no primary promotion. Monitor ReplicationLatency, test conflicting writes, and validate PITR across matching replica topology. Current multi-Region feature/KMS/TTL constraints must be checked before selection.

Required tests: query by full partition key, rejected unsupported consistency/query, hot/unbounded partition, TTL/tombstone behavior, capacity throttle/retry, multi-Region conflict/routing, and PITR restore to a new table.

Amazon Timestream for InfluxDB

Timestream for InfluxDB runs managed InfluxDB-compatible 2.x/current supported engines inside a VPC for time-series ingestion and millisecond queries. A DB instance contains Influx organizations/buckets and is reached with Influx API/UI/client tools; there is no host access. The AWS control-plane user/password initializes the service, while Influx operator/API tokens and authorization must be protected separately.

Select DB instance class, up to current supported storage, IOPS-Included tier, subnet/SG, private/public setting, IPv4/dual-stack, parameter group, log delivery, maintenance, and Single-AZ/Multi-AZ. Multi-AZ standby is synchronous HA and not readable. New read-replica cluster deployments use a writer and asynchronous reader with an InfluxData licensed add-on; they provide read scaling/failover but add lag, licensing, endpoint, and cost considerations. Deployment choices can be one-way at creation/modification - verify before building.

Influx schema design controls measurement, tag, field, timestamp, series cardinality, shard/bucket retention, downsampling/tasks, and query performance. High-cardinality tags can exhaust resources; putting query dimensions in fields causes scans. Use line-protocol validation, precision/time zones, duplicate-point semantics, late data, retention, and export/backup plans.

AWS documents internal backups for service recovery and newer customer-managed backup/PITR capabilities by engine generation. Do not promise self-service restore without checking exact Timestream for InfluxDB version. Test backup restore, token/secret recovery, Multi-AZ failover, replica lag, maintenance, and log delivery.

AWS states Timestream for LiveAnalytics is no longer accepting new customers. Existing customers can continue supported workloads, but this course does not direct a new account to create it. LiveAnalytics' memory/magnetic-store architecture is not the architecture of Timestream for InfluxDB.

Cross-service decision table

RequirementSelect directionNearest wrong choice
MongoDB API/document migrationDocumentDB after compatibility testDynamoDB document attributes do not provide MongoDB API/query compatibility
Deep relationship traversalNeptuneRelational joins/key-value reads become complex or unbounded
Cassandra CQL wide-column at managed scaleKeyspacesDocumentDB does not implement Cassandra partition/clustering semantics
Influx line protocol/Flux or supported SQL time-series ecosystemTimestream for InfluxDBLiveAnalytics is closed to new customers and has a different API/model
Simple known-key document lookupDynamoDB may be simplerDocumentDB cluster cost/query flexibility may be unnecessary
BI warehouse aggregationsRedshiftGraph/document/time-series service is not a columnar warehouse

Complete evaluation lab

For one workload per service, create a data/query contract, compatibility matrix, topology diagram, capacity/cardinality estimate, failure plan, backup restore, migration method, and monthly cost. Run at least three representative queries, one unsupported/negative operation, one network/identity denial, one capacity/skew failure, and one restore validation using an approved disposable environment or supplied exact evidence.

Cost must include every instance/node/replica/license, storage/I/O/request units, backup/snapshot/export, transfer/multi-Region, KMS, logging, and idle retention. Cleanup must remove replicas/clusters/instances/tables, snapshots/backups, secrets/tokens, roles, subnet/parameter groups, SG rules, log buckets/groups, and KMS retained state in reviewed order.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open DocumentDB clusters, Neptune databases, Keyspaces tables, and Timestream for InfluxDB DB instances in separate read-only tabs.
  2. For each supplied example, record engine or API compatibility, network placement, encryption, backup, status, and endpoint shape.
  3. Confirm that no new-customer lesson asks the learner to create Timestream for LiveAnalytics.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws docdb describe-db-clusters --query 'DBClusters[].{Id:DBClusterIdentifier,Status:Status,Encrypted:StorageEncrypted}' --output table
aws neptune describe-db-clusters --query 'DBClusters[].{Id:DBClusterIdentifier,Status:Status,Encrypted:StorageEncrypted}' --output table
aws keyspaces list-keyspaces --query 'keyspaces[].keyspaceName' --output table
aws timestream-influxdb list-db-instances --query 'items[].{Id:id,Name:name,Status:status,Deployment:deploymentType}' --output table

Expected interpretation

The inventory identifies purpose-built control-plane objects. It does not prove query compatibility, client libraries, graph traversal efficiency, partition design, or time-series retention behavior.

Practical work

Map product documents, fraud relationships, device-by-time rows, and Cassandra-style IoT keys to the four current services. Write one sample query pattern and one rejection reason for each.

Diagnose this topic from its own evidence

First identify which engine and API contract is failing; compatibility claims are not identity claims. For DocumentDB, inspect supported MongoDB API behavior, indexes, connections, replica lag and profiler/slow-query evidence. For Neptune, identify property-graph versus RDF model, query language, loader/IAM/S3 path and whether a supernode or unbounded traversal caused the problem. For Keyspaces, inspect CQL support, partition-key distribution, consistency, capacity and regional replication lag. For Timestream for InfluxDB, inspect VPC/DNS/TLS/token access, write cardinality, storage/retention, instance/cluster role and Multi-AZ/read-replica status.

Negative test: list one familiar upstream feature that each managed compatible service does not promise, verify it in current official compatibility documentation, and show the application test that prevents accidental dependence on it.

Cost and cleanup

Instances or capacity, storage, I/O or requests, backups, data transfer, monitoring, and multi-AZ or replica choices vary by service.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Compare document, graph, wide-column, and current managed InfluxDB time-series workloads without directing new accounts to closed Timestream for LiveAnalytics.

  1. Which scope or ownership boundary must be proved first?

Expected direction: DocumentDB offers MongoDB-compatible document workloads, Neptune supports graph models, Keyspaces supports Apache Cassandra-compatible wide-column access, and Timestream for InfluxDB provides managed InfluxDB instances.

  1. What evidence is strong enough to accept the result?

Expected direction: The data model, query language, compatibility boundary, network placement, availability, backup, encryption, monitoring, migration, and client behavior are tested.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid compatibility assumptions, using a specialized database only because its label matches a noun, or selecting Timestream for LiveAnalytics for a new account.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Instances or capacity, storage, I/O or requests, backups, data transfer, monitoring, and multi-AZ or replica choices vary by service.

Lesson acceptance

Pass when the learner selects among document, graph, wide-column and InfluxDB time-series models from concrete access patterns; states compatibility boundaries; designs keys/indexes/labels/tags and cardinality; and explains networking, authentication, encryption, scaling, availability, backups and recovery for the selected service. Include migration tests, quotas, monitoring, failure evidence, current regional/service availability, cost and cleanup. The learner must explicitly note that Timestream for LiveAnalytics is unavailable to new customers and must not substitute it silently. Fail if product names replace workload evidence or “compatible” is interpreted as feature-identical.

Official sources

Advertisement