Lesson 131 · AWS Learning Path

AWS 131: DynamoDB keys and secondary indexes

· Published · 11 min read

Labelled process diagram for AWS 131: Access pattern to Key condition to Base table or index to Distribution and projection cost, with decision, proof and rejection evidence.

Why this lesson matters

Design partition keys, sort keys, GSIs, and LSIs from exact access patterns and distribution needs.

A DynamoDB schema is a set of deliberately supported access patterns. Partition and sort keys control data placement and ordering; secondary indexes create additional physical query views with their own consistency, write, storage, backfill, hot-key, and failure consequences.

What you will be able to do

By the end, you can:

  • explain dynamodb keys and secondary indexes in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeDesign partition keys, sort keys, GSIs, and LSIs from exact access patterns and distribution needs.
Scope and boundaryThe primary key places and orders items. A GSI can use different keys and scales separately. An LSI shares the partition key, is created with the table, and has item-collection constraints.
Evidence of successEvery query names its target key or index, key condition, projection, consistency need, expected item count, and hot-key risk.
Cost modelIndex writes and storage, projected attributes, reads, backfills, and over-provisioned capacity increase cost.
Safe rejection ruleAvoid low-cardinality hot keys, write sharding without a read plan, ALL projection by habit, and LSIs when item collections may exceed their limits.

How the request flows

+----------------------+
|    Access pattern    |
+----------------------+
           |
           v
+----------------------+
|    Key condition     |
+----------------------+
           |
           v
+-----------------------+
|  Base table or index  |
+-----------------------+
           |
           v
+------------------------------------+
|  Distribution and projection cost  |
+------------------------------------+

For DynamoDB keys and secondary indexes, the important boundary is this: The primary key places and orders items. A GSI can use different keys and scales separately. An LSI shares the partition key, is created with the table, and has item-collection constraints. Every query names its target key or index, key condition, projection, consistency need, expected item count, and hot-key risk. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.

Architecture decision table

SituationDirectionReason
Requirement matchesUse a GSI when an important query needs a different partition key; use a sort key for ordered item collections and range conditions.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid low-cardinality hot keys, write sharding without a read plan, ALL projection by habit, and LSIs when item collections may exceed their limits.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

Translate each question into a key operation

Business questionTarget key designEfficient operation
Get order 991PK=ORDER#991, known sort keyGetItem or bounded Query
List customer 42 orders newest firstPK=CUSTOMER#42, SK=ORDER#timestamp#idQuery partition, descending sort order, page
Orders in status PENDING by timeGSI GSI1PK=STATUS#PENDING, GSI1SK=timestamp#idQuery GSI; mitigate hot status key if traffic exceeds one partition
Unique email lookupGSI or transactional uniqueness item keyed by normalized hashQuery/Get plus conditional transaction; GSI alone is eventually consistent and cannot enforce uniqueness
Events between two times for devicePK=DEVICE#id#bucket, SK=timestamp#eventidQuery one/few time buckets and merge pages

Write expected maximum items/bytes per partition-key value, reads/writes per second, page size, growth, and consistency. A design is incomplete when the query says “filter by anything.”

Sort-key patterns

Sort keys support equality, inequalities, BETWEEN, and begins_with in key conditions. Hierarchical values such as COUNTRY#IN#STATE#MH#CITY#PUNE can support prefix queries. Time values must sort lexicographically in the intended order - use normalized UTC ISO-8601 or fixed-width epoch representations and add a unique suffix for collisions.

Common patterns include:

  • type prefixes (ORDER#, PAYMENT#) to colocate related items;
  • version-control pattern with a stable v0_ latest item plus immutable historical versions;
  • adjacency lists for bounded relationship access;
  • time bucketing to cap item collections and distribute writes;
  • write sharding (STATUS#PENDING#00..N) for hot low-cardinality indexes, with explicit parallel read/merge logic.

Do not embed unbounded mutable arrays in one 400-KB item. Split child records into item collections and use transactions/conditions where cross-item invariants matter.

GSI versus LSI

PropertyGlobal secondary indexLocal secondary index
Partition keyCan differ from tableMust equal table partition key
CreationCan be added/updated after table creation within service rulesMust be defined when table is created; cannot be added later
Storage/throughputSeparate distributed index and capacity behavior; on-demand follows table modeShares base-table partition placement/throughput accounting behavior
ConsistencyEventually consistent reads onlyCan request strong consistency
Item collectionNew partition distributionBase table + all LSIs under one partition key subject to 10-GB item-collection limit
Typical useInvert/query by another entity or attributeAlternate ordering/filter within the same bounded item collection

A GSI update is asynchronous. Immediately querying the index after writing the base table can miss the new item. If uniqueness or read-your-write is mandatory, use a base-table transaction/condition or read path designed for it.

Sparse indexes, overloading, and projection

Only items containing the index key attributes appear in an index. This makes a sparse GSI useful for rare workflow states: add EscalatedAt/index keys only to escalated orders. Removing the key removes the index entry asynchronously.

Index overloading reuses generic GSI key attributes for several entity/access-pattern types. It can reduce index count but increases naming, permission, operational, and migration complexity. Document every discriminator and collision rule.

Projection choices:

  • KEYS_ONLY: index and base key attributes; lowest index storage/write size but base-table fetch may be needed.
  • INCLUDE: keys plus selected non-key attributes; balance covered queries against write/storage amplification.
  • ALL: copies every attribute; simplest covered reads but maximum storage/write amplification.

For GSIs, fetching unprojected attributes requires a separate base-table read by the application; DynamoDB does not transparently fetch them as LSI querying can under supported behavior. Projection size contributes to index item size and write capacity. Every base-table write can update zero or several indexes.

Index lifecycle and backfill

Adding a GSI starts asynchronous resource allocation and backfill. The table remains usable, but index creation consumes resources and can be affected by write capacity, key violations, data skew, or quotas. Monitor index status, backfill progress, throttling, online index consumed capacity, and application latency. Do not route production reads until ACTIVE and data correctness checks pass.

Deleting a GSI is destructive to that query view. Preserve IaC/configuration, verify no application/role/dashboard depends on its ARN/name, remove traffic, then delete. Recreating performs another backfill and does not restore historical index behavior instantly.

Key-schema attributes must use supported scalar types. Existing data that lacks a sparse key is fine, but incompatible key values can create violations during backfill. Run violation detection/correction procedures from current documentation where applicable.

Hot-key and sharding worked example

GSI1PK=STATUS#PENDING receives 8,000 writes/second. One constant GSI partition key concentrates traffic beyond one partition's write limit. Add deterministic shards, for example STATUS#PENDING#00 through #15, chosen from order ID hash. Readers query all 16 partitions in parallel with bounded concurrency, merge by sort key, paginate each shard, and deduplicate/retry safely.

The tradeoff is write scalability for read fan-out and complexity. If status dashboards tolerate aggregation delay, a stream-fed aggregate may be better than querying every pending order.

Security and tenancy

IAM can constrain dynamodb:LeadingKeys and attributes for supported direct table operations, but index access and single-table mixed entities complicate tenant isolation. Never accept a client-supplied tenant key without authorization binding. Test table and every GSI ARN; an index can expose projected attributes through a new access path.

Separate tables may be safer when tenants/teams need independent KMS keys, backup/restore, capacity, lifecycle, policy, or blast radius. Single-table design is a performance/model choice, not a certification requirement.

Required design and failure tests

  1. For at least ten access patterns, name table/index, exact key condition, projection, max page, consistency, and capacity.
  2. Prove a key-based Query reads fewer items/capacity than a Scan+filter for the same result.
  3. Write then immediately query a GSI and explain any eventual visibility; do not add sleeps as a correctness guarantee.
  4. Attempt a duplicate unique attribute using a transaction/condition and prove rejection.
  5. Demonstrate a sparse index by adding/removing index-key attributes.
  6. Load a skewed status GSI, observe throttling, implement sharding, and measure read fan-out/cost.
  7. Model GSI backfill and rollback with status/throttle/correctness gates.
  8. Calculate base item plus every projected index item, write amplification, reads, streams, backup, and global replication.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open DynamoDB Tables, choose supplied evidence, and inspect the primary key in General information.
  2. Open Indexes and compare partition key, sort key, projection, status, and capacity for each GSI or LSI.
  3. Open Explore items and configure a Query by exact key rather than executing an unbounded Scan.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws dynamodb describe-table --table-name replace-with-table-name --query 'Table.{Primary:KeySchema,Attributes:AttributeDefinitions,GSIs:GlobalSecondaryIndexes[].{Name:IndexName,Keys:KeySchema,Projection:Projection.ProjectionType,Status:IndexStatus},LSIs:LocalSecondaryIndexes[].{Name:IndexName,Keys:KeySchema}}' --output json

Expected interpretation

The schema proves declared keys and indexes. It does not prove balanced cardinality, correct application queries, or acceptable backfill and write amplification.

Practical work

Create an access-pattern table for get order, list customer orders by date, find orders by status, and idempotency lookup. Map each to base table or GSI and estimate key cardinality.

Diagnose this topic from its own evidence

For an empty or slow query, record the table/index name, key condition, filter, ScannedCount, Count, consumed capacity and pagination token. A filter is applied after DynamoDB reads matching key data, so a small returned count does not imply low cost. If a new GSI is CREATING or backfilling, inspect index status, backfilling and throttling metrics; do not assume it is ready. For throttling on one tenant or date bucket, compare key frequency and item sizes to expose skew. For missing index fields, inspect projection configuration and remember that GSI propagation is eventually consistent.

Negative test: issue a query for an access pattern not represented by a table or index key. Reject a scan-based “solution,” add the proposed key/index to the design only, and calculate its write amplification, storage, backfill and hot-key consequences before creation.

Cost and cleanup

Index writes and storage, projected attributes, reads, backfills, and over-provisioned capacity increase cost.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Design partition keys, sort keys, GSIs, and LSIs from exact access patterns and distribution needs.

  1. Which scope or ownership boundary must be proved first?

Expected direction: The primary key places and orders items. A GSI can use different keys and scales separately. An LSI shares the partition key, is created with the table, and has item-collection constraints.

  1. What evidence is strong enough to accept the result?

Expected direction: Every query names its target key or index, key condition, projection, consistency need, expected item count, and hot-key risk.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid low-cardinality hot keys, write sharding without a read plan, ALL projection by habit, and LSIs when item collections may exceed their limits.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Index writes and storage, projected attributes, reads, backfills, and over-provisioned capacity increase cost.

Lesson acceptance

Pass when every required access pattern maps to an exact table/GSI/LSI key condition, uniqueness and adjacency semantics are explicit, sparse-index behavior and projection are demonstrated, and partition distribution is justified with sample cardinalities. The learner must compare GSI and LSI lifecycle, consistency, capacity, size and creation constraints; plan online GSI backfill and rollback; test pagination and empty results; and account for duplicated writes/storage. Fail if normal traffic needs a full scan, filters are mistaken for key selection, or a single popular key creates an unexplained hot partition.

Official sources

Advertisement