AWS 131: DynamoDB keys and secondary indexes
Why this lesson matters
Design partition keys, sort keys, GSIs, and LSIs from exact access patterns and distribution needs.
A DynamoDB schema is a set of deliberately supported access patterns. Partition and sort keys control data placement and ordering; secondary indexes create additional physical query views with their own consistency, write, storage, backfill, hot-key, and failure consequences.
What you will be able to do
By the end, you can:
- explain dynamodb keys and secondary indexes in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Design partition keys, sort keys, GSIs, and LSIs from exact access patterns and distribution needs. |
| Scope and boundary | The primary key places and orders items. A GSI can use different keys and scales separately. An LSI shares the partition key, is created with the table, and has item-collection constraints. |
| Evidence of success | Every query names its target key or index, key condition, projection, consistency need, expected item count, and hot-key risk. |
| Cost model | Index writes and storage, projected attributes, reads, backfills, and over-provisioned capacity increase cost. |
| Safe rejection rule | Avoid low-cardinality hot keys, write sharding without a read plan, ALL projection by habit, and LSIs when item collections may exceed their limits. |
How the request flows
+----------------------+
| Access pattern |
+----------------------+
|
v
+----------------------+
| Key condition |
+----------------------+
|
v
+-----------------------+
| Base table or index |
+-----------------------+
|
v
+------------------------------------+
| Distribution and projection cost |
+------------------------------------+
For DynamoDB keys and secondary indexes, the important boundary is this: The primary key places and orders items. A GSI can use different keys and scales separately. An LSI shares the partition key, is created with the table, and has item-collection constraints. Every query names its target key or index, key condition, projection, consistency need, expected item count, and hot-key risk. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use a GSI when an important query needs a different partition key; use a sort key for ordered item collections and range conditions. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid low-cardinality hot keys, write sharding without a read plan, ALL projection by habit, and LSIs when item collections may exceed their limits. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
Translate each question into a key operation
| Business question | Target key design | Efficient operation |
|---|---|---|
| Get order 991 | PK=ORDER#991, known sort key | GetItem or bounded Query |
| List customer 42 orders newest first | PK=CUSTOMER#42, SK=ORDER#timestamp#id | Query partition, descending sort order, page |
| Orders in status PENDING by time | GSI GSI1PK=STATUS#PENDING, GSI1SK=timestamp#id | Query GSI; mitigate hot status key if traffic exceeds one partition |
| Unique email lookup | GSI or transactional uniqueness item keyed by normalized hash | Query/Get plus conditional transaction; GSI alone is eventually consistent and cannot enforce uniqueness |
| Events between two times for device | PK=DEVICE#id#bucket, SK=timestamp#eventid | Query one/few time buckets and merge pages |
Write expected maximum items/bytes per partition-key value, reads/writes per second, page size, growth, and consistency. A design is incomplete when the query says “filter by anything.”
Sort-key patterns
Sort keys support equality, inequalities, BETWEEN, and begins_with in key conditions. Hierarchical values such as COUNTRY#IN#STATE#MH#CITY#PUNE can support prefix queries. Time values must sort lexicographically in the intended order - use normalized UTC ISO-8601 or fixed-width epoch representations and add a unique suffix for collisions.
Common patterns include:
- type prefixes (
ORDER#,PAYMENT#) to colocate related items; - version-control pattern with a stable
v0_latest item plus immutable historical versions; - adjacency lists for bounded relationship access;
- time bucketing to cap item collections and distribute writes;
- write sharding (
STATUS#PENDING#00..N) for hot low-cardinality indexes, with explicit parallel read/merge logic.
Do not embed unbounded mutable arrays in one 400-KB item. Split child records into item collections and use transactions/conditions where cross-item invariants matter.
GSI versus LSI
| Property | Global secondary index | Local secondary index |
|---|---|---|
| Partition key | Can differ from table | Must equal table partition key |
| Creation | Can be added/updated after table creation within service rules | Must be defined when table is created; cannot be added later |
| Storage/throughput | Separate distributed index and capacity behavior; on-demand follows table mode | Shares base-table partition placement/throughput accounting behavior |
| Consistency | Eventually consistent reads only | Can request strong consistency |
| Item collection | New partition distribution | Base table + all LSIs under one partition key subject to 10-GB item-collection limit |
| Typical use | Invert/query by another entity or attribute | Alternate ordering/filter within the same bounded item collection |
A GSI update is asynchronous. Immediately querying the index after writing the base table can miss the new item. If uniqueness or read-your-write is mandatory, use a base-table transaction/condition or read path designed for it.
Sparse indexes, overloading, and projection
Only items containing the index key attributes appear in an index. This makes a sparse GSI useful for rare workflow states: add EscalatedAt/index keys only to escalated orders. Removing the key removes the index entry asynchronously.
Index overloading reuses generic GSI key attributes for several entity/access-pattern types. It can reduce index count but increases naming, permission, operational, and migration complexity. Document every discriminator and collision rule.
Projection choices:
KEYS_ONLY: index and base key attributes; lowest index storage/write size but base-table fetch may be needed.INCLUDE: keys plus selected non-key attributes; balance covered queries against write/storage amplification.ALL: copies every attribute; simplest covered reads but maximum storage/write amplification.
For GSIs, fetching unprojected attributes requires a separate base-table read by the application; DynamoDB does not transparently fetch them as LSI querying can under supported behavior. Projection size contributes to index item size and write capacity. Every base-table write can update zero or several indexes.
Index lifecycle and backfill
Adding a GSI starts asynchronous resource allocation and backfill. The table remains usable, but index creation consumes resources and can be affected by write capacity, key violations, data skew, or quotas. Monitor index status, backfill progress, throttling, online index consumed capacity, and application latency. Do not route production reads until ACTIVE and data correctness checks pass.
Deleting a GSI is destructive to that query view. Preserve IaC/configuration, verify no application/role/dashboard depends on its ARN/name, remove traffic, then delete. Recreating performs another backfill and does not restore historical index behavior instantly.
Key-schema attributes must use supported scalar types. Existing data that lacks a sparse key is fine, but incompatible key values can create violations during backfill. Run violation detection/correction procedures from current documentation where applicable.
Hot-key and sharding worked example
GSI1PK=STATUS#PENDING receives 8,000 writes/second. One constant GSI partition key concentrates traffic beyond one partition's write limit. Add deterministic shards, for example STATUS#PENDING#00 through #15, chosen from order ID hash. Readers query all 16 partitions in parallel with bounded concurrency, merge by sort key, paginate each shard, and deduplicate/retry safely.
The tradeoff is write scalability for read fan-out and complexity. If status dashboards tolerate aggregation delay, a stream-fed aggregate may be better than querying every pending order.
Security and tenancy
IAM can constrain dynamodb:LeadingKeys and attributes for supported direct table operations, but index access and single-table mixed entities complicate tenant isolation. Never accept a client-supplied tenant key without authorization binding. Test table and every GSI ARN; an index can expose projected attributes through a new access path.
Separate tables may be safer when tenants/teams need independent KMS keys, backup/restore, capacity, lifecycle, policy, or blast radius. Single-table design is a performance/model choice, not a certification requirement.
Required design and failure tests
- For at least ten access patterns, name table/index, exact key condition, projection, max page, consistency, and capacity.
- Prove a key-based Query reads fewer items/capacity than a Scan+filter for the same result.
- Write then immediately query a GSI and explain any eventual visibility; do not add sleeps as a correctness guarantee.
- Attempt a duplicate unique attribute using a transaction/condition and prove rejection.
- Demonstrate a sparse index by adding/removing index-key attributes.
- Load a skewed status GSI, observe throttling, implement sharding, and measure read fan-out/cost.
- Model GSI backfill and rollback with status/throttle/correctness gates.
- Calculate base item plus every projected index item, write amplification, reads, streams, backup, and global replication.
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Open DynamoDB Tables, choose supplied evidence, and inspect the primary key in General information.
- Open Indexes and compare partition key, sort key, projection, status, and capacity for each GSI or LSI.
- Open Explore items and configure a Query by exact key rather than executing an unbounded Scan.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws dynamodb describe-table --table-name replace-with-table-name --query 'Table.{Primary:KeySchema,Attributes:AttributeDefinitions,GSIs:GlobalSecondaryIndexes[].{Name:IndexName,Keys:KeySchema,Projection:Projection.ProjectionType,Status:IndexStatus},LSIs:LocalSecondaryIndexes[].{Name:IndexName,Keys:KeySchema}}' --output json
Expected interpretation
The schema proves declared keys and indexes. It does not prove balanced cardinality, correct application queries, or acceptable backfill and write amplification.
Practical work
Create an access-pattern table for get order, list customer orders by date, find orders by status, and idempotency lookup. Map each to base table or GSI and estimate key cardinality.
Diagnose this topic from its own evidence
For an empty or slow query, record the table/index name, key condition, filter, ScannedCount, Count, consumed capacity and pagination token. A filter is applied after DynamoDB reads matching key data, so a small returned count does not imply low cost. If a new GSI is CREATING or backfilling, inspect index status, backfilling and throttling metrics; do not assume it is ready. For throttling on one tenant or date bucket, compare key frequency and item sizes to expose skew. For missing index fields, inspect projection configuration and remember that GSI propagation is eventually consistent.
Negative test: issue a query for an access pattern not represented by a table or index key. Reject a scan-based “solution,” add the proposed key/index to the design only, and calculate its write amplification, storage, backfill and hot-key consequences before creation.
Cost and cleanup
Index writes and storage, projected attributes, reads, backfills, and over-provisioned capacity increase cost.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Design partition keys, sort keys, GSIs, and LSIs from exact access patterns and distribution needs.
- Which scope or ownership boundary must be proved first?
Expected direction: The primary key places and orders items. A GSI can use different keys and scales separately. An LSI shares the partition key, is created with the table, and has item-collection constraints.
- What evidence is strong enough to accept the result?
Expected direction: Every query names its target key or index, key condition, projection, consistency need, expected item count, and hot-key risk.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid low-cardinality hot keys, write sharding without a read plan, ALL projection by habit, and LSIs when item collections may exceed their limits.
- Which cost dimensions and retained resources need an owner?
Expected direction: Index writes and storage, projected attributes, reads, backfills, and over-provisioned capacity increase cost.
Lesson acceptance
Pass when every required access pattern maps to an exact table/GSI/LSI key condition, uniqueness and adjacency semantics are explicit, sparse-index behavior and projection are demonstrated, and partition distribution is justified with sample cardinalities. The learner must compare GSI and LSI lifecycle, consistency, capacity, size and creation constraints; plan online GSI backfill and rollback; test pagination and empty results; and account for duplicated writes/storage. Fail if normal traffic needs a full scan, filters are mistaken for key selection, or a single popular key creates an unexplained hot partition.