AWS 134: Amazon ElastiCache
Why this lesson matters
Use a managed cache for a measured latency or load problem while preserving a correct source of truth.
A cache trades freshness, memory/capacity, and failure complexity for lower latency and reduced source load. ElastiCache manages cache infrastructure; the application still owns keys, TTL, invalidation, stampede prevention, serialization, fallback, and whether cached data may be lost.
What you will be able to do
By the end, you can:
- explain amazon elasticache in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Use a managed cache for a measured latency or load problem while preserving a correct source of truth. |
| Scope and boundary | ElastiCache supports Valkey, Redis OSS, and Memcached choices with different data structures, persistence, replication, clustering, and failover behavior. |
| Evidence of success | The design defines cache-aside or write path, key and TTL policy, eviction, encryption, authentication, topology, hit ratio, fallback, and warm-up behavior. |
| Cost model | Node or serverless capacity, data processed, snapshots, transfer, Global Datastore, monitoring, and idle clusters can charge. |
| Safe rejection rule | Avoid storing the only copy of durable business data in a cache or using long TTLs without an invalidation and staleness plan. |
How the request flows
+----------------------+
| Application read |
+----------------------+
|
v
+----------------------+
| Cache key lookup |
+----------------------+
|
v
+------------------------------------+
| Hit or source database miss path |
+------------------------------------+
|
v
+-----------------------------------------+
| TTL, eviction, and hit-ratio evidence |
+-----------------------------------------+
For Amazon ElastiCache, the important boundary is this: ElastiCache supports Valkey, Redis OSS, and Memcached choices with different data structures, persistence, replication, clustering, and failover behavior. The design defines cache-aside or write path, key and TTL policy, eviction, encryption, authentication, topology, hit ratio, fallback, and warm-up behavior. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use ElastiCache when measured repeated reads, session or rate-limit data, or supported data structures justify a managed low-latency cache. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid storing the only copy of durable business data in a cache or using long TTLs without an invalidation and staleness plan. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
Choose the engine and deployment model
| Choice | Capabilities | Main fit/boundary |
|---|---|---|
| Valkey | Open-source Redis-compatible data structures and commands with current ElastiCache feature direction | Preferred for new compatible workloads after client/command/version testing |
| Redis OSS | Redis OSS compatibility for supported versions and existing workloads | Verify licensing/version lifecycle, command/modules and migration to Valkey where appropriate |
| Memcached | Simple multithreaded distributed key-value cache | Easy horizontal node distribution; no replication/persistence/rich Redis structures |
| Serverless cache | Managed elastic capacity/data processing for Valkey, Redis OSS, or Memcached support according to current matrix | Variable workloads; fixed parameter/topology control is reduced and limits/pricing differ |
| Node-based cluster | Explicit node types, shards, replicas, AZs, parameter groups, reservations | Predictable steady workloads or features/configuration requiring topology control |
MemoryDB is a different durable Valkey/Redis-compatible database service. ElastiCache is normally treated as a cache even when Valkey/Redis persistence/backups can aid recovery.
Cache patterns
Cache-aside
GET key -> hit: return
-> miss: read source -> populate with TTL -> return
write source -> invalidate/update cache after authoritative commit
Protect against cache stampede with request coalescing/locks, jittered TTLs, stale-while-revalidate where business-safe, prewarming, and source backpressure. A cache outage must not create an uncontrolled database storm.
Write-through/write-behind
Write-through updates source and cache in a controlled order but needs failure compensation. Write-behind acknowledges before durable source update and can lose data; use only with an explicit queue/durability/replay design. Never invent durability by calling a cache “fast database.”
Session, counters, locks, and rate limits
Session data needs accepted loss/failover behavior and secure TTL. Atomic counters/rate limits need shard and expiry design. Distributed locks require token ownership, bounded lease, fencing token, clock/failure analysis, and safe unlock - not a bare SETNX assumption.
Valkey/Redis topology and endpoints
Cluster mode disabled uses one primary shard with optional read replicas; primary endpoint follows writes and reader endpoint distributes reads. Cluster mode enabled partitions hash slots across shards, each with primary/replicas. Clients must support cluster discovery/redirection and multi-key commands require compatible hash-slot design, often using hash tags deliberately.
Automatic failover promotes a replica when eligible. Existing connections/transactions/scripts can fail, replicas can lag, and cache data may be missing. Test reconnect and source fallback. Multi-AZ means replica placement/failover - not durable source-of-truth status.
Memcached clients discover node endpoints and distribute keys client-side. Adding/removing nodes changes key mapping unless consistent hashing is used and causes cache churn/misses. There is no replica promotion.
TTL, eviction, and memory
TTL is a freshness and memory policy. Distinguish no-expiry keys, application expiry, lazy/active expiration, and eviction under memory pressure. Valkey/Redis eviction policies include no-eviction and LRU/LFU/random/TTL variants over all keys or volatile keys; select from business semantics. noeviction rejects writes when full instead of preserving performance magically.
Monitor bytes used, fragmentation, swap (avoid), evictions, expirations, hit/miss ratio, key count, connections, CPU/engine CPU, network, command latency, replication lag, and rejected connections. Large keys and expensive commands can block shards and create latency. Use bounded scans rather than production KEYS *.
Data tiering on supported r6gd Valkey/Redis node clusters keeps keys in memory and moves less-used values to local SSD using LRU. It fits datasets with a small hot set and tolerable extra SSD access latency; it does not make SSD a durable database copy.
Persistence, backup, and global behavior
Valkey/Redis node-based options can use snapshot and append-only persistence behavior according to configuration. ElastiCache backups/snapshots and restore support vary by engine/deployment. Persistence can improve recovery but consumes resources and has RPO; test failover and restore.
Global Datastore provides cross-Region replication for supported Valkey/Redis node-based clusters, typically one primary Region and read-only secondary clusters. It is asynchronous; promotion/failover, DNS/client routing, lag, data loss, and failback need a runbook. It does not make every Region a simultaneous writer.
Security
ElastiCache is VPC-only. Use cache subnet groups across intended AZs, SG ingress from application SG, TLS in transit, at-rest encryption where supported, and Valkey/Redis authentication token or role-based access control/users/user groups. Store secrets securely and rotate through supported workflow.
IAM controls the ElastiCache API and supported IAM authentication options; cache commands still need engine identity/ACL semantics. Parameter groups can enable dangerous behavior; restrict administrative commands. Protect snapshots, logs, KMS keys, and network paths.
Performance and price
Use long-lived pooled connections, pipelining only with bounded response memory, cluster-aware clients, appropriate timeouts/retries, and serialization/compression measured against CPU. Do not retry non-idempotent list/counter operations blindly after timeout.
Node-based cost includes every node/hour, data tiering/storage where applicable, backups, transfer, Global Datastore, monitoring, and reserved commitments. Serverless includes capacity/data processed and storage dimensions according to current pricing. Calculate miss cost on the source and failure surge too.
Worked decisions
- Read-heavy catalog: cache-aside with short jittered TTL and event invalidation; checkout revalidates source price/stock.
- Login sessions: replicated Valkey with Multi-AZ only if session-loss semantics fit; encrypt/authenticate, TTL every session, and test failover.
- Simple disposable object cache: Memcached can fit when no persistence/replication/rich structures are needed and clients handle node churn.
- Steady large cold dataset: data tiering may reduce memory cost if hot set is small; benchmark SSD-hit latency.
- Variable unpredictable cache: compare serverless with node-based reserved economics and feature/parameter constraints.
Required tests
- Measure source latency/load before cache, cold miss, warm hit, hit ratio, and total p99/cost.
- Update source through and around the normal write path; prove invalidation/staleness bounds.
- Expire/evict keys under bounded pressure and prove application fallback without stampede.
- Fail a primary/node or supplied dependency; measure reconnect, data loss, replica promotion, and source surge.
- Deny SG/TLS/auth user separately and identify each error without opening public access.
- Test a large key/blocking command and correct data/command design.
- Restore a snapshot or rebuild cache from source and document which is authoritative.
- Delete replication/global relationships, snapshots, cluster/serverless cache, users/groups, subnet/parameter groups, SG rules, alarms, and secrets in reviewed order.
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Open ElastiCache and identify Valkey or Redis OSS caches and Memcached caches separately.
- Inspect engine, deployment option, nodes or shards, replicas, AZs, encryption, security groups, endpoint, maintenance, and backups.
- Open CloudWatch metrics and compare cache hits, misses, evictions, memory, connections, and replication lag.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws elasticache describe-replication-groups --query 'ReplicationGroups[].{Id:ReplicationGroupId,Status:Status,Engine:Engine,Clustered:ClusterEnabled,AutoFailover:AutomaticFailover,Encrypted:AtRestEncryptionEnabled}' --output table
aws elasticache describe-cache-clusters --show-cache-node-info --query 'CacheClusters[].{Id:CacheClusterId,Engine:Engine,Status:CacheClusterStatus,Nodes:NumCacheNodes}' --output table
Expected interpretation
The inventory proves cache topology. It does not prove hit ratio, correct invalidation, source-of-truth safety, or application fallback during cache loss.
Practical work
Design cache-aside for a product catalog. Define keys, TTL, invalidation event, miss path, stampede control, failure bypass, hit-ratio target, and a stale-price safety rule.
Diagnose this topic from its own evidence
Begin with endpoint type, DNS, port, TLS/auth mode, security-group path and engine/client error. Then inspect connections, CPU, engine CPU, memory, evictions, hit ratio, replication lag, swap/network, command latency and node/slot health. A low hit ratio can mean poor key reuse, premature TTLs, oversized churn or wrong workload - not merely “add memory.” A cache stampede appears as simultaneous misses and backend load; mitigate with request coalescing, jittered TTLs, bounded stale serving or prewarming according to correctness needs.
Negative test: delete/expire a non-production cache key or use a disposable namespace and prove the application rebuilds it from the durable owner. Simulate an unavailable endpoint in the client and prove timeouts/circuit breaking prevent cache failure from exhausting the application.
Cost and cleanup
Node or serverless capacity, data processed, snapshots, transfer, Global Datastore, monitoring, and idle clusters can charge.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Use a managed cache for a measured latency or load problem while preserving a correct source of truth.
- Which scope or ownership boundary must be proved first?
Expected direction: ElastiCache supports Valkey, Redis OSS, and Memcached choices with different data structures, persistence, replication, clustering, and failover behavior.
- What evidence is strong enough to accept the result?
Expected direction: The design defines cache-aside or write path, key and TTL policy, eviction, encryption, authentication, topology, hit ratio, fallback, and warm-up behavior.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid storing the only copy of durable business data in a cache or using long TTLs without an invalidation and staleness plan.
- Which cost dimensions and retained resources need an owner?
Expected direction: Node or serverless capacity, data processed, snapshots, transfer, Global Datastore, monitoring, and idle clusters can charge.
Lesson acceptance
Pass when the learner selects Valkey/Redis OSS versus Memcached and serverless versus node-based from required data structures, durability, scaling and operations; identifies endpoint behavior in clustered and failover modes; and proves TTL, eviction and failure semantics. The design must keep authoritative data elsewhere unless persistence limitations are explicitly accepted, protect the data plane with private networking/TLS/auth, monitor saturation and hit quality, prevent stampedes, estimate minimum/transfer/backup costs and define cleanup. Fail if cache availability or persistence is confused with database durability.