Lesson 134 · AWS Learning Path

AWS 134: Amazon ElastiCache

· Published · 11 min read

Labelled process diagram for AWS 134: Application read to Cache key lookup to Hit or source database miss path to TTL, eviction, and hit-ratio evidence, with decision, proof and rejection evidence.

Why this lesson matters

Use a managed cache for a measured latency or load problem while preserving a correct source of truth.

A cache trades freshness, memory/capacity, and failure complexity for lower latency and reduced source load. ElastiCache manages cache infrastructure; the application still owns keys, TTL, invalidation, stampede prevention, serialization, fallback, and whether cached data may be lost.

What you will be able to do

By the end, you can:

  • explain amazon elasticache in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeUse a managed cache for a measured latency or load problem while preserving a correct source of truth.
Scope and boundaryElastiCache supports Valkey, Redis OSS, and Memcached choices with different data structures, persistence, replication, clustering, and failover behavior.
Evidence of successThe design defines cache-aside or write path, key and TTL policy, eviction, encryption, authentication, topology, hit ratio, fallback, and warm-up behavior.
Cost modelNode or serverless capacity, data processed, snapshots, transfer, Global Datastore, monitoring, and idle clusters can charge.
Safe rejection ruleAvoid storing the only copy of durable business data in a cache or using long TTLs without an invalidation and staleness plan.

How the request flows

+----------------------+
|   Application read   |
+----------------------+
           |
           v
+----------------------+
|   Cache key lookup   |
+----------------------+
           |
           v
+------------------------------------+
|  Hit or source database miss path  |
+------------------------------------+
                  |
                  v
+-----------------------------------------+
|  TTL, eviction, and hit-ratio evidence  |
+-----------------------------------------+

For Amazon ElastiCache, the important boundary is this: ElastiCache supports Valkey, Redis OSS, and Memcached choices with different data structures, persistence, replication, clustering, and failover behavior. The design defines cache-aside or write path, key and TTL policy, eviction, encryption, authentication, topology, hit ratio, fallback, and warm-up behavior. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.

Architecture decision table

SituationDirectionReason
Requirement matchesUse ElastiCache when measured repeated reads, session or rate-limit data, or supported data structures justify a managed low-latency cache.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid storing the only copy of durable business data in a cache or using long TTLs without an invalidation and staleness plan.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

Choose the engine and deployment model

ChoiceCapabilitiesMain fit/boundary
ValkeyOpen-source Redis-compatible data structures and commands with current ElastiCache feature directionPreferred for new compatible workloads after client/command/version testing
Redis OSSRedis OSS compatibility for supported versions and existing workloadsVerify licensing/version lifecycle, command/modules and migration to Valkey where appropriate
MemcachedSimple multithreaded distributed key-value cacheEasy horizontal node distribution; no replication/persistence/rich Redis structures
Serverless cacheManaged elastic capacity/data processing for Valkey, Redis OSS, or Memcached support according to current matrixVariable workloads; fixed parameter/topology control is reduced and limits/pricing differ
Node-based clusterExplicit node types, shards, replicas, AZs, parameter groups, reservationsPredictable steady workloads or features/configuration requiring topology control

MemoryDB is a different durable Valkey/Redis-compatible database service. ElastiCache is normally treated as a cache even when Valkey/Redis persistence/backups can aid recovery.

Cache patterns

Cache-aside

GET key -> hit: return
        -> miss: read source -> populate with TTL -> return
write source -> invalidate/update cache after authoritative commit

Protect against cache stampede with request coalescing/locks, jittered TTLs, stale-while-revalidate where business-safe, prewarming, and source backpressure. A cache outage must not create an uncontrolled database storm.

Write-through/write-behind

Write-through updates source and cache in a controlled order but needs failure compensation. Write-behind acknowledges before durable source update and can lose data; use only with an explicit queue/durability/replay design. Never invent durability by calling a cache “fast database.”

Session, counters, locks, and rate limits

Session data needs accepted loss/failover behavior and secure TTL. Atomic counters/rate limits need shard and expiry design. Distributed locks require token ownership, bounded lease, fencing token, clock/failure analysis, and safe unlock - not a bare SETNX assumption.

Valkey/Redis topology and endpoints

Cluster mode disabled uses one primary shard with optional read replicas; primary endpoint follows writes and reader endpoint distributes reads. Cluster mode enabled partitions hash slots across shards, each with primary/replicas. Clients must support cluster discovery/redirection and multi-key commands require compatible hash-slot design, often using hash tags deliberately.

Automatic failover promotes a replica when eligible. Existing connections/transactions/scripts can fail, replicas can lag, and cache data may be missing. Test reconnect and source fallback. Multi-AZ means replica placement/failover - not durable source-of-truth status.

Memcached clients discover node endpoints and distribute keys client-side. Adding/removing nodes changes key mapping unless consistent hashing is used and causes cache churn/misses. There is no replica promotion.

TTL, eviction, and memory

TTL is a freshness and memory policy. Distinguish no-expiry keys, application expiry, lazy/active expiration, and eviction under memory pressure. Valkey/Redis eviction policies include no-eviction and LRU/LFU/random/TTL variants over all keys or volatile keys; select from business semantics. noeviction rejects writes when full instead of preserving performance magically.

Monitor bytes used, fragmentation, swap (avoid), evictions, expirations, hit/miss ratio, key count, connections, CPU/engine CPU, network, command latency, replication lag, and rejected connections. Large keys and expensive commands can block shards and create latency. Use bounded scans rather than production KEYS *.

Data tiering on supported r6gd Valkey/Redis node clusters keeps keys in memory and moves less-used values to local SSD using LRU. It fits datasets with a small hot set and tolerable extra SSD access latency; it does not make SSD a durable database copy.

Persistence, backup, and global behavior

Valkey/Redis node-based options can use snapshot and append-only persistence behavior according to configuration. ElastiCache backups/snapshots and restore support vary by engine/deployment. Persistence can improve recovery but consumes resources and has RPO; test failover and restore.

Global Datastore provides cross-Region replication for supported Valkey/Redis node-based clusters, typically one primary Region and read-only secondary clusters. It is asynchronous; promotion/failover, DNS/client routing, lag, data loss, and failback need a runbook. It does not make every Region a simultaneous writer.

Security

ElastiCache is VPC-only. Use cache subnet groups across intended AZs, SG ingress from application SG, TLS in transit, at-rest encryption where supported, and Valkey/Redis authentication token or role-based access control/users/user groups. Store secrets securely and rotate through supported workflow.

IAM controls the ElastiCache API and supported IAM authentication options; cache commands still need engine identity/ACL semantics. Parameter groups can enable dangerous behavior; restrict administrative commands. Protect snapshots, logs, KMS keys, and network paths.

Performance and price

Use long-lived pooled connections, pipelining only with bounded response memory, cluster-aware clients, appropriate timeouts/retries, and serialization/compression measured against CPU. Do not retry non-idempotent list/counter operations blindly after timeout.

Node-based cost includes every node/hour, data tiering/storage where applicable, backups, transfer, Global Datastore, monitoring, and reserved commitments. Serverless includes capacity/data processed and storage dimensions according to current pricing. Calculate miss cost on the source and failure surge too.

Worked decisions

  • Read-heavy catalog: cache-aside with short jittered TTL and event invalidation; checkout revalidates source price/stock.
  • Login sessions: replicated Valkey with Multi-AZ only if session-loss semantics fit; encrypt/authenticate, TTL every session, and test failover.
  • Simple disposable object cache: Memcached can fit when no persistence/replication/rich structures are needed and clients handle node churn.
  • Steady large cold dataset: data tiering may reduce memory cost if hot set is small; benchmark SSD-hit latency.
  • Variable unpredictable cache: compare serverless with node-based reserved economics and feature/parameter constraints.

Required tests

  1. Measure source latency/load before cache, cold miss, warm hit, hit ratio, and total p99/cost.
  2. Update source through and around the normal write path; prove invalidation/staleness bounds.
  3. Expire/evict keys under bounded pressure and prove application fallback without stampede.
  4. Fail a primary/node or supplied dependency; measure reconnect, data loss, replica promotion, and source surge.
  5. Deny SG/TLS/auth user separately and identify each error without opening public access.
  6. Test a large key/blocking command and correct data/command design.
  7. Restore a snapshot or rebuild cache from source and document which is authoritative.
  8. Delete replication/global relationships, snapshots, cluster/serverless cache, users/groups, subnet/parameter groups, SG rules, alarms, and secrets in reviewed order.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open ElastiCache and identify Valkey or Redis OSS caches and Memcached caches separately.
  2. Inspect engine, deployment option, nodes or shards, replicas, AZs, encryption, security groups, endpoint, maintenance, and backups.
  3. Open CloudWatch metrics and compare cache hits, misses, evictions, memory, connections, and replication lag.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws elasticache describe-replication-groups --query 'ReplicationGroups[].{Id:ReplicationGroupId,Status:Status,Engine:Engine,Clustered:ClusterEnabled,AutoFailover:AutomaticFailover,Encrypted:AtRestEncryptionEnabled}' --output table
aws elasticache describe-cache-clusters --show-cache-node-info --query 'CacheClusters[].{Id:CacheClusterId,Engine:Engine,Status:CacheClusterStatus,Nodes:NumCacheNodes}' --output table

Expected interpretation

The inventory proves cache topology. It does not prove hit ratio, correct invalidation, source-of-truth safety, or application fallback during cache loss.

Practical work

Design cache-aside for a product catalog. Define keys, TTL, invalidation event, miss path, stampede control, failure bypass, hit-ratio target, and a stale-price safety rule.

Diagnose this topic from its own evidence

Begin with endpoint type, DNS, port, TLS/auth mode, security-group path and engine/client error. Then inspect connections, CPU, engine CPU, memory, evictions, hit ratio, replication lag, swap/network, command latency and node/slot health. A low hit ratio can mean poor key reuse, premature TTLs, oversized churn or wrong workload - not merely “add memory.” A cache stampede appears as simultaneous misses and backend load; mitigate with request coalescing, jittered TTLs, bounded stale serving or prewarming according to correctness needs.

Negative test: delete/expire a non-production cache key or use a disposable namespace and prove the application rebuilds it from the durable owner. Simulate an unavailable endpoint in the client and prove timeouts/circuit breaking prevent cache failure from exhausting the application.

Cost and cleanup

Node or serverless capacity, data processed, snapshots, transfer, Global Datastore, monitoring, and idle clusters can charge.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Use a managed cache for a measured latency or load problem while preserving a correct source of truth.

  1. Which scope or ownership boundary must be proved first?

Expected direction: ElastiCache supports Valkey, Redis OSS, and Memcached choices with different data structures, persistence, replication, clustering, and failover behavior.

  1. What evidence is strong enough to accept the result?

Expected direction: The design defines cache-aside or write path, key and TTL policy, eviction, encryption, authentication, topology, hit ratio, fallback, and warm-up behavior.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid storing the only copy of durable business data in a cache or using long TTLs without an invalidation and staleness plan.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Node or serverless capacity, data processed, snapshots, transfer, Global Datastore, monitoring, and idle clusters can charge.

Lesson acceptance

Pass when the learner selects Valkey/Redis OSS versus Memcached and serverless versus node-based from required data structures, durability, scaling and operations; identifies endpoint behavior in clustered and failover modes; and proves TTL, eviction and failure semantics. The design must keep authoritative data elsewhere unless persistence limitations are explicitly accepted, protect the data plane with private networking/TLS/auth, monitor saturation and hit quality, prevent stampedes, estimate minimum/transfer/backup costs and define cleanup. Fail if cache availability or persistence is confused with database durability.

Official sources

Advertisement