Lesson 125 · AWS Learning Path

AWS 125: RDS read replicas

· Published · 10 min read

Labelled process diagram for AWS 125: Primary commit to Asynchronous replication to Replica endpoint read to Lag alarm or controlled promotion, with decision, proof and rejection evidence.

Why this lesson matters

Use read replicas for read scaling and selected disaster-recovery patterns without confusing asynchronous replication with Multi-AZ failover.

Read replicas are independently addressed database copies fed asynchronously from a source. They can scale eligible reads and support migration/recovery patterns, but the application must own endpoint routing, stale data, lag, replication errors, and promotion.

What you will be able to do

By the end, you can:

  • explain rds read replicas in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeUse read replicas for read scaling and selected disaster-recovery patterns without confusing asynchronous replication with Multi-AZ failover.
Scope and boundaryReplicas have separate endpoints and can be in another AZ or Region where supported. Replication lag and engine behavior affect read-after-write results and promotion.
Evidence of successThe application sends only safe reads to replica endpoints, measures lag, handles stale data, and has a tested promotion and DNS or connection change plan.
Cost modelReplica instance hours, storage, I/O, monitoring, and cross-Region data transfer can charge independently of the source.
Safe rejection ruleAvoid sending writes to a read-only replica or promising zero data loss when replication lag can exist.

How the request flows

+----------------------+
|    Primary commit    |
+----------------------+
           |
           v
+----------------------------+
|  Asynchronous replication  |
+----------------------------+
              |
              v
+-------------------------+
|  Replica endpoint read  |
+-------------------------+
            |
            v
+-------------------------------------+
|  Lag alarm or controlled promotion  |
+-------------------------------------+

For RDS read replicas, the important boundary is this: Replicas have separate endpoints and can be in another AZ or Region where supported. Replication lag and engine behavior affect read-after-write results and promotion. The application sends only safe reads to replica endpoints, measures lag, handles stale data, and has a tested promotion and DNS or connection change plan. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.

Architecture decision table

SituationDirectionReason
Requirement matchesUse replicas for scalable read traffic, reporting isolation, or a documented cross-Region recovery design that tolerates asynchronous replication.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid sending writes to a read-only replica or promising zero data loss when replication lag can exist.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

Read replica data path

application writes -> source endpoint -> source commit/log
                                           |
                                           | asynchronous engine/service replication
                                           v
read application -> replica endpoint -> apply stream -> replica data
                                           |
                                           +-> lag/errors/CloudWatch events
                                           +-> optional independent promotion

The source can acknowledge a commit before the replica applies it. Therefore a replica read can be stale even when both resources are available. Replication mechanisms, supported number/topology, cascading behavior, storage, engine version, encryption, backups, and cross-Region behavior vary by engine; check the exact current matrix.

A read replica has its own endpoint, instance class, storage, parameters, maintenance, SGs, metrics, and bill. It is not the hidden Multi-AZ standby. A source can be Multi-AZ while also having read replicas; those solve separate availability and scaling requirements.

Application read routing

Classify every query:

  • Writer-only: INSERT/UPDATE/DELETE, transactions, locks, DDL, and consistency-sensitive reads.
  • Replica-safe: stale-tolerant catalog, reporting, historical dashboards, or fan-out reads within a defined lag budget.
  • Session-sensitive: read-after-write, monotonic reads, temporary/session tables, or functions requiring the same connection state.

RDS does not automatically rewrite application connections from source to replicas. Use separate connection pools/endpoints and explicit query routing. A load balancer designed for HTTP is not placed in front of a database. DNS-based application routing or a database-aware proxy must preserve protocol/session behavior and be tested.

Read-after-write choices include reading from writer for a bounded interval, tracking commit position where engine/application supports it, returning the committed value from the write API, or accepting documented staleness. Sleeping an arbitrary number of seconds is not a correctness guarantee.

Lag, failure, and source impact

Measure engine-specific replica lag metrics plus application marker lag. A metric can be missing, -1, stale, or semantically different during error. Causes include:

  • long-running transactions delaying log availability/apply;
  • write bursts or large DDL;
  • replica CPU, memory, storage IOPS/throughput, locks, or query contention;
  • network/cross-Region transfer interruption;
  • incompatible schema/engine settings or replication error;
  • insufficient source log retention/storage;
  • maintenance/reboot or replica stopped state.

A heavily loaded replica can fall behind. Scale it, tune/limit reporting, fix indexes/queries, or add replicas according to evidence. Source replication also consumes resources; monitor source logs/storage and transaction performance.

Set a lag SLO, for example “99.9% of replica reads observe commits within 10 seconds.” Alarm before the business stale-data limit, remove unhealthy replicas from routing, and define catch-up/rebuild steps.

Cross-Region replicas

Cross-Region replicas can place read data near users and provide a recovery candidate. Replication remains asynchronous and incurs destination compute/storage, data transfer/replication, KMS, backup, and monitoring charges. Encryption requires supported encrypted-copy behavior and destination-Region KMS design.

Promotion is not managed transparent Multi-AZ failover. Promoting normally breaks replication and creates an independent writable database with a different endpoint. Applications, secrets, DNS/configuration, dependent replicas, backup policy, monitoring, and writes must be redirected. RPO equals unreplicated/apply lag plus ambiguous transactions; RTO includes decision, promotion, routing, validation, and reconciliation.

Promotion runbook

  1. Declare planned migration or incident; identify authoritative source and change authority.
  2. Quiesce/fence source writes when possible for zero-loss planned promotion.
  3. Record source commit position/time, replica lag, errors, health, capacity, backups, and last marker.
  4. Wait for lag to reach acceptable threshold; stop if RPO would be violated.
  5. Promote the exact replica with a bounded waiter.
  6. Apply independent backup, maintenance, parameter, monitoring, secret, SG, and deletion-protection settings as required.
  7. Redirect clients through controlled configuration/DNS, discard old pools, and run positive/negative transactions.
  8. Reconcile ambiguous or missing operations. Prevent writes to the old source.
  9. Establish a new replication/DR topology and test failback; promotion does not automatically reverse replication.

Worked examples

Reporting isolation

Create a sized same-Region replica and route only reporting credentials/queries to its endpoint. Limit query concurrency, monitor lag and source impact, and remove it from routing above the stale-data threshold. Multi-AZ remains required separately if source availability demands it.

Global product catalog

Use cross-Region replicas for local stale-tolerant catalog reads. Checkout retrieves authoritative price/inventory from the writer Region. If the business requires local writes, ordinary RDS read replicas are not active-active; redesign or evaluate a service built for that topology.

Version upgrade migration

Where supported, use a replica/blue-green approach to test target version, monitor replication, quiesce writes, catch up, promote/switchover, and redirect. Snapshot before change and define rollback data reconciliation because new writes on the target prevent a simple DNS reversal.

Disaster recovery

A cross-Region replica can lower RTO versus snapshot restore, but still has nonzero RPO and continuous cost. Test promotion with marker transactions and infrastructure dependencies. If zero data loss is mandatory across Region loss, this asynchronous design cannot promise it.

Required tests

  1. Write timestamped unique markers to source and poll replica; measure distribution of apply lag.
  2. Route a stale-tolerant read successfully, then show why an immediate consistency-sensitive read uses writer.
  3. Attempt a write using a replica-only role/endpoint and preserve expected rejection.
  4. Apply bounded reporting load, observe lag/resources, make one tuning/capacity change, and repeat.
  5. Inject/analyze a replication dependency failure, preserve engine/RDS events, and rebuild only if correction/catch-up is unsafe.
  6. Perform an approved promotion, redirect a test client, prove write/read, measure RPO/RTO, and establish the new backup/topology.
  7. Delete only owned replicas after proving no application/DNS/DR dependency remains; retained promoted resources need an owner and cost review.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open RDS Databases and inspect the Replication section for an approved source.
  2. Compare source and replica endpoints, Regions, classes, encryption, and status.
  3. Open CloudWatch metrics and locate ReplicaLag or the engine-specific lag signal.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws rds describe-db-instances --query 'DBInstances[?ReadReplicaSourceDBInstanceIdentifier!=null].{Replica:DBInstanceIdentifier,Source:ReadReplicaSourceDBInstanceIdentifier,Status:DBInstanceStatus,Endpoint:Endpoint.Address}' --output table
aws rds describe-db-instances --query 'DBInstances[].{Id:DBInstanceIdentifier,Replicas:ReadReplicaDBInstanceIdentifiers}' --output json

Expected interpretation

A replica listed as available is not proof of zero lag or safe read routing. Promotion creates an independent database and does not automatically redirect clients.

Practical work

Design read routing for catalog searches and order-status reads. Mark which reads may be stale, define a lag alarm, and document promotion, endpoint change, and split-brain prevention.

Diagnose this topic from its own evidence

Start with source commit time/position, replica apply position, lag metric semantics, replica resource saturation, engine replication error, network/Region path, and application endpoint. Remove a replica from reads when its staleness exceeds the contract; do not promote it merely to clear an alarm.

Cost and cleanup

Replica instance hours, storage, I/O, monitoring, and cross-Region data transfer can charge independently of the source.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Use read replicas for read scaling and selected disaster-recovery patterns without confusing asynchronous replication with Multi-AZ failover.

  1. Which scope or ownership boundary must be proved first?

Expected direction: Replicas have separate endpoints and can be in another AZ or Region where supported. Replication lag and engine behavior affect read-after-write results and promotion.

  1. What evidence is strong enough to accept the result?

Expected direction: The application sends only safe reads to replica endpoints, measures lag, handles stale data, and has a tested promotion and DNS or connection change plan.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid sending writes to a read-only replica or promising zero data loss when replication lag can exist.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Replica instance hours, storage, I/O, monitoring, and cross-Region data transfer can charge independently of the source.

Lesson acceptance

Pass with explicit writer/replica query routing, measured lag, stale-read policy, read-only denial, source-impact evidence, one repaired replication failure, controlled promotion with RPO/RTO and reconciliation, new topology/backup plan, cost, and cleanup.

Official sources

Advertisement