Lesson 124 · AWS Learning Path

AWS 124: RDS Multi-AZ deployments

· Published · 10 min read

Labelled process diagram for AWS 124: Writer transaction to Synchronous copy in another AZ to Failure detection and endpoint change to Client reconnect and verification, with decision, proof and rejection evidence.

Why this lesson matters

Separate database availability from read scaling by understanding current Multi-AZ deployment choices.

“Multi-AZ” names two materially different RDS deployment architectures. Learners must identify which one is running, how writes are replicated, whether standby members can serve reads, how endpoints change, and what the application experiences during failover.

What you will be able to do

By the end, you can:

  • explain rds multi-az deployments in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeSeparate database availability from read scaling by understanding current Multi-AZ deployment choices.
Scope and boundaryA Multi-AZ DB instance has a synchronous standby that is not a read endpoint. A Multi-AZ DB cluster has a writer and readable instances across three AZs for supported engines.
Evidence of successA failure test shows endpoint behavior, failover timing, transaction impact, client retry behavior, and the difference between infrastructure recovery and application recovery.
Cost modelAdditional instances, storage, I/O, cross-service transfer, and monitoring are billed even when a standby is not serving application reads.
Safe rejection ruleDo not present a classic standby as a read-scaling target or use Multi-AZ as a substitute for backups and tested restore.

How the request flows

+----------------------+
|  Writer transaction  |
+----------------------+
           |
           v
+----------------------------------+
|  Synchronous copy in another AZ  |
+----------------------------------+
                 |
                 v
+-----------------------------------------+
|  Failure detection and endpoint change  |
+-----------------------------------------+
                    |
                    v
+-------------------------------------+
|  Client reconnect and verification  |
+-------------------------------------+

For RDS Multi-AZ, the important boundary is this: A Multi-AZ DB instance has a synchronous standby that is not a read endpoint. A Multi-AZ DB cluster has a writer and readable instances across three AZs for supported engines. A failure test shows endpoint behavior, failover timing, transaction impact, client retry behavior, and the difference between infrastructure recovery and application recovery. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.

Architecture decision table

SituationDirectionReason
Requirement matchesUse Multi-AZ for production availability when the database must survive an infrastructure or AZ problem with managed failover.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchDo not present a classic standby as a read-scaling target or use Multi-AZ as a substitute for backups and tested restore.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

Two Multi-AZ architectures

PropertyMulti-AZ DB instance deploymentMulti-AZ DB cluster deployment
TopologyOne primary DB instance plus one standby in another AZOne writer plus two readable replicas in three AZs
ReplicationSynchronous standby replication managed by RDS/engine architectureSemisynchronous native-engine replication; commit needs acknowledgement from at least one reader, not necessarily execution on both
Read trafficStandby is not an application read endpointReader instances can serve reads through reader endpoint
FailoverStandby becomes primary and endpoint DNS is updatedMost up-to-date eligible reader is promoted and writer endpoint follows it
Resource presentationOne DB instance row with Multi-AZ enabledCluster row plus three instance rows/roles
Engine/version supportBroader but engine-specificSupported RDS MySQL/PostgreSQL versions/classes/Regions only; verify current matrix
CostPrimary and standby capacity/storage; standby is paid even without readsThree instance capacities plus cluster storage/I/O/monitoring

Neither is Aurora. An RDS Multi-AZ DB cluster uses engine-native replication and separate local instance storage architecture; an Aurora cluster uses shared distributed Aurora cluster storage. Do not merge diagrams or failover claims.

Commit and failover paths

DB instance deployment:
client -> stable DB endpoint -> primary (AZ-a) ==synchronous==> standby (AZ-b)
                                  failure -> RDS promotes standby -> endpoint DNS changes

DB cluster deployment:
client writer -> writer endpoint -> writer (AZ-a)
                                      | semisynchronous replication
                                      +-> readable replica (AZ-b)
                                      +-> readable replica (AZ-c)
client readers -> reader endpoint -----^ new connections distributed to readers

Infrastructure failover interrupts connections. An in-flight transaction can have an ambiguous result from the client's viewpoint: the server might have committed before the connection broke. Retrying a payment/order blindly can duplicate it. Use idempotency keys, transaction status checks, bounded backoff/jitter, DNS refresh, short connection timeouts, and a pool that discards failed sessions.

Failover triggers include AZ/host/storage/network failure, instance class changes, certain maintenance, reboot with failover, or a manual test. The endpoint name usually remains but its resolved address changes. Clients that cache DNS too long or pin IP/instance endpoints defeat managed failover.

Availability is not read scaling or backup

The standby in a Multi-AZ DB instance deployment cannot be used for normal reads. Add a read replica for read scaling, accepting asynchronous lag and separate promotion behavior. The two readers in a Multi-AZ DB cluster can serve reads, but lag and read consistency remain application concerns.

Multi-AZ protects service availability from defined infrastructure failures. It does not restore a table dropped by an authorized user, recover yesterday's data, provide cross-Region DR, or guarantee application health. Automated backups/PITR, snapshots/copies, validation, and potentially cross-Region replicas remain separate.

Design and build requirements

  • Use a DB subnet group spanning suitable subnets/AZs with adequate free IP addresses.
  • Permit only approved clients through SGs and preserve DNS/return paths.
  • Size failover capacity for production load; a recovery target must not be immediately saturated.
  • Set backup and maintenance windows so they do not conflict with business peaks; monitor events and pending maintenance.
  • Use TLS and database least privilege; protect secret/KMS dependencies across all AZ paths.
  • Alarm on CPU, memory, free storage, connections, replica lag for DB clusters, transaction/storage latency, and failover events.
  • Test application reconnection from every client platform/driver and connection pool.
  • Represent topology and settings as IaC and review modifications that can cause failover/downtime.

Controlled failover test

  1. Establish baseline: topology, roles/AZs, endpoint DNS, database marker row, open long transaction, and p95 transaction latency.
  2. Start a continuous idempotent transaction/read probe that logs UTC, operation ID, connection/member identity, and result.
  3. Invoke only an approved reboot-with-failover or cluster failover during a maintenance exercise.
  4. Capture RDS events, DNS answers, disconnect errors, ambiguous operations, new writer, and time until successful validated transaction.
  5. Reconcile every operation ID and prove no duplicate/lost committed business effect.
  6. Compare observed interruption with RTO/SLO. Correct client DNS/retry/pool or topology capacity and rerun.

Do not test by terminating infrastructure outside the supported RDS operation, changing production SGs, or creating uncontrolled load.

Worked decisions

OLTP requiring HA but no read scaling

A Multi-AZ DB instance can fit when the supported engine and failover behavior meet RTO. The standby cost buys availability, not read throughput. Add backups and restore tests separately.

MySQL application needing HA and two readable members

Evaluate a Multi-AZ DB cluster. Route writes and consistency-sensitive reads to writer, stale-tolerant reads to reader endpoint, monitor lag, and test driver recovery. Compare Aurora based on compatibility, performance, storage, feature, and price - not the word “cluster.”

Cross-Region disaster recovery

Multi-AZ alone fails the requirement because all members are in one Region. Add supported cross-Region read replica/snapshot-copy/backup strategy or evaluate Aurora Global Database according to RPO/RTO.

Reporting query overload

Do not expect a classic standby to help. Add a read replica or analytics destination, then govern lag and query load. On a Multi-AZ DB cluster, use reader endpoint but still diagnose query/index problems.

Failure diagnosis

SymptomProveCorrect boundary
App offline after RDS says availableDNS cache, pool sessions, retry logs, SG/secret/TLS, new writer transactionClient recovery/configuration
Failover unusually slowevent timeline, replica lag, recovery target health/capacity, long transactionsDatabase/topology/client contributor
Read fails on standbydeployment type and endpointStandby is not readable; use proper replica architecture
Duplicate transaction after retryoperation IDs, DB commit, client timeoutIdempotency/reconciliation, not infrastructure rollback
DB cluster reader staleReplicaLag, transaction time, endpoint memberConsistency routing or lag reduction
AZ resilience not provensubnet group versus actual member AZsPlace/verify members across distinct AZs

Acceptance evidence

The learner must identify the exact Multi-AZ type, diagram its replication/endpoints, run a positive transaction, preserve an unauthorized/network denial, execute or analyze a controlled failover, reconcile an ambiguous transaction, restore a backup separately, calculate all paid members/storage/I/O, and show final retained-state ownership.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open RDS Databases and inspect Multi-AZ and Availability Zone fields.
  2. For a supplied Multi-AZ instance and cluster, compare endpoints, members, roles, and failover configuration.
  3. Open Events and Monitoring and identify the evidence that would timestamp a failover.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws rds describe-db-instances --query 'DBInstances[].{Id:DBInstanceIdentifier,MultiAZ:MultiAZ,PrimaryAZ:AvailabilityZone,SecondaryAZ:SecondaryAvailabilityZone,Endpoint:Endpoint.Address}' --output table
aws rds describe-db-clusters --query 'DBClusters[].{Id:DBClusterIdentifier,Engine:Engine,MultiAZ:MultiAZ,Members:DBClusterMembers[].DBInstanceIdentifier}' --output json

Expected interpretation

MultiAZ true describes deployment configuration. It does not prove an application reconnects, retries safely, or meets a measured recovery objective.

Practical work

Draw the writer, standby or readers, DNS endpoint, application pool, synchronous replication, failover event, and reconnect path. Set an RTO hypothesis and list the evidence needed to test it.

Diagnose this topic from its own evidence

Use the failure table and preserve RDS events, DNS answers, client errors, transaction IDs, member roles/AZs, lag, and timestamps. Separate RDS control-plane recovery from client reconnection and business-transaction reconciliation.

Cost and cleanup

Additional instances, storage, I/O, cross-service transfer, and monitoring are billed even when a standby is not serving application reads.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Separate database availability from read scaling by understanding current Multi-AZ deployment choices.

  1. Which scope or ownership boundary must be proved first?

Expected direction: A Multi-AZ DB instance has a synchronous standby that is not a read endpoint. A Multi-AZ DB cluster has a writer and readable instances across three AZs for supported engines.

  1. What evidence is strong enough to accept the result?

Expected direction: A failure test shows endpoint behavior, failover timing, transaction impact, client retry behavior, and the difference between infrastructure recovery and application recovery.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Do not present a classic standby as a read-scaling target or use Multi-AZ as a substitute for backups and tested restore.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Additional instances, storage, I/O, cross-service transfer, and monitoring are billed even when a standby is not serving application reads.

Lesson acceptance

Pass only when the learner distinguishes DB instance from DB cluster Multi-AZ, proves standby/read behavior, executes or analyzes a controlled failover, measures interruption, reconciles ambiguous transactions, tests backup recovery separately, and documents all paid resources.

Official sources

Advertisement