AWS 124: RDS Multi-AZ deployments
Why this lesson matters
Separate database availability from read scaling by understanding current Multi-AZ deployment choices.
“Multi-AZ” names two materially different RDS deployment architectures. Learners must identify which one is running, how writes are replicated, whether standby members can serve reads, how endpoints change, and what the application experiences during failover.
What you will be able to do
By the end, you can:
- explain rds multi-az deployments in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Separate database availability from read scaling by understanding current Multi-AZ deployment choices. |
| Scope and boundary | A Multi-AZ DB instance has a synchronous standby that is not a read endpoint. A Multi-AZ DB cluster has a writer and readable instances across three AZs for supported engines. |
| Evidence of success | A failure test shows endpoint behavior, failover timing, transaction impact, client retry behavior, and the difference between infrastructure recovery and application recovery. |
| Cost model | Additional instances, storage, I/O, cross-service transfer, and monitoring are billed even when a standby is not serving application reads. |
| Safe rejection rule | Do not present a classic standby as a read-scaling target or use Multi-AZ as a substitute for backups and tested restore. |
How the request flows
+----------------------+
| Writer transaction |
+----------------------+
|
v
+----------------------------------+
| Synchronous copy in another AZ |
+----------------------------------+
|
v
+-----------------------------------------+
| Failure detection and endpoint change |
+-----------------------------------------+
|
v
+-------------------------------------+
| Client reconnect and verification |
+-------------------------------------+
For RDS Multi-AZ, the important boundary is this: A Multi-AZ DB instance has a synchronous standby that is not a read endpoint. A Multi-AZ DB cluster has a writer and readable instances across three AZs for supported engines. A failure test shows endpoint behavior, failover timing, transaction impact, client retry behavior, and the difference between infrastructure recovery and application recovery. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use Multi-AZ for production availability when the database must survive an infrastructure or AZ problem with managed failover. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Do not present a classic standby as a read-scaling target or use Multi-AZ as a substitute for backups and tested restore. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
Two Multi-AZ architectures
| Property | Multi-AZ DB instance deployment | Multi-AZ DB cluster deployment |
|---|---|---|
| Topology | One primary DB instance plus one standby in another AZ | One writer plus two readable replicas in three AZs |
| Replication | Synchronous standby replication managed by RDS/engine architecture | Semisynchronous native-engine replication; commit needs acknowledgement from at least one reader, not necessarily execution on both |
| Read traffic | Standby is not an application read endpoint | Reader instances can serve reads through reader endpoint |
| Failover | Standby becomes primary and endpoint DNS is updated | Most up-to-date eligible reader is promoted and writer endpoint follows it |
| Resource presentation | One DB instance row with Multi-AZ enabled | Cluster row plus three instance rows/roles |
| Engine/version support | Broader but engine-specific | Supported RDS MySQL/PostgreSQL versions/classes/Regions only; verify current matrix |
| Cost | Primary and standby capacity/storage; standby is paid even without reads | Three instance capacities plus cluster storage/I/O/monitoring |
Neither is Aurora. An RDS Multi-AZ DB cluster uses engine-native replication and separate local instance storage architecture; an Aurora cluster uses shared distributed Aurora cluster storage. Do not merge diagrams or failover claims.
Commit and failover paths
DB instance deployment:
client -> stable DB endpoint -> primary (AZ-a) ==synchronous==> standby (AZ-b)
failure -> RDS promotes standby -> endpoint DNS changes
DB cluster deployment:
client writer -> writer endpoint -> writer (AZ-a)
| semisynchronous replication
+-> readable replica (AZ-b)
+-> readable replica (AZ-c)
client readers -> reader endpoint -----^ new connections distributed to readers
Infrastructure failover interrupts connections. An in-flight transaction can have an ambiguous result from the client's viewpoint: the server might have committed before the connection broke. Retrying a payment/order blindly can duplicate it. Use idempotency keys, transaction status checks, bounded backoff/jitter, DNS refresh, short connection timeouts, and a pool that discards failed sessions.
Failover triggers include AZ/host/storage/network failure, instance class changes, certain maintenance, reboot with failover, or a manual test. The endpoint name usually remains but its resolved address changes. Clients that cache DNS too long or pin IP/instance endpoints defeat managed failover.
Availability is not read scaling or backup
The standby in a Multi-AZ DB instance deployment cannot be used for normal reads. Add a read replica for read scaling, accepting asynchronous lag and separate promotion behavior. The two readers in a Multi-AZ DB cluster can serve reads, but lag and read consistency remain application concerns.
Multi-AZ protects service availability from defined infrastructure failures. It does not restore a table dropped by an authorized user, recover yesterday's data, provide cross-Region DR, or guarantee application health. Automated backups/PITR, snapshots/copies, validation, and potentially cross-Region replicas remain separate.
Design and build requirements
- Use a DB subnet group spanning suitable subnets/AZs with adequate free IP addresses.
- Permit only approved clients through SGs and preserve DNS/return paths.
- Size failover capacity for production load; a recovery target must not be immediately saturated.
- Set backup and maintenance windows so they do not conflict with business peaks; monitor events and pending maintenance.
- Use TLS and database least privilege; protect secret/KMS dependencies across all AZ paths.
- Alarm on CPU, memory, free storage, connections, replica lag for DB clusters, transaction/storage latency, and failover events.
- Test application reconnection from every client platform/driver and connection pool.
- Represent topology and settings as IaC and review modifications that can cause failover/downtime.
Controlled failover test
- Establish baseline: topology, roles/AZs, endpoint DNS, database marker row, open long transaction, and p95 transaction latency.
- Start a continuous idempotent transaction/read probe that logs UTC, operation ID, connection/member identity, and result.
- Invoke only an approved reboot-with-failover or cluster failover during a maintenance exercise.
- Capture RDS events, DNS answers, disconnect errors, ambiguous operations, new writer, and time until successful validated transaction.
- Reconcile every operation ID and prove no duplicate/lost committed business effect.
- Compare observed interruption with RTO/SLO. Correct client DNS/retry/pool or topology capacity and rerun.
Do not test by terminating infrastructure outside the supported RDS operation, changing production SGs, or creating uncontrolled load.
Worked decisions
OLTP requiring HA but no read scaling
A Multi-AZ DB instance can fit when the supported engine and failover behavior meet RTO. The standby cost buys availability, not read throughput. Add backups and restore tests separately.
MySQL application needing HA and two readable members
Evaluate a Multi-AZ DB cluster. Route writes and consistency-sensitive reads to writer, stale-tolerant reads to reader endpoint, monitor lag, and test driver recovery. Compare Aurora based on compatibility, performance, storage, feature, and price - not the word “cluster.”
Cross-Region disaster recovery
Multi-AZ alone fails the requirement because all members are in one Region. Add supported cross-Region read replica/snapshot-copy/backup strategy or evaluate Aurora Global Database according to RPO/RTO.
Reporting query overload
Do not expect a classic standby to help. Add a read replica or analytics destination, then govern lag and query load. On a Multi-AZ DB cluster, use reader endpoint but still diagnose query/index problems.
Failure diagnosis
| Symptom | Prove | Correct boundary |
|---|---|---|
| App offline after RDS says available | DNS cache, pool sessions, retry logs, SG/secret/TLS, new writer transaction | Client recovery/configuration |
| Failover unusually slow | event timeline, replica lag, recovery target health/capacity, long transactions | Database/topology/client contributor |
| Read fails on standby | deployment type and endpoint | Standby is not readable; use proper replica architecture |
| Duplicate transaction after retry | operation IDs, DB commit, client timeout | Idempotency/reconciliation, not infrastructure rollback |
| DB cluster reader stale | ReplicaLag, transaction time, endpoint member | Consistency routing or lag reduction |
| AZ resilience not proven | subnet group versus actual member AZs | Place/verify members across distinct AZs |
Acceptance evidence
The learner must identify the exact Multi-AZ type, diagram its replication/endpoints, run a positive transaction, preserve an unauthorized/network denial, execute or analyze a controlled failover, reconcile an ambiguous transaction, restore a backup separately, calculate all paid members/storage/I/O, and show final retained-state ownership.
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Open RDS Databases and inspect Multi-AZ and Availability Zone fields.
- For a supplied Multi-AZ instance and cluster, compare endpoints, members, roles, and failover configuration.
- Open Events and Monitoring and identify the evidence that would timestamp a failover.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws rds describe-db-instances --query 'DBInstances[].{Id:DBInstanceIdentifier,MultiAZ:MultiAZ,PrimaryAZ:AvailabilityZone,SecondaryAZ:SecondaryAvailabilityZone,Endpoint:Endpoint.Address}' --output table
aws rds describe-db-clusters --query 'DBClusters[].{Id:DBClusterIdentifier,Engine:Engine,MultiAZ:MultiAZ,Members:DBClusterMembers[].DBInstanceIdentifier}' --output json
Expected interpretation
MultiAZ true describes deployment configuration. It does not prove an application reconnects, retries safely, or meets a measured recovery objective.
Practical work
Draw the writer, standby or readers, DNS endpoint, application pool, synchronous replication, failover event, and reconnect path. Set an RTO hypothesis and list the evidence needed to test it.
Diagnose this topic from its own evidence
Use the failure table and preserve RDS events, DNS answers, client errors, transaction IDs, member roles/AZs, lag, and timestamps. Separate RDS control-plane recovery from client reconnection and business-transaction reconciliation.
Cost and cleanup
Additional instances, storage, I/O, cross-service transfer, and monitoring are billed even when a standby is not serving application reads.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Separate database availability from read scaling by understanding current Multi-AZ deployment choices.
- Which scope or ownership boundary must be proved first?
Expected direction: A Multi-AZ DB instance has a synchronous standby that is not a read endpoint. A Multi-AZ DB cluster has a writer and readable instances across three AZs for supported engines.
- What evidence is strong enough to accept the result?
Expected direction: A failure test shows endpoint behavior, failover timing, transaction impact, client retry behavior, and the difference between infrastructure recovery and application recovery.
- Which tempting design or shortcut must be rejected?
Expected direction: Do not present a classic standby as a read-scaling target or use Multi-AZ as a substitute for backups and tested restore.
- Which cost dimensions and retained resources need an owner?
Expected direction: Additional instances, storage, I/O, cross-service transfer, and monitoring are billed even when a standby is not serving application reads.
Lesson acceptance
Pass only when the learner distinguishes DB instance from DB cluster Multi-AZ, proves standby/read behavior, executes or analyzes a controlled failover, measures interruption, reconciles ambiguous transactions, tests backup recovery separately, and documents all paid resources.