Lesson 234 · AWS Learning Path

AWS 234: RDS and DynamoDB backup, recovery, and failover operations

· Published · 12 min read

Labelled process diagram for AWS 234: Data failure and recovery point to New RDS or DynamoDB restore target to Integrity and application test to Controlled cutover and cleanup, with decision, proof and rejection...

Why this lesson matters

Availability and historical recovery solve different failures. RDS Multi-AZ, Aurora replicas and DynamoDB Global Tables can keep serving after infrastructure or regional failure. They can also replicate a bad update, accidental deletion or malicious write. Backups and point-in-time recovery preserve older state but restore to a new target that must be secured, validated and cut over.

The architect's job is not “press Restore.” It is choosing a point before the fault, proving keys and dependencies, restoring into isolation, validating business invariants, reconciling legitimate later writes exactly once, fencing writers, controlling cutover/rollback, and proving temporary-resource cleanup.

Outcomes

You will be able to:

  • distinguish RDS automated backups, transaction logs, manual snapshots, retained

automated backups, AWS Backup points and snapshot exports;

  • distinguish Single-AZ, Multi-AZ DB instance, Multi-AZ DB cluster, read replica,

Aurora Replica and Aurora Global Database behavior;

  • calculate actual RPO from earliest/latest restorable times and RTO end to end;
  • plan RDS/Aurora snapshot and PITR restoration to a new endpoint;
  • distinguish DynamoDB PITR, on-demand backup, AWS Backup, export/import, Streams

and Global Tables;

  • identify every DynamoDB setting not recreated automatically after restore;
  • explain Global Table replication/conflicts and why it is not a backup;
  • validate schema, indexes, data invariants, security, performance and application;
  • design idempotent delta replay, controlled cutover, rollback and cleanup;
  • diagnose role, KMS, quota, topology, DNS, replication and restore failures.

Safety boundary and workbook

  • This lesson is no-create. Commands are read-only. Do not reboot/fail over a DB,

promote a replica, change Global Table routing, restore, delete or cut over.

  • Never restore over the source: both services create a new target. Preserve the

source and forensic evidence until the incident/data owner approves disposal.

  • Never connect a restored database to production writers or event consumers

during validation. Use isolated networking, safe secrets and restricted IAM.

  • Confirm account, Region, UTC and resource ARN privately. Redact customer data,

endpoints, account IDs, keys and secrets from submissions.

Download the database recovery workbook or complete archive.

First principle: classify the failure

FailureAvailability feature helps?Historical restore helps?
host/AZ outageMulti-AZ/replica usuallynot normally first response
writer instance failuremanaged failoverif failover data is unusable
Region unavailablecross-Region topology/routingcross-Region copy/restore
bad deployment/schema changelikely replicates damagePITR/snapshot before change
accidental delete/corruptionlikely replicates damagePITR/backup plus reconciliation
compromised credentialsreplica can spread writesisolated immutable/independent point
KMS key unavailablemay make replicas/backups inaccessibleonly with usable independent key/copy

Do not fail over a healthy replica to “repair” logical corruption. Do not restore an old point to solve a brief host failure without calculating data loss.

Shared recovery workflow

detect incident -> fence risky writers -> preserve timeline/audit
  -> determine last known good business state
  -> verify eligible recovery point + KMS/role
  -> restore to unique isolated target
  -> validate control plane, schema, data, security, app, performance
  -> reconcile trusted post-point transactions exactly once
  -> final writer fence and delta
  -> controlled endpoint/routing cutover
  -> monitor and retain rollback target
  -> enable protection on new target
  -> clean temporary resources with exact inventory

RPO is incident time minus the newest validated good point, plus any unrecoverable delta. RTO includes detection, decision, restore queue/runtime, network/IAM/KMS, engine/table availability, data validation, delta replay, cutover and acceptance.

RDS automated backups and snapshots

For DB instances, backup retention can be 0–35 days; zero disables automated backups. Multi-AZ DB clusters require 1–35 days. API/CLI-created DB instances default to one day if omitted, while console creation defaults differ, so always inspect actual configuration. Changing instance retention between zero and nonzero causes an outage.

RDS creates a daily storage snapshot during the backup window and captures transaction logs for PITR; for DB instances logs are uploaded about every five minutes. Restore can target any supported time between EarliestRestorableTime and LatestRestorableTime, not the wall clock “now.” Times must be handled in UTC. Backups require an eligible state such as available; stopped or storage-full databases can create protection gaps.

MechanismRetention/behaviorRecovery use
Automated backup + logsrolling configured windowrestore DB to selected second/new target
Manual snapshotretained until explicitly deletednamed long-lived point/new target
Retained automated backupoptionally retained after source deletionlimited recovery artifact; inspect feature constraints
AWS Backupcentralized plan/vault/copy controlsorganization-wide policy where supported
Snapshot export to S3Parquet analytical export for supported enginesnot a native RDS snapshot restore substitute
Logical dump/binlog/CDCdatabase-level/selective recoverycomplements snapshots; workload-owned consistency

Automated backup storage and manual snapshots are regional. Cross-Region automated-backup replication or snapshot copy requires destination retention, KMS and monitoring. A copied encrypted snapshot is independent only if the destination key and account controls remain usable.

RDS restoration creates a different database

PITR and snapshot restoration produce a new DB instance or cluster with a new identifier/endpoint. They do not overwrite the source. Explicitly review:

  • engine/version compatibility and parameter/option groups;
  • DB instance class, storage type/size/IOPS/throughput/autoscaling;
  • Multi-AZ topology, subnet group, VPC security groups and public accessibility;
  • KMS key, CA certificate, port, authentication, Secrets Manager rotation;
  • maintenance/backup windows, deletion protection and monitoring/log exports;
  • tags, Performance/Database Insights, enhanced monitoring and alarms;
  • RDS Proxy/custom/read/write endpoints and application DNS/configuration;
  • dependent files, queues, caches, search indexes and application release.

Restores can initially use default parameter/option groups unless custom groups are specified/applied. A status of available proves the control plane, not that the schema, extension, collation, users, data or application is correct.

Validate read-only first: database identity/time, migrations/schema version, row counts/checksums, foreign keys, business totals, missing/duplicate orders, latest transaction timestamp, query plans and audit log. Never use customer credentials in a shared test environment.

RDS availability topologies and failover

Multi-AZ DB instance

A synchronous standby in another AZ supports automatic failover; it is not a read endpoint. RDS changes the DB endpoint's DNS record, and existing connections must reconnect. Typical failover is documented as 60–120 seconds but long transactions/recovery and client DNS caching can extend observed RTO. Java DNS TTL should generally be no more than 60 seconds per RDS guidance. Test connection pool retry/backoff, transaction ambiguity and DNS refresh.

Multi-AZ DB cluster

This topology has a writer and readable instances across AZs for supported engines, with cluster endpoints and different performance/failover behavior. Do not apply single-standby assumptions; inventory exact engine/topology.

Read replica

Read replicas are generally asynchronous and serve scale/read/DR patterns. Replication lag is potential data loss. Promotion creates an independent writer and requires application endpoint/routing changes; it is not ordinary automatic Multi-AZ failover. Logical corruption can replicate before promotion.

Aurora

Aurora separates distributed cluster storage from DB instances and uses writer/ reader/custom endpoints. Aurora Replicas can be promoted during instance failure. Aurora Global Database replicates to secondary Regions for DR; managed switchover and unplanned failover have different data-loss and topology consequences. Backtrack, where supported, is not a replacement for independent backups.

For every test capture RDS events, old/new AZ or writer, endpoint DNS answers, connection failures, retry duration, transaction outcome, replica lag and actual RTO. Never trigger a failover in this no-create lesson.

Read-only RDS inventory

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --output json
aws rds describe-db-instances --output json
aws rds describe-db-clusters --output json
aws rds describe-db-snapshots --snapshot-type manual --output json
aws rds describe-db-snapshots --snapshot-type automated --output json
aws rds describe-db-snapshot-attributes --db-snapshot-identifier "$snapshot_id" --output json
aws rds describe-events --duration 1440 --output json
aws rds describe-db-proxies --output json

For instances record BackupRetentionPeriod, backup window, EarliestRestorableTime, LatestRestorableTime, MultiAZ, replica identifiers/ lag metrics, endpoint, subnet/security/parameter/option groups, encryption/key, certificate, deletion protection and pending modifications. For clusters record members/roles, endpoints, global membership and backtrack/PITR fields where valid.

Inspect automated backup replication and AWS Backup separately; one inventory API does not prove every copy. Follow pagination and retain a period exceeding the required RPO rather than checking one recent success.

DynamoDB backup and recovery mechanisms

Point-in-time recovery

PITR continuously protects table data for a configurable 1–35-day window and restores to a selected second between EarliestRestorableDateTime and LatestRestorableDateTime. It does not consume the source table's provisioned throughput. Disabling and re-enabling PITR resets the recoverable start window.

On-demand backup and AWS Backup

On-demand backups are full logical recovery points retained until deleted and run without consuming source throughput. AWS Backup adds scheduled/vault/copy/ audit controls where supported. Record which tool owns retention and deletion.

Export/import and Streams

PITR export to S3 can produce full/incremental analytical data without consuming source RCUs, but importing creates a new table and is a different reconstruction workflow. DynamoDB Streams is a short-lived change stream for CDC/event handling, not a durable backup. It can support carefully audited delta replay only within its retention and consumer evidence.

DynamoDB restore: what returns and what does not

PITR or on-demand restore always creates a new table; source remains available. A full restore can recreate data, key schema, LSIs/GSIs, historical provisioned capacity and encryption settings. You may alter destination billing/capacity, encryption and include/exclude indexes; fewer indexes can reduce restore time.

Manually re-establish and verify after restore:

  • Auto Scaling policies;
  • IAM identity/resource policies;
  • CloudWatch metrics/alarms and Contributor Insights;
  • tags;
  • DynamoDB Streams settings and every event-source mapping/consumer;
  • TTL attribute/enablement;
  • deletion protection;
  • PITR on the new table;
  • application aliases/config/routing and Global Table membership.

Restore duration is not guaranteed solely by table size; indexes and service conditions matter. Cross-Region restore is supported with documented regional exceptions and data-transfer charges - verify current source/destination support.

Validate ACTIVE, exact table ARN/Region/key schema/index status, item counts, sampled/full checksums, business invariants, encryption, denied/allowed access, capacity/throttling and application reads. Do not enable Streams consumers until you prove they cannot replay old business side effects.

Global Tables are availability replication, not historical backup

A Global Table has one replica per Region and propagates writes/deletes. MREC uses multi-Region eventual consistency and last-writer-wins conflict resolution; cross-Region reads can be stale. MRSC has different consistency/Region/topology constraints and must be reviewed separately. Replication lag, quotas, KMS access and application routing determine actual failover behavior.

An accidental delete or corruption is a valid write and can replicate globally. Enable/verify PITR and independent recovery strategy according to risk; current multi-account Global Tables do not replicate PITR settings, so protection must be configured where required. Loss of KMS permission can make a replica inaccessible and, if prolonged, can cause irreversible topology consequences.

Failover is an application-routing and write-authority operation:

  1. detect regional/service failure and check replication status/lag;
  2. fence writes or choose conflict/data-loss policy;
  3. route clients to healthy replica with retries and idempotency;
  4. verify reads/writes and replication/conflict metrics;
  5. reconcile old Region before accepting writes again;
  6. distinguish planned failback from emergency failover.

Never delete/remove a replica or change keys during an incident without the Global Table mode's current documented sequencing and owner approval.

Read-only DynamoDB inventory

aws dynamodb list-tables --output json
aws dynamodb describe-table --table-name "$table_name" --output json
aws dynamodb describe-continuous-backups --table-name "$table_name" --output json
aws dynamodb list-backups --table-name "$table_name" --backup-type ALL --output json
aws dynamodb describe-time-to-live --table-name "$table_name" --output json
aws dynamodb describe-kinesis-streaming-destination --table-name "$table_name" --output json
aws dynamodb list-tags-of-resource --resource-arn "$table_arn" --output json

Inspect table/global status, replica descriptions, consistency mode, capacity, indexes, encryption/key, deletion protection, Streams, TTL, PITR earliest/latest, backups and resource policy. Query CloudWatch consumed/throttled capacity, system/user errors and replication latency for every relevant Region.

Corruption recovery and delta reconciliation

Choose a restore time before the first bad transaction, not merely before alert time. Determine propagation using database audit/binlogs, CloudTrail data events where enabled, DynamoDB Streams/CDC, application journal and immutable business events. Clock skew and late detection widen uncertainty.

Never bulk-copy the damaged current source over the clean restore. Classify each post-point change:

  • trusted and replayable with stable idempotency key;
  • already represented in restored state;
  • derived/rebuildable from authoritative source;
  • corrupted/malicious and excluded;
  • ambiguous and quarantined for owner decision.

Run reconciliation in isolation, compare counts/checksums/invariants, freeze old writers, process final delta, then cut over. Preserve source read-only for an approved rollback/forensic window. A rollback after new writes requires a defined reverse-delta strategy; DNS reversal alone can lose data.

Failure diagnosis

SymptomFirst evidenceLikely boundary
no RDS PITR pointretention and earliest/latest/status/eventsbackups disabled, stopped/unavailable state, lag/window
RDS restore failsevent/status and target configKMS, snapshot state/share, quota, subnet/option/engine incompatibility
DB available, app failsendpoint DNS/TLS/SG/secret/parameter/schemarestored config/client/dependency mismatch
failover slowRDS events, recovery, DNS, pool logstransaction recovery, cache TTL, reconnect/backoff
replica promoted with missing datalag/transaction positionasynchronous replication RPO accepted incorrectly
DynamoDB point unavailablePITR status/window historydisabled/re-enabled or chosen time outside window
table restore active but writes failpolicy/capacity/key/routeomitted IAM/autoscaling/config or KMS deny
duplicate side effects after cutoverStreams mappings/idempotency journalconsumer enabled before safe checkpoint
global replica diverges/unavailablereplica/KMS/latency/conflict metricskey access, quota, routing or conflict model

Fix one proven boundary and use a new restore/test ID. Never delete the only good point or source forensic state while diagnosing.

Cost model

RDS/Aurora costs can include backup/snapshot storage beyond allowances, cross- Region replication/transfer, snapshot export, restored DB instances/clusters, I/O, Performance/Database Insights, proxies, extended support and temporary test environments. DynamoDB costs include PITR/on-demand storage, restore/export/import, cross-Region transfer, restored table storage/capacity/index writes, Global Table replicated writes, Streams and monitoring. KMS, NAT, logs and retained rollback targets add cost. Price the full restore-test duration and cleanup lag.

Practical work and acceptance

Complete the workbook for two incidents: an RDS bad schema/data deployment and a DynamoDB accidental write replicated to another Region. Include exact last-good UTC, point eligibility, RPO/RTO, isolated targets, omitted settings, validation, trusted delta replay, fencing, cutover, rollback and cleanup.

Acceptance requires topology and feature evidence, not product names; correct availability-versus-backup classification; every KMS/network/IAM/config dependency; business-level integrity tests; Global Table conflict/replication reasoning; protection enabled on the new target; exact negative inventory; and steady-state plus incident cost. Reject restore-over-source, “replica is backup,” available/ACTIVE as sole proof, uncontrolled consumer activation, or DNS-only rollback without delta handling.

Knowledge check

  1. Does Multi-AZ repair logical corruption? No; synchronous availability can

preserve/replicate the bad state. Restore an earlier validated point.

  1. Why is LatestRestorableTime not now? Log processing/upload creates a gap;

choose only within observed earliest/latest UTC.

  1. Does RDS PITR keep the same endpoint? No; it creates a new DB target and

endpoint that requires complete configuration and controlled cutover.

  1. What DynamoDB controls are omitted after restore? Autoscaling, IAM policy,

alarms/Contributor Insights, tags, Streams, TTL, deletion protection and PITR, plus application/Global Table configuration.

  1. Why aren't Global Tables backups? Valid bad writes/deletes replicate and

conflict resolution converges current state rather than preserving history.

  1. What makes delta replay safe? Authoritative journal, stable idempotency,

classification, writer fencing, invariant checks and a rollback strategy.

Official sources

Advertisement