Lesson 126 · AWS Learning Path

AWS 126: RDS backups and point-in-time recovery

· Published · 10 min read

Labelled process diagram for AWS 126: Committed data to Automated backup and transaction logs to New restored database to Integrity test and controlled cutover, with decision, proof and rejection evidence.

Why this lesson matters

Plan automated backups, snapshots, retention, restore, and evidence around an explicit RPO and RTO.

A green backup job proves that RDS produced a recovery artifact. It does not prove the correct database, transaction time, KMS key, networking, parameter configuration, credentials, dependencies, or application can be restored inside the required RTO.

What you will be able to do

By the end, you can:

  • explain rds backups and point-in-time recovery in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposePlan automated backups, snapshots, retention, restore, and evidence around an explicit RPO and RTO.
Scope and boundaryAutomated backups support point-in-time recovery inside the retention window. Manual snapshots persist until deleted. A restore creates a new DB instance or cluster.
Evidence of successA restore drill proves the chosen time, new resource settings, integrity checks, application cutover decision, cleanup, and measured recovery time.
Cost modelBackup storage beyond the service allowance, retained manual snapshots, cross-Region copies, restored databases, KMS keys, and transfer can charge.
Safe rejection ruleDo not overwrite the source during a restore or count untested backup existence as a successful recovery.

How the request flows

+----------------------+
|    Committed data    |
+----------------------+
           |
           v
+-----------------------------------------+
|  Automated backup and transaction logs  |
+-----------------------------------------+
                    |
                    v
+-------------------------+
|  New restored database  |
+-------------------------+
            |
            v
+-----------------------------------------+
|  Integrity test and controlled cutover  |
+-----------------------------------------+

For RDS backups and point-in-time recovery, the important boundary is this: Automated backups support point-in-time recovery inside the retention window. Manual snapshots persist until deleted. A restore creates a new DB instance or cluster. A restore drill proves the chosen time, new resource settings, integrity checks, application cutover decision, cleanup, and measured recovery time. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.

Architecture decision table

SituationDirectionReason
Requirement matchesUse automated backups for time-based operational recovery and manual or copied snapshots for defined retention and migration needs.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchDo not overwrite the source during a restore or count untested backup existence as a successful recovery.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

Recovery vocabulary

TermMeaning
RPOMaximum acceptable amount of committed data loss, expressed as time or transactions
RTOMaximum acceptable time from incident to validated service recovery
Automated backupService-managed snapshots plus transaction logs retained for the configured window and used for PITR
Latest restorable timeLatest point currently available for PITR; it can trail current time
Manual snapshotUser-retained point-in-time snapshot that persists until explicitly deleted or retention automation removes it
Final snapshotOptional snapshot requested when deleting a DB; skipping it is a deliberate irreversible choice
Snapshot copyIndependent copy subject to Region/account/KMS and engine constraints; copy completion must be monitored
RestoreCreation of a new DB instance/cluster from snapshot or PITR; not an in-place overwrite

Backup and restore paths

source commits -> automated storage snapshots + transaction logs
                         | retention window
                         +-> PITR to chosen time -> NEW DB endpoint

source/manual action -> DB snapshot -> copy/share/export paths
                         |
                         +-> snapshot restore -> NEW DB endpoint

new DB -> subnet/SG + parameter/option + secret + validation -> controlled cutover

Automated backup retention can be set according to engine/deployment rules; setting retention to zero disables automated backups where supported and can cause an outage when changing to/from zero. Stopped databases can still have retained backup/storage cost and automatically restart after service limits. Manual snapshots are not removed with the source by default, so they need lifecycle ownership.

Choose the recovery mechanism by incident

IncidentLikely mechanismKey limitation
Accidental row/table change at known timePITR to immediately before change, then extract/reconcile or cut overRestores a whole new DB, not one table in place
Need exact retained release baselineManual snapshotRecovery point only; transactions after snapshot are absent
Source Region unavailableCross-Region snapshot/automated-backup replication/copy or replica/Global Database designCopy age and completion determine RPO; restore infrastructure still required
Logical corruption copied to standby/replicaHistorical PITR/snapshotMust select a point before corruption and prevent replay
Analytics exportSnapshot export to S3 where supportedExport is not a directly bootable database backup
Fast test copySnapshot restore, Aurora clone, or blue/green according to engineCost, data masking, dependency isolation and cleanup

Multi-AZ and read replicas are not historical backups. A standby can reproduce bad writes; a read replica can apply a dropped table. AWS Backup can centralize RDS backup policy, copy, vault and restore testing, but underlying engine/resource behavior and validation remain.

Point-in-time recovery timeline

Create a UTC timeline with transaction markers:

10:00 marker A committed
10:05 bad deployment begins
10:07 marker B / corrupt update committed
10:12 incident detected
latest restorable time observed: 10:10
chosen restore time: 10:04:59

The restore target must precede the unwanted transaction while preserving required prior commits. Server clocks, application timestamps, transaction commit time, CloudTrail time, and operator local time are not interchangeable. Record UTC and verify with database rows/business events.

PITR normally accepts latest-restorable or a specific valid time in the retention window. It creates a new endpoint and may require explicit class, storage, subnet group, SG, parameter/option group, public access, port, encryption, monitoring, and log settings. Never assume all current-source modifications are inherited.

Snapshot encryption, copy, and sharing

Encrypted snapshots remain encrypted. Cross-Region/account design requires a usable KMS key and key policy/grants in the destination workflow; snapshots encrypted under an AWS managed key have sharing constraints. Copying can re-encrypt with an eligible destination key. Preserve the key for as long as any recovery point must remain restorable.

Snapshot sharing grants restoration access; it does not copy ownership or expose a running endpoint. Public snapshot sharing is inappropriate for private data. Cross-account recipients should copy an approved shared snapshot into their account when independent retention is required, then validate access and remove sharing according to policy.

Snapshot export to S3 writes data in an analytics format for supported engines; it is not a replacement snapshot that RDS can directly restore. Secure the export role, bucket, KMS key, data catalog, retention, and deletion separately.

Complete restore drill

  1. Select a protected marker and a recovery point/time; record expected rows and RPO.
  2. Confirm recovery artifact status, Region/account, engine/version, KMS key state, and restore permissions.
  3. Build isolated subnet/SG and a recovery secret. Prevent production applications from connecting accidentally.
  4. Restore with explicit instance/storage/configuration. Start the RTO timer before request submission.
  5. Wait with a bound while collecting RDS events; do not poll infinitely or launch duplicates.
  6. Connect with TLS through the recovery path; run engine integrity checks, row counts/checksums, marker queries, migrations, and application smoke tests.
  7. Decide extraction versus cutover. For cutover, freeze/reconcile writes, update secrets/config/DNS, clear pools/caches, and verify positive/negative behavior.
  8. Record achieved RPO/RTO, missing dependencies, cost, and corrective actions.
  9. Delete the restored DB and temporary network/secret/log resources only after evidence approval, or record retained ownership/expiry.

Worked recoveries

Dropped table

PITR to before DROP, validate the table, export/copy only the missing data back through a reviewed reconciliation, and avoid rolling the entire application backward when newer valid transactions must remain.

Ransomware/credential compromise

Restore from a recovery point predating compromise into an isolated account/network, rotate secrets and possibly keys, patch the access path, validate audit evidence, and only then cut over. Restoring into the compromised trust boundary can repeat the incident.

Regional recovery

Use a completed cross-Region copy/recovery point and destination KMS/network/IaC. Measure copy lag as part of RPO and environment creation/application validation as RTO. A snapshot visible in the console is not an application.

Upgrade rollback

Restore the pre-upgrade snapshot to a new DB. Writes accepted after upgrade require forward migration/reconciliation; switching endpoints back without handling them causes data loss.

Failure diagnosis

SymptomEvidenceLikely boundary
Chosen PITR time rejectedretention window, earliest/latest restorable, UTC formatInvalid recovery point/time
Snapshot restore deniedcaller, snapshot ownership/share, KMS key policy/stateIAM/resource/KMS authorization
Restore remains creatingRDS events, quota, subnet IPs, class/engine availabilityCapacity/dependency
Restored DB unreachableendpoint, SG/NACL/route/DNS, port, TLSNetwork/client path
App schema differssnapshot time, parameter/option, migration historyLogical/configuration recovery
Bill remains after source deletionsnapshots, retained automated backups, restored test DB, exports/logsRetained resource

Acceptance evidence

Pass only with a real or supplied restore that proves marker inclusion/exclusion, exact recovery artifact and KMS path, database integrity, application smoke test, measured RPO/RTO, failed-path diagnosis, cutover/rollback reasoning, and final cost/cleanup inventory.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open RDS Databases, choose a supplied database, and inspect Automated backups and Latest restorable time.
  2. Open Snapshots and Automated backups, then distinguish manual, automated, and retained records.
  3. Start the restore wizard only far enough to identify required new-instance values, then cancel without creating it.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws rds describe-db-instances --query 'DBInstances[].{Id:DBInstanceIdentifier,Retention:BackupRetentionPeriod,Earliest:EarliestRestorableTime,Latest:LatestRestorableTime,Encrypted:StorageEncrypted}' --output table
aws rds describe-db-snapshots --snapshot-type manual --query 'DBSnapshots[].{Id:DBSnapshotIdentifier,Source:DBInstanceIdentifier,State:Status,Created:SnapshotCreateTime}' --output table

Expected interpretation

Restorable timestamps and snapshot availability show recovery material exists. They do not prove schema consistency, usable credentials, application integrity, or measured RTO.

Practical work

Write a restore runbook with a sample corruption time, acceptable restore point, new identifier, subnet and SG, parameter groups, validation queries, cutover, rollback, evidence, and deletion order.

Diagnose this topic from its own evidence

Use the recovery timeline and failure table. A missing row can mean the wrong restore time, transaction never committed, wrong database/schema, migration mismatch, replica lag before backup, or application cache - not necessarily backup corruption. Compare exact marker transactions and engine integrity evidence.

Cost and cleanup

Backup storage beyond the service allowance, retained manual snapshots, cross-Region copies, restored databases, KMS keys, and transfer can charge.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Plan automated backups, snapshots, retention, restore, and evidence around an explicit RPO and RTO.

  1. Which scope or ownership boundary must be proved first?

Expected direction: Automated backups support point-in-time recovery inside the retention window. Manual snapshots persist until deleted. A restore creates a new DB instance or cluster.

  1. What evidence is strong enough to accept the result?

Expected direction: A restore drill proves the chosen time, new resource settings, integrity checks, application cutover decision, cleanup, and measured recovery time.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Do not overwrite the source during a restore or count untested backup existence as a successful recovery.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Backup storage beyond the service allowance, retained manual snapshots, cross-Region copies, restored databases, KMS keys, and transfer can charge.

Lesson acceptance

Pass only with an exact RPO/RTO, automated/manual/copy decision, successful isolated restore, marker and integrity validation, one failed dependency diagnosis, cutover and write-reconciliation plan, KMS/secret/network proof, and deletion or retained-snapshot ownership.

Official sources

Advertisement