AWS 186: Backup and restore strategy
Why this lesson matters
Build recovery from business impact, RPO, RTO, consistency, retention, immutability, copy, access, restore order, and repeated testing.
What you will be able to do
By the end, you can:
- explain backup and restore strategy in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Build recovery from business impact, RPO, RTO, consistency, retention, immutability, copy, access, restore order, and repeated testing. |
| Scope and boundary | The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Backup and restore strategy. |
| Evidence of success | Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Backup and restore strategy. |
| Cost model | Backup storage, warm and cold retention, copies, restores, cross-Region transfer, restore-test resources, and logs can charge. |
| Safe rejection rule | Avoid counting replication as backup, deleting the source before restore validation, or locking a vault in a training account. |
How the request flows
+----------------------+
| Protected workload |
+----------------------+
|
v
+------------------------------------+
| Backup policy and recovery point |
+------------------------------------+
|
v
+----------------------+
| Isolated restore |
+----------------------+
|
v
+---------------------------------------------+
| Integrity test, measured RTO, and cleanup |
+---------------------------------------------+
Define recovery before choosing services
RPO is the maximum acceptable data-loss window measured backward from disruption. RTO is the maximum acceptable time to restore the business capability. Both come from business impact and must include detection, decision, infrastructure deployment, data restore, validation, DNS/traffic change and dependent systems - not only an AWS job duration.
High availability handles expected component/AZ faults continuously. Disaster recovery handles a declared disaster that prevents business objectives in the primary location. Backup protects a point in time; replication reduces lag but can copy corruption or deletion. A resilient design usually needs both.
Backup-and-restore architecture
- Inventory state: databases, objects, volumes, file systems, application artifacts, IaC, configuration, secrets/key dependencies and external integrations.
- Select backup frequency/continuous recovery to meet RPO, retention for legal/business needs, and copy destination for account/Region isolation.
- Protect recovery points with least privilege, separate backup administration, Vault Lock or logically air-gapped vault where justified, and monitored deletion/change events.
- Pre-stage the recovery account/Region baseline: identity, quotas, networking, DNS, KMS keys/policies, artifact repositories and deployment pipeline.
- During recovery, deploy infrastructure from versioned IaC, restore data to new resources, deploy code/configuration, validate integrity/security/business transactions, then shift traffic.
- Record actual recovery point, elapsed RTO, exceptions, cleanup and failback plan.
Backups encrypted under a key that is deleted or inaccessible in the recovery account are unusable. Cross-account/cross-Region support differs by resource type, and copies may use a destination key. Check AWS Backup's current feature matrix rather than assuming every protected resource supports every copy, PITR, cold tier or restore-testing capability.
Recovery-point and corruption decisions
Choose the newest known-good point, not automatically the newest point. Preserve multiple versions and enough retention to reach before delayed ransomware, operator deletion or application corruption. Define how database consistency, transaction logs, object versions and dependent data stores align. Restoring each tier from unrelated timestamps can create a technically successful but logically corrupt application.
AWS Backup restore testing can periodically select eligible recovery points, launch restore jobs and measure duration. A completed restore job still needs validation such as filesystem mount, database consistency query, application smoke transaction, malware/security check and teardown. Tag and isolate test restores so they cannot send production messages or access production secrets.
Control plane and runbook risk
Recovery that creates resources depends on regional control planes, quotas and capacity. Prefer pre-created data-plane mechanisms for the most stringent targets, and pre-approve alternate instance types/Regions. Store runbooks and credentials outside the failed environment. Every manual approval must have an owner, deadline and alternate.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use centralized backup policy where supported, but preserve application-specific consistency and restore procedures. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid counting replication as backup, deleting the source before restore validation, or locking a vault in a training account. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Use the Console service search and open AWS Backup, Backup plans, Backup vaults, Protected resources, and Jobs; confirm the account and Region before reading the page.
- Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
- Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
- Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws backup list-backup-plans --query 'BackupPlansList[].{Name:BackupPlanName,Id:BackupPlanId,Version:VersionId}' --output table
aws backup list-backup-vaults --query 'BackupVaultList[].{Name:BackupVaultName,Points:NumberOfRecoveryPoints,Locked:Locked}' --output table
aws backup list-backup-jobs --by-state FAILED --max-results 20 --output table
Expected interpretation
A successful backup job proves a recovery point was written under that job. Only a validated restore, integrity test, measured time, and application cutover prove recoverability.
Practical work
Create a backup matrix for S3, EBS, RDS, DynamoDB, EFS, and application configuration. Set RPO/RTO, consistency method, schedule, lifecycle, vault, cross-account/Region copy, restore order, quarterly test, and evidence owner.
Diagnose this topic from its own evidence
- Backup job
COMPLETEDbut no usable restore: inspect recovery-point ARN, resource support, KMS/key policy, restore metadata, IAM role and application validation. - RPO missed: compare last successful recovery point/PITR window with disaster time and detection delay.
- RTO missed: split elapsed time into declaration, access, quota, IaC, restore, deployment, validation and traffic stages.
- Cross-account copy absent: inspect organization relationship, vault/access policy, key policy, resource feature support and destination Region.
- Restore works but application fails: check dependency order, secrets, DNS, certificates, identity, data consistency and outbound integrations.
Cost and cleanup
Backup storage, warm and cold retention, copies, restores, cross-Region transfer, restore-test resources, and logs can charge.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Build recovery from business impact, RPO, RTO, consistency, retention, immutability, copy, access, restore order, and repeated testing.
- Which scope or ownership boundary must be proved first?
Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Backup and restore strategy.
- What evidence is strong enough to accept the result?
Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Backup and restore strategy.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid counting replication as backup, deleting the source before restore validation, or locking a vault in a training account.
- Which cost dimensions and retained resources need an owner?
Expected direction: Backup storage, warm and cold retention, copies, restores, cross-Region transfer, restore-test resources, and logs can charge.
Lesson acceptance
- Calculate RPO loss and RTO elapsed time from a supplied timeline.
- Produce a state/dependency inventory and per-resource backup/copy/retention matrix.
- Explain backup versus replication and select a known-good recovery point after corruption.
- Design isolated automated restore testing with application-level validation and cleanup.
- Prove recovery-account identity, network, key, quota, artifact and runbook readiness.