Lesson 232 · AWS Learning Path

AWS 232: AWS Backup plans, vaults, lifecycle, cross-Region copies, and restore testing

· Published · 13 min read

Labelled process diagram for AWS 232: Protected resource and policy to Backup and copy jobs to Isolated restore test to Integrity, RPO/RTO, and audit evidence, with decision, proof and rejection evidence.

Why this lesson matters

A successful backup job is not the same as recoverability. The recovery point can be too old, crash-inconsistent, encrypted with an unavailable key, copied to an inaccessible vault, missing dependent resources, too slow to restore, or never validated by the application owner.

AWS Backup centralizes policy, jobs and evidence for supported services. It does not discover business RPO/RTO, quiesce every application, validate business data, or decide legal retention. Those remain architecture and ownership decisions.

Outcomes

You will be able to:

  • translate business impact into measurable RPO, RTO, retention and restore scope;
  • distinguish plans, rules, selections, assignments, vaults, recovery points,

backup/copy/restore jobs, indexes and reports;

  • calculate schedule, start/completion windows and lifecycle constraints;
  • distinguish snapshot backups from continuous point-in-time recovery;
  • design tag/ARN selections without silently missing resources;
  • explain vault access policy, KMS ownership, Vault Lock modes and logically

air-gapped vaults;

  • trace cross-Region and cross-account copy authorization;
  • design isolated restore testing and application validation;
  • diagnose failed, expired, partial, copy, restore, validation and cleanup states;
  • estimate storage, transfer, retrieval, restore-test and retained-resource cost.

Before you start

  • This lesson is no-create. Every provided AWS command is read-only. Complete the

workbook locally; do not create plans, vaults, locks, copies or restores.

  • Never enable Vault Lock - especially compliance mode - from a training exercise.
  • Use a non-root audit role and confirm account/Region. Do not expose recovery-

point ARNs, account IDs, key policies or restored data in shared submissions.

  • Feature support differs by resource type, Region, backup type and account

setting. Check the current AWS feature matrix before approving a design.

  • Never restore sensitive data into a shared/default VPC merely to test it.

Download the design artifact

The workbook forces every schedule, selection, identity, vault, key, copy, restore, validation, exception, cost and cleanup decision to have an owner.

Start with the recovery contract

business service and data dependencies
  -> maximum tolerable data loss (RPO)
  -> maximum tolerable outage (RTO)
  -> backup frequency/type/consistency
  -> retention and independent copy
  -> isolated restore and validation
  -> measured point age + end-to-end recovery duration

RPO is the age of the newest usable and sufficiently consistent recovery point at the incident time. An hourly schedule does not guarantee a one-hour RPO if jobs start late, expire, fail, overlap, or produce unusable data.

RTO is not just RestoreJob.CompletedAt - CreatedAt. Include incident decision, approval, recovery-point discovery, restore queue/runtime, KMS/network/IAM and dependency provisioning, database/application startup, validation, DNS/traffic cutover and business acceptance.

Define recovery granularity: one object/table/volume, whole resource, application stack, account, or Region. Restoring a database without its secrets, parameter versions, file data, queues, DNS and compatible application release is not service recovery.

Core resource model

ObjectPurposeCritical evidence
Backup planversioned policy containerplan ID/version/advanced settings
Backup ruleschedule, windows, vault, lifecycle, copy, continuous flageffective values and overlap
Selectionresources/tags and backup roleincluded/excluded current inventory
Vaultregional recovery-point containertype, key, policy, lock, point count
Recovery pointrestorable backup representationstatus, creation/completion, type, key, lifecycle
Backup jobproduces recovery pointstate, resource, role, bytes, times/error
Copy jobcreates independent destination pointsource/destination, key, state/error
Restore jobmaterializes resource from pointmetadata, role, state/times/error
Restore-testing planschedule and recovery-point algorithm/windowtimezone/start window/vaults
Restore-testing selectionresource type and ARN or condition scoperestore role/metadata/cleanup hours
Audit framework/reportevaluates configured controlscontrol parameters/scope/evidence period

Plans are versioned. Inventory the plan version attached to observed jobs, not only today's edited plan. Recovery points remain governed by lifecycle assigned when created; changing a plan does not prove older points changed.

End-to-end identity and data path

plan rule + selection
  -> AWS Backup service assumes backup role
  -> source service APIs + source KMS key policy/grants
  -> source vault/recovery point
  -> copy role: CopyFrom source + CopyInto destination
  -> destination vault access policy + destination KMS key
  -> restore-testing role + restore metadata
  -> isolated restored resource
  -> EventBridge validation workflow
  -> validation result + deletion after validation window

The human/automation caller needs control-plane permissions. The backup role needs source-service operations. Vault policies govern supported vault access; KMS key policy/grants must authorize cryptographic use. Restore uses a restore role and service-specific metadata. An allow in one layer cannot override an explicit deny, SCP, boundary, key policy or incompatible vault retention.

Backup rules: schedule and windows

A rule contains a schedule (cron/rate where supported), optional timezone, start window, completion window, target vault, lifecycle, copy actions, recovery-point tags and optional continuous backup.

  • The start window is how long AWS Backup may begin the job. For retryable start

errors, it retries at least every ten minutes until running or EXPIRED when the start window closes.

  • The completion window begins when the job starts. Expiry/failure inside either

window affects achieved RPO even if the schedule itself is correct.

  • Overlapping rules can intentionally create different retentions/copies, but

can also create duplicate points and cost. Record precedence/effective result.

  • Monitor CREATED, PENDING, RUNNING, COMPLETED, FAILED, ABORTED,

EXPIRED and partial states where exposed; don't filter only FAILED.

Snapshot versus continuous recovery

Snapshot backups provide discrete points. Continuous backup maintains PITR for supported resources and can offer a recovery window, but it is not universally supported and does not mean zero data loss. Current documented retention is:

  • snapshot recovery points: 1 day to 100 years, or indefinite when configured;
  • continuous recovery points: 1 to 35 days.

Creation date for lifecycle is based on backup job start, not completion. Verify service-native behavior: incremental storage, transaction consistency, PITR granularity, retention and restore semantics differ across EBS, EC2, EFS, RDS/Aurora, DynamoDB, S3 and other supported resources.

Lifecycle and cold storage

MoveToColdStorageAfterDays and DeleteAfterDays are not arbitrary. A backup moved to cold must remain there at least 90 days, in addition to warm time. AWS recommends waiting at least eight days before transition. Only supported resource types can transition, and cold retrieval delay/cost can break RTO.

backup starts day 0
 -> warm period
 -> optional cold transition day N
 -> at least 90 days in cold
 -> expiration/deletion day >= N + 90

Some lifecycle values cannot be changed after transition. Vault Lock min/max retention can reject jobs whose rule lifecycle falls outside its bounds. Legal hold, retention, deletion and cost must be modelled together.

Resource selection and opt-in

Selections can use explicit ARNs, resource patterns and tags/conditions, with an IAM role. Organization backup policies can select resource types/tags broadly, but not individual resource ARNs; local plans can target individuals.

Tag governance questions:

  • Who can set/remove Backup, environment and data-classification tags?
  • Is matching AND or legacy OR behavior for the exact selection structure?
  • What happens when a resource is untagged, mistagged, newly created, moved,

unsupported, opted out, or in another Region?

  • Does the backup role have service/KMS/tag permissions for every selected type?
  • Is there a daily control comparing application inventory with protected resources?

Never infer coverage from a plan name. Prove current selection, Region settings, effective Organization policy and list-protected-resources/job evidence.

Vaults, encryption and access policy

Standard vaults are regional containers. Encryption behavior depends on whether AWS Backup fully manages that resource type's backup: some recovery points use the vault key independently, while others retain encryption relationships with the source service/key. Copies can change encryption boundaries only according to service support.

For each vault record:

  • ARN/account/Region and standard versus logically air-gapped type;
  • KMS key ARN, key policy, grants, rotation/deletion state and recovery owner;
  • vault access policy, IAM/SCP/boundary and denied public/cross-account access;
  • lock mode/date/minimum/maximum retention;
  • recovery-point inventory, lifecycle, legal hold and cost owner;
  • EventBridge/CloudWatch monitoring and break-glass process.

The default vault/key is not automatically the right isolation boundary. Disabling/deleting a KMS key can make retained recovery points unusable even while backup inventory still looks healthy.

Vault Lock and logically air-gapped vaults

Vault Lock enforces write-once/read-many retention:

  • Governance mode: sufficiently authorized identities can remove/change lock.
  • Compliance mode: specifying ChangeableForDays creates a grace period

(currently 3–36,500 days). After it expires, no customer, root user or AWS can remove/change the lock; protected points cannot be deleted before lifecycle.

Minimum/maximum lock retention constrains newly written points; existing points from before lock creation are not retroactively changed. A point retained “always” in a compliance-locked vault can create permanent, unavoidable storage.

Logically air-gapped vaults add isolation/sharing capabilities, encryption with an AWS-owned or customer-managed key, and are always compliance-mode locked. They are not a classroom substitute for threat modelling, multi-party approval, separate-account controls, restore permissions and ransomware recovery tests.

Cross-Region and cross-account copies

Cross-Region protects against a regional failure but not source-account compromise by itself. Cross-account copy can provide administrative isolation.

Current cross-account fundamentals:

  • source and destination accounts must belong to the same AWS Organization;
  • the Organizations management account enables the global cross-account setting;
  • source identity/role needs scoped CopyFromBackupVault and CopyIntoBackupVault;
  • destination non-default vault policy must allow CopyIntoBackupVault;
  • source and destination KMS policies/grants and service-specific sharing must work;
  • resource and Region feature support must be checked;
  • cross-account copy into cold storage is not supported;
  • SCPs should limit approved destination accounts/vault tags and prevent an

isolated backup account from casually leaving the Organization.

If a destination account leaves, it can retain copied backups - an isolation benefit and data-governance risk. Disabling cross-account backup can still allow in-flight behavior for a short eventual-consistency period. Never delete a copy role while jobs are running; service cleanup/unsharing can fail.

Copy success is separate from source backup success. Monitor both job IDs and prove destination recovery-point status/key/lifecycle. Then restore from the destination; an inventory-only copy test is insufficient.

Restore testing: the proof layer

AWS Backup restore testing schedules real restore jobs from real eligible points. A plan defines schedule/timezone/start window and recovery-point selection; each resource-type selection defines protected resources/conditions, restore role, metadata overrides and cleanup/validation duration.

Important semantics:

  • choose LATEST_WITHIN_WINDOW or RANDOM_WITHIN_WINDOW from included vaults;
  • selection window is evaluated from actual job execution, which can occur in

the start window; make it larger than backup interval plus start-window time;

  • a selection uses protected-resource ARNs or wildcard with conditions, not both;
  • conditions identify protected resources using their latest point's tags; they

do not force the ultimately chosen restore point to carry those tags;

  • at most one eligible recovery point is restored per selected protected resource;
  • continuous points are eligible only when explicitly included and supported;
  • no eligible point is evidence of a coverage/window defect, not a passed test;
  • restored resources supporting tag-on-restore receive awsbackup-restore-test;
  • optional validation retention is 1–168 hours; after validation or window close,

AWS Backup deletes test resources according to service SLAs.

Use a dedicated test account/VPC with denied production routes, safe DNS, non-production secrets, quotas and an exact cleanup inventory. Inferred restore metadata is a starting point; override subnet, security group, IAM profile, instance class, names and other service-specific fields to keep isolation/cost.

Validation is more than COMPLETED

An EventBridge rule can invoke a validator after restore job completion. Validate:

  1. correct recovery-point ID/time and restore target identity;
  2. resource available and encrypted under approved key;
  3. mount/connect/read using least-privilege test identity;
  4. checksum, row counts, referential/business invariants and application version;
  5. dependent-resource and smoke-transaction behavior;
  6. denied production connectivity and no message/email/external side effects;
  7. measured RPO and end-to-end RTO;
  8. validation result captured before cleanup; exact negative inventory afterward.

Submitting a restore validation status without executing business tests creates false assurance. A failed validation can be more important than a completed job.

Read-only operational inventory

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --output json
aws backup get-region-settings --output json
aws backup get-global-settings --output json
aws backup list-backup-plans --include-deleted --output json
aws backup list-backup-vaults --output json
aws backup list-protected-resources --output json
aws backup list-backup-jobs --max-results 100 --output json
aws backup list-copy-jobs --max-results 100 --output json
aws backup list-restore-jobs --max-results 100 --output json
aws backup list-restore-testing-plans --output json
aws backup list-frameworks --output json
aws backup list-report-plans --output json

For every plan, retrieve effective rules and selections:

aws backup get-backup-plan --backup-plan-id "$plan_id" \
  --version-id "$version_id" --output json
aws backup list-backup-selections --backup-plan-id "$plan_id" --output json
aws backup get-backup-selection --backup-plan-id "$plan_id" \
  --selection-id "$selection_id" --output json

For every vault, inspect - not modify - policy, lock and points:

aws backup describe-backup-vault --backup-vault-name "$vault_name" --output json
aws backup get-backup-vault-access-policy --backup-vault-name "$vault_name" --output json
aws backup list-recovery-points-by-backup-vault \
  --backup-vault-name "$vault_name" --output json
aws kms describe-key --key-id "$vault_key_arn" --output json
aws kms get-key-policy --key-id "$vault_key_arn" --policy-name default --output text

Absence of a vault policy response, an empty current page, or one recent COMPLETED job does not establish coverage. Follow pagination and observation windows spanning required RPO/retention periods.

For restore testing:

aws backup get-restore-testing-plan --restore-testing-plan-name "$test_plan" --output json
aws backup list-restore-testing-selections \
  --restore-testing-plan-name "$test_plan" --output json
aws backup get-restore-testing-selection \
  --restore-testing-plan-name "$test_plan" \
  --restore-testing-selection-name "$selection_name" --output json

Then correlate restore-testing job, restore job, selected point, validator result, resource tags and exact cleanup through job IDs and UTC timestamps.

Failure diagnosis

SymptomFirst evidenceCommon boundary
EXPIREDschedule/start window/status messageno successful start before window, throttling or retryable dependency
backup failedjob status/code/message and CloudTrailbackup role, service/KMS policy, opt-in, quota, state
expected resource absenteffective selection/tags/Region settingstag drift, wrong ARN/type/Region, opt-out, role
copy failedsource/destination job, vault/key policiesOrganization setting, CopyFrom/Into, KMS, unsupported cold/type
point exists but restore deniedrestore metadata/role/key/service policyrestore identity, KMS, network/name/quota
no eligible restore-test pointplan vaults/window and protected selectioninterval/start-window edge, ARN/condition, continuous flag
restore completed, validation failedvalidator logs and data invariantinconsistent/incomplete data or dependencies/application
test resource remainsvalidation window/job/deletion event/inventorycleanup delay/failure, unsupported dependency or manually created child
lock rejects joblock min/max versus rule lifecycleincompatible retention design

Always separate control-plane inventory, actual data-plane recovery, application validation and cleanup. Repair one boundary and rerun with new job IDs; never delete the only recovery point while diagnosing.

Audit Manager and evidence limitations

AWS Backup Audit Manager frameworks evaluate configured controls such as backup frequency, retention, cross-account/cross-Region copy and restore objectives. Scope, control parameters, resource support, evidence time and report delivery must be inspected. A compliant control does not independently prove application consistency, complete dependency restoration or successful business validation.

Use EventBridge for backup/copy/restore job state changes and route to monitored targets with retry/DLQ. Alarm on failed/aborted/expired jobs, no recent recovery point, copy lag, restore/validation failure and cleanup orphaning. Test alarms; missing metrics must not silently mean healthy.

Cost model

Estimate by resource type and Region:

  • warm backup storage and incremental/full semantics;
  • cold storage, minimum duration and early deletion where applicable;
  • cross-Region/account data transfer and destination storage;
  • restore and cold retrieval charges;
  • temporary restore resources: EC2/RDS/EFS/FSx/S3 requests, NAT and data transfer;
  • KMS requests/keys, EventBridge, Lambda, logs, alarms and Audit reports;
  • indefinite/locked retention and failed-cleanup resources.

Deduplication/incremental behavior is service-specific; do not multiply source size by retention blindly or promise savings without measured change rate.

Practical work and acceptance

Complete every section of the workbook for EC2/EBS, EFS, RDS/Aurora and DynamoDB. For each, verify current feature support and distinguish native service backup from AWS Backup ownership. Include daily and continuous needs only where valid.

Acceptance requires:

  • explicit business RPO/RTO and measured calculation method;
  • resource/dependency/consistency inventory and authoritative backup tool;
  • rule schedule/windows/lifecycle/copy math with cold constraints;
  • positive, negative, missing and conflicting selection-tag tests;
  • source/destination account, vault, KMS, Organization and IAM path;
  • Vault Lock/air-gap threat model without enabling a lock;
  • restore plan/selection/metadata, latest/random rationale and isolation;
  • workload validation, cleanup and negative inventory;
  • failure alarms, Audit Manager limitations, exception expiry and monthly cost.

Reject “backup completed” as sole proof, assumed tag coverage, inaccessible KMS keys, untested destination copies, compliance Vault Lock in training, production- connected restore tests, omitted dependencies or unowned retained cost.

Knowledge check

  1. Hourly schedule equals one-hour RPO? No; point usability, consistency,

start delay/failure and validation determine achieved RPO.

  1. Cold lifecycle minimum? A transitioned backup must remain cold at least

90 days, and service support/retrieval RTO must be checked.

  1. Governance versus compliance lock? Authorized users can remove governance;

compliance becomes immutable after grace expiry.

  1. Does restore COMPLETED prove recovery? No; data/application/security

validation and dependency checks remain.

  1. What authorizes cross-account copy? Same Organization/global setting,

source identity permissions, destination vault policy, KMS/service permissions and absence of overriding denies.

  1. Why enlarge restore selection window? Eligibility is evaluated from actual

job execution within the start window, not only scheduled start.

Official sources

Advertisement