AWS 119: AWS Backup and cross-account backup
The real problem
A team recognizes the name AWS Backup and cross-account backup but has not connected the feature to a real requirement, identity boundary, network or data path, failure mode, price dimension, and cleanup owner. A plausible configuration could still fail the workload.
Final outcome
The learner will produce a requirement-led artifact for AWS Backup and cross-account backup, inspect the matching AWS control plane in the Management Console, run a matching CloudShell or AWS CLI query, interpret the output, diagnose one failure, defend one architecture choice, and prove cleanup or approved retained state.
The practical outcome is not a command transcript. It must show what was expected, what happened, what the result proves, what it does not prove, and which evidence would change the decision.
Learning objectives
By the end of this lesson, the learner can:
- explain central policy;
- explain vault boundary;
- explain cross-account copy;
- explain cross-region copy;
- explain vault lock;
- connect control-plane state to the real data, network, identity, or application behavior;
- identify cost and cleanup ownership before any optional mutation;
- troubleshoot from evidence without opening broad access or adding broad permissions.
Relationship model
Requirement
|
v
Identity and policy -> AWS configuration -> network or data path -> workload behavior
| | | |
+--------------------+----------------------+--------------------+
|
v
monitoring, cost, recovery, cleanup
Use this model to separate an AWS object that exists from a result that actually works. Every arrow is a verification boundary.
Prerequisites, permissions, Region, and safety
- Learning baseline: This sequence assumes practical Linux knowledge but no prior cloud-computing or AWS knowledge. Cloud, networking, security, data, automation, and architecture concepts must come from completed earlier lessons. If a prerequisite checkpoint is incomplete, return to its linked lesson before continuing.
- Confirm a non-root caller with
aws sts get-caller-identityand keep the account number private. - Use
ap-south-1unless this lesson explicitly names a second Region. - Confirm the intended profile and Region with
aws configure listbefore interpreting an empty result. - Use read-only List, Get, and Describe permissions for the named services. Design exercises run locally and require no resource-creation permission.
- This is a no-create lesson. Console and CLI work is read-only, and every design artifact is created locally.
- Never publish account IDs, public addresses, ARNs containing private account data, session IDs, presigned URLs, object data, credentials, or KMS material.
- Do not use root, world-open SSH or RDP, disabled TLS verification, unowned resources, or irreversible retention controls in a training exercise.
Core model
| Concept | What the learner must understand |
|---|---|
| Central policy | AWS Backup uses plans, rules, schedules, windows, lifecycle, vaults, selections, IAM roles, copy actions, and recovery points to centralize supported-service protection. Coverage varies by service and Region. |
| Vault boundary | A backup vault organizes recovery points and has encryption and access-policy behavior. Source-service permissions can still matter for some backup types. |
| Cross-account copy | AWS Backup can copy supported backups to another account in an AWS Organization when vault policy, organization settings, IAM, encryption keys, and service support align. |
| Cross-Region copy | Copy actions can create recovery points in another Region. RPO, copy completion, encryption, data transfer, retention, and restore location need monitoring. |
| Vault Lock | Governance mode allows authorized management. Compliance mode becomes immutable after its grace time and can force retained storage cost until recovery points expire. |
| Logically air-gapped vault | This specialized vault includes compliance-mode Vault Lock, service-account isolation, sharing and recovery features, and additional availability and resource-type considerations. |
How it works
A backup is useful only when the correct resource, point in time, encryption keys, metadata, role, network, dependencies, and application validation can be restored within RTO. Schedule restore tests and record evidence, not only backup-job success.
Begin with business impact. RPO is the maximum acceptable data-loss window; it drives frequency or continuous-backup needs. RTO is the maximum acceptable time to restore usable service; it includes authorization, infrastructure reconstruction, data restore, application validation, DNS/routing and decision time. Retention states how long recovery points must remain. Availability replicas and data replication can reduce outage time but may copy corruption or deletion; backups provide historical recovery points. Neither replaces the other.
Control and data model
An AWS Backup plan contains rules for schedule, start/completion windows, target vault, lifecycle and copy actions. A selection maps resources explicitly or by tags and uses a service role. A job attempts the backup/copy/restore. A recovery point is the retained result in a vault. Frameworks and Audit Manager controls assess evidence; reports and EventBridge/CloudWatch alarms expose failed or missing jobs. Coverage and exact feature behavior vary by resource type and Region, so the architecture must maintain a dated support matrix.
Encryption behavior is not uniform. Some resources are fully managed by AWS Backup while others retain source-service snapshot behavior; source and destination KMS keys, key policies, grants and account ownership affect copy and restore. Deleting or disabling a KMS key can make retained encrypted recovery points unusable. Cross-account copy normally requires supported resource type, Organizations trust/settings, destination vault access policy, source/copy roles and compatible encryption. Cross-Region and cross-account are separate dimensions and produce separate recovery points with transfer/storage and completion lag.
Immutability and isolation
Standard vaults organize and policy-protect recovery points. Vault Lock governance mode can be changed by sufficiently privileged identities; compliance mode becomes immutable after its cooling-off/grace period. Once locked, neither the customer nor AWS can shorten retention or delete protected recovery points until expiry. This is a legal, operational and cost commitment - never enable it as an exploratory lab action.
A logically air-gapped vault is a specialized vault with compliance-mode lock and backup data isolated in an AWS Backup service-owned account. Current capabilities include RAM-based sharing/restore patterns and Multi-party approval options, subject to Region/resource restrictions. It does not mean a physical offline tape and does not excuse testing identities, KMS behavior, sharing and restore access. Separate backup/security and recovery accounts reduce the blast radius of compromised production administrators only if SCPs, break-glass roles, keys and approval paths are independently controlled.
Restore is the product
Restore to a separate, quarantined destination. Reconstruct dependencies - VPC/subnets/SGs, IAM roles, KMS keys, secrets, parameter groups, DNS and application configuration - then validate schema, row/object/file counts, checksums and business transactions. Scan for malware where required before reconnecting. Measure requested-to-usable time and clean up the restored resources. Automated restore testing proves that AWS can perform configured restore steps for supported resources; application-level validation and dependency exercises remain your responsibility.
Read the result in layers:
- Scope: account, Region, VPC, bucket, AZ, endpoint, principal, object version, or resource ARN.
- Control plane: the requested configuration exists and reached an expected state.
- Behavior: the request, connection, health check, replication, restore, or application result meets the requirement.
- Operations: monitoring, failure owner, cost, retention, rollback, and cleanup are known.
Control-plane success is necessary but not sufficient. A resource can be available while policy, routing, DNS, health, data, or application behavior remains wrong.
Architecture decision table
| Requirement | Preferred direction | Why |
|---|---|---|
| Central schedules across supported services | AWS Backup plan | One policy coordinates selections, vaults, lifecycle, and copies. |
| Reduce production-account deletion blast radius | Cross-account copy or logically air-gapped vault by requirement | Administrative isolation adds recovery protection. |
| Non-bypassable regulatory retention | Compliance Vault Lock after formal review | The setting becomes immutable and cost persists until retention completes. |
| Need evidence that backup works | Automated or scheduled restore test | A completed backup job alone does not prove recovery. |
Professional questions normally contain several valid services. State the requirement that selects one option, why the nearest alternative fails it, and what changed requirement would reverse the choice.
AWS Management Console guided practice
Before opening a service page, write the expected account, Region, starting state, and evidence. Do not choose Create, Save, Purchase, Lock, or Delete unless the lesson explicitly authorizes the live track.
- Open AWS Backup Dashboard, Backup plans, Backup vaults, Protected resources, Jobs, and Restore testing; inspect supplied evidence without creating a plan.
- Open a vault and record type, encryption key, access policy, lock mode and date, retention boundaries, recovery points, copy state, and sharing scope.
- Trace one recovery point through restore role, metadata, destination network, validation, cleanup, measured RTO, and audit record.
For each step, capture the field name and value in text. A screenshot may support the record but does not replace the explanation. Console labels can evolve, so use the service search and current documentation if a navigation label differs.
CloudShell and AWS CLI practice
CloudShell is the default browser-based command environment taught in AWS 028. AWS 029 and AWS 030 cover local CLI installation and authentication. This lesson therefore does not assume that an unconfigured local shell is ready.
Start every session with:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account portion of the ARN before sharing. Then perform the topic query:
Inventory plans, vaults, lock state, recovery-point count, and recent backup and restore jobs without creating one.
aws backup list-backup-plans --query 'BackupPlansList[].{Id:BackupPlanId,Name:BackupPlanName,Version:VersionId}' --output table
aws backup list-backup-vaults --query 'BackupVaultList[].{Name:BackupVaultName,Type:VaultType,Locked:Locked,Points:NumberOfRecoveryPoints,Key:EncryptionKeyArn}' --output table
aws backup list-restore-jobs --by-status COMPLETED --output table
Expected interpretation:
Recovery points and completed jobs support backup evidence. They do not prove the right data, application consistency, cross-account access, dependency reconstruction, or achieved RTO.
Replace every replace-with-... sample value before running its command, and use only an explicitly owned resource. Explain each option first. These queries are read-only; a successful response does not authorize a later create or delete operation.
Practical work
Create p06-backup-architecture.md for production S3, EBS, EFS, and RDS resources. Define tag-based selection, schedules, windows, continuous versus snapshot behavior where supported, standard and isolated vaults, KMS ownership, cross-account and cross-Region copies, retention, Vault Lock review, restore-test schedule, RPO/RTO evidence, alarm ownership, cost, and recovery cleanup. Do not enable Vault Lock or create a recovery point.
Create a coverage matrix with resource ARN pattern/tag, business owner, data classification, RPO/RTO, native/AWS Backup mechanism, schedule, vault/account/Region, KMS owner, retention, copy lag alarm, restore frequency and latest successful application validation. Include deliberate tests for an untagged new resource, denied KMS grant, failed copy, expired recovery point, unavailable production account and restore into a clean recovery account. Reconcile inventory to protected-resource and job evidence; a 100% successful-job rate can still hide resources never selected.
Write a recovery runbook that starts from an incident timestamp, selects the right recovery point, obtains independent approval, restores into isolation, validates data and application dependencies, records achieved RPO/RTO, controls cutover/failback and deletes temporary recovery infrastructure. Estimate warm and cold storage, copy, transfer, restore, testing-resource and immutable-retention costs.
The evidence package must contain:
- the problem and final requirement in the learner's own words;
- caller type and Region with private identifiers redacted;
- exact planned values, ownership, and cost class;
- one Console observation and matching CLI or API evidence;
- one behavior result or supplied data-plane record;
- one denied, failed, or counterexample result and evidence-led diagnosis;
- one architecture choice plus the rejected alternative;
- cleanup proof or explicit retained-state owner, expiry, and next lesson.
Verification standard
Use expected state before observed state. Record timestamps in UTC and preserve the original failure before changing anything. A passing submission answers all four questions:
- What exact requirement was tested?
- Which evidence proves the AWS configuration?
- Which evidence proves the workload behavior?
- What remains unproven or requires later monitoring?
If AWS returns no rows, verify account, Region, permission, filters, pagination, resource type, and deletion state before concluding that nothing exists.
Common failures and troubleshooting
| Symptom | Evidence first | Likely boundary | Smallest safe response |
|---|---|---|---|
| object appears missing | caller, Region, filters, pagination, tags | scope or read permission | align scope before creating a duplicate |
| state remains pending or unavailable | service state, events, dependencies, quotas | dependency or capacity | correct the named dependency and wait with a bound |
| AccessDenied | principal, action, resource, explicit-deny context | identity, resource, endpoint, organization, or KMS policy | change only the proven policy layer |
| configuration exists but behavior fails | route, DNS, security, listener, health, logs, object version | data path or application | test the next boundary and change one control |
| bill is higher than expected | hours, bytes, requests, AZs, addresses, retention | cost model or retained resource | stop optional work and reconcile the ledger |
| cleanup is blocked | dependency inventory and owning service | deletion order or immutable state | remove owned dependants in reviewed reverse order |
| resource has no recovery point | inventory versus selection tags/conditions and plan scope | coverage gap | correct selection and alert on unprotected inventory; do not fabricate historical coverage |
| backup/copy job fails | job status message, service role, vault policy, KMS and support matrix | permission/encryption/unsupported path | repair the named dependency and rerun within RPO bounds |
| recovery point exists but restore is denied | restore role, destination account policy, SCP and KMS key state | recovery authorization | use tested break-glass path; never weaken organization-wide controls blindly |
| restored infrastructure is unusable | dependency manifest, network/secret/config and app integrity checks | application recovery | reconstruct versioned dependencies and validate above resource status |
| retention cannot be shortened | vault lock mode, min/max retention and cooling-off state | intentional immutability | escalate to governance/legal/cost owners; compliance lock is not bypassable |
| dashboard is green but a resource is absent | independent inventory and assignment evidence | observability blind spot | add coverage reconciliation and Audit Manager control |
Do not troubleshoot by attaching administrator access, opening administration ports to the internet, disabling encryption, retrying uncontrolled creation, deleting unknown resources, or weakening retention.
Cost, cleanup, and retained state
No AWS resource is created. Close CloudShell and remove or redact downloaded evidence.
Cleanup evidence requires terminal state and an after-inventory. Search related ENIs, public IPv4 addresses, EBS volumes and snapshots, load balancers, target groups, Auto Scaling instances, endpoints, logs, S3 versions and delete markers, backup recovery points, and global IAM roles when they apply. Billing data can lag, so schedule a later review.
Architecture and certification decisions
- Certification coverage: SAA-C03; SOA-C03; SAP-C02; DOP-C02.
- Exam mapping: SAA D1-D4.
- Explain service scope, failure boundary, consistency, recovery, security, operations, and price rather than matching a keyword.
- Treat availability and durability, encryption and authorization, routing and filtering, health and lifecycle, backup and replication, and discount and capacity as separate concepts.
- Do not reproduce protected certification questions.
Knowledge check
- Does a completed backup prove recovery?
Expected direction: No. Perform and validate a restore.
- What adds administrative isolation?
Expected direction: A supported cross-account copy or logically air-gapped vault design.
- Can compliance Vault Lock be removed after grace time?
Expected direction: No. Treat it as an irreversible governance action.
- What must cross-account backup also consider?
Expected direction: Organizations, vault policy, IAM, KMS, service support, and restore procedure.
Completion gate and assessment
| Area | Points | Passing evidence |
|---|---|---|
| Requirement and model | 15 | Correct scope, terminology, and final outcome |
| Console evidence | 15 | Current path and interpreted fields |
| CLI or API evidence | 15 | Scoped command, expected result, and limitations |
| Behavior or decision exercise | 20 | Reproducible result or defensible architecture reasoning |
| Troubleshooting | 15 | Original symptom, hypothesis, one change, retest, rollback |
| Security and cost | 10 | Least privilege, data protection, current price dimensions |
| Cleanup and handoff | 10 | Terminal-state proof or approved retained-state record |
Pass at 80 out of 100 with no critical safety failure. A missing practical artifact, unexplained output, unsafe access, destructive action outside the owned scope, unplanned billed resource, or false cleanup claim requires remediation and a changed retest.