Lesson 119 · AWS Learning Path

AWS 119: AWS Backup and cross-account backup

· Published · 12 min read

A trusted principal crosses an account boundary by assuming a role and receives temporary credentials

The real problem

A team recognizes the name AWS Backup and cross-account backup but has not connected the feature to a real requirement, identity boundary, network or data path, failure mode, price dimension, and cleanup owner. A plausible configuration could still fail the workload.

Final outcome

The learner will produce a requirement-led artifact for AWS Backup and cross-account backup, inspect the matching AWS control plane in the Management Console, run a matching CloudShell or AWS CLI query, interpret the output, diagnose one failure, defend one architecture choice, and prove cleanup or approved retained state.

The practical outcome is not a command transcript. It must show what was expected, what happened, what the result proves, what it does not prove, and which evidence would change the decision.

Learning objectives

By the end of this lesson, the learner can:

  • explain central policy;
  • explain vault boundary;
  • explain cross-account copy;
  • explain cross-region copy;
  • explain vault lock;
  • connect control-plane state to the real data, network, identity, or application behavior;
  • identify cost and cleanup ownership before any optional mutation;
  • troubleshoot from evidence without opening broad access or adding broad permissions.

Relationship model

Requirement
   |
   v
Identity and policy -> AWS configuration -> network or data path -> workload behavior
        |                    |                      |                    |
        +--------------------+----------------------+--------------------+
                                      |
                                      v
                         monitoring, cost, recovery, cleanup

Use this model to separate an AWS object that exists from a result that actually works. Every arrow is a verification boundary.

Prerequisites, permissions, Region, and safety

  • Learning baseline: This sequence assumes practical Linux knowledge but no prior cloud-computing or AWS knowledge. Cloud, networking, security, data, automation, and architecture concepts must come from completed earlier lessons. If a prerequisite checkpoint is incomplete, return to its linked lesson before continuing.
  • Confirm a non-root caller with aws sts get-caller-identity and keep the account number private.
  • Use ap-south-1 unless this lesson explicitly names a second Region.
  • Confirm the intended profile and Region with aws configure list before interpreting an empty result.
  • Use read-only List, Get, and Describe permissions for the named services. Design exercises run locally and require no resource-creation permission.
  • This is a no-create lesson. Console and CLI work is read-only, and every design artifact is created locally.
  • Never publish account IDs, public addresses, ARNs containing private account data, session IDs, presigned URLs, object data, credentials, or KMS material.
  • Do not use root, world-open SSH or RDP, disabled TLS verification, unowned resources, or irreversible retention controls in a training exercise.

Core model

ConceptWhat the learner must understand
Central policyAWS Backup uses plans, rules, schedules, windows, lifecycle, vaults, selections, IAM roles, copy actions, and recovery points to centralize supported-service protection. Coverage varies by service and Region.
Vault boundaryA backup vault organizes recovery points and has encryption and access-policy behavior. Source-service permissions can still matter for some backup types.
Cross-account copyAWS Backup can copy supported backups to another account in an AWS Organization when vault policy, organization settings, IAM, encryption keys, and service support align.
Cross-Region copyCopy actions can create recovery points in another Region. RPO, copy completion, encryption, data transfer, retention, and restore location need monitoring.
Vault LockGovernance mode allows authorized management. Compliance mode becomes immutable after its grace time and can force retained storage cost until recovery points expire.
Logically air-gapped vaultThis specialized vault includes compliance-mode Vault Lock, service-account isolation, sharing and recovery features, and additional availability and resource-type considerations.

How it works

A backup is useful only when the correct resource, point in time, encryption keys, metadata, role, network, dependencies, and application validation can be restored within RTO. Schedule restore tests and record evidence, not only backup-job success.

Begin with business impact. RPO is the maximum acceptable data-loss window; it drives frequency or continuous-backup needs. RTO is the maximum acceptable time to restore usable service; it includes authorization, infrastructure reconstruction, data restore, application validation, DNS/routing and decision time. Retention states how long recovery points must remain. Availability replicas and data replication can reduce outage time but may copy corruption or deletion; backups provide historical recovery points. Neither replaces the other.

Control and data model

An AWS Backup plan contains rules for schedule, start/completion windows, target vault, lifecycle and copy actions. A selection maps resources explicitly or by tags and uses a service role. A job attempts the backup/copy/restore. A recovery point is the retained result in a vault. Frameworks and Audit Manager controls assess evidence; reports and EventBridge/CloudWatch alarms expose failed or missing jobs. Coverage and exact feature behavior vary by resource type and Region, so the architecture must maintain a dated support matrix.

Encryption behavior is not uniform. Some resources are fully managed by AWS Backup while others retain source-service snapshot behavior; source and destination KMS keys, key policies, grants and account ownership affect copy and restore. Deleting or disabling a KMS key can make retained encrypted recovery points unusable. Cross-account copy normally requires supported resource type, Organizations trust/settings, destination vault access policy, source/copy roles and compatible encryption. Cross-Region and cross-account are separate dimensions and produce separate recovery points with transfer/storage and completion lag.

Immutability and isolation

Standard vaults organize and policy-protect recovery points. Vault Lock governance mode can be changed by sufficiently privileged identities; compliance mode becomes immutable after its cooling-off/grace period. Once locked, neither the customer nor AWS can shorten retention or delete protected recovery points until expiry. This is a legal, operational and cost commitment - never enable it as an exploratory lab action.

A logically air-gapped vault is a specialized vault with compliance-mode lock and backup data isolated in an AWS Backup service-owned account. Current capabilities include RAM-based sharing/restore patterns and Multi-party approval options, subject to Region/resource restrictions. It does not mean a physical offline tape and does not excuse testing identities, KMS behavior, sharing and restore access. Separate backup/security and recovery accounts reduce the blast radius of compromised production administrators only if SCPs, break-glass roles, keys and approval paths are independently controlled.

Restore is the product

Restore to a separate, quarantined destination. Reconstruct dependencies - VPC/subnets/SGs, IAM roles, KMS keys, secrets, parameter groups, DNS and application configuration - then validate schema, row/object/file counts, checksums and business transactions. Scan for malware where required before reconnecting. Measure requested-to-usable time and clean up the restored resources. Automated restore testing proves that AWS can perform configured restore steps for supported resources; application-level validation and dependency exercises remain your responsibility.

Read the result in layers:

  1. Scope: account, Region, VPC, bucket, AZ, endpoint, principal, object version, or resource ARN.
  2. Control plane: the requested configuration exists and reached an expected state.
  3. Behavior: the request, connection, health check, replication, restore, or application result meets the requirement.
  4. Operations: monitoring, failure owner, cost, retention, rollback, and cleanup are known.

Control-plane success is necessary but not sufficient. A resource can be available while policy, routing, DNS, health, data, or application behavior remains wrong.

Architecture decision table

RequirementPreferred directionWhy
Central schedules across supported servicesAWS Backup planOne policy coordinates selections, vaults, lifecycle, and copies.
Reduce production-account deletion blast radiusCross-account copy or logically air-gapped vault by requirementAdministrative isolation adds recovery protection.
Non-bypassable regulatory retentionCompliance Vault Lock after formal reviewThe setting becomes immutable and cost persists until retention completes.
Need evidence that backup worksAutomated or scheduled restore testA completed backup job alone does not prove recovery.

Professional questions normally contain several valid services. State the requirement that selects one option, why the nearest alternative fails it, and what changed requirement would reverse the choice.

AWS Management Console guided practice

Before opening a service page, write the expected account, Region, starting state, and evidence. Do not choose Create, Save, Purchase, Lock, or Delete unless the lesson explicitly authorizes the live track.

  1. Open AWS Backup Dashboard, Backup plans, Backup vaults, Protected resources, Jobs, and Restore testing; inspect supplied evidence without creating a plan.
  2. Open a vault and record type, encryption key, access policy, lock mode and date, retention boundaries, recovery points, copy state, and sharing scope.
  3. Trace one recovery point through restore role, metadata, destination network, validation, cleanup, measured RTO, and audit record.

For each step, capture the field name and value in text. A screenshot may support the record but does not replace the explanation. Console labels can evolve, so use the service search and current documentation if a navigation label differs.

CloudShell and AWS CLI practice

CloudShell is the default browser-based command environment taught in AWS 028. AWS 029 and AWS 030 cover local CLI installation and authentication. This lesson therefore does not assume that an unconfigured local shell is ready.

Start every session with:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account portion of the ARN before sharing. Then perform the topic query:

Inventory plans, vaults, lock state, recovery-point count, and recent backup and restore jobs without creating one.

aws backup list-backup-plans --query 'BackupPlansList[].{Id:BackupPlanId,Name:BackupPlanName,Version:VersionId}' --output table
aws backup list-backup-vaults --query 'BackupVaultList[].{Name:BackupVaultName,Type:VaultType,Locked:Locked,Points:NumberOfRecoveryPoints,Key:EncryptionKeyArn}' --output table
aws backup list-restore-jobs --by-status COMPLETED --output table

Expected interpretation:

Recovery points and completed jobs support backup evidence. They do not prove the right data, application consistency, cross-account access, dependency reconstruction, or achieved RTO.

Replace every replace-with-... sample value before running its command, and use only an explicitly owned resource. Explain each option first. These queries are read-only; a successful response does not authorize a later create or delete operation.

Practical work

Create p06-backup-architecture.md for production S3, EBS, EFS, and RDS resources. Define tag-based selection, schedules, windows, continuous versus snapshot behavior where supported, standard and isolated vaults, KMS ownership, cross-account and cross-Region copies, retention, Vault Lock review, restore-test schedule, RPO/RTO evidence, alarm ownership, cost, and recovery cleanup. Do not enable Vault Lock or create a recovery point.

Create a coverage matrix with resource ARN pattern/tag, business owner, data classification, RPO/RTO, native/AWS Backup mechanism, schedule, vault/account/Region, KMS owner, retention, copy lag alarm, restore frequency and latest successful application validation. Include deliberate tests for an untagged new resource, denied KMS grant, failed copy, expired recovery point, unavailable production account and restore into a clean recovery account. Reconcile inventory to protected-resource and job evidence; a 100% successful-job rate can still hide resources never selected.

Write a recovery runbook that starts from an incident timestamp, selects the right recovery point, obtains independent approval, restores into isolation, validates data and application dependencies, records achieved RPO/RTO, controls cutover/failback and deletes temporary recovery infrastructure. Estimate warm and cold storage, copy, transfer, restore, testing-resource and immutable-retention costs.

The evidence package must contain:

  • the problem and final requirement in the learner's own words;
  • caller type and Region with private identifiers redacted;
  • exact planned values, ownership, and cost class;
  • one Console observation and matching CLI or API evidence;
  • one behavior result or supplied data-plane record;
  • one denied, failed, or counterexample result and evidence-led diagnosis;
  • one architecture choice plus the rejected alternative;
  • cleanup proof or explicit retained-state owner, expiry, and next lesson.

Verification standard

Use expected state before observed state. Record timestamps in UTC and preserve the original failure before changing anything. A passing submission answers all four questions:

  1. What exact requirement was tested?
  2. Which evidence proves the AWS configuration?
  3. Which evidence proves the workload behavior?
  4. What remains unproven or requires later monitoring?

If AWS returns no rows, verify account, Region, permission, filters, pagination, resource type, and deletion state before concluding that nothing exists.

Common failures and troubleshooting

SymptomEvidence firstLikely boundarySmallest safe response
object appears missingcaller, Region, filters, pagination, tagsscope or read permissionalign scope before creating a duplicate
state remains pending or unavailableservice state, events, dependencies, quotasdependency or capacitycorrect the named dependency and wait with a bound
AccessDeniedprincipal, action, resource, explicit-deny contextidentity, resource, endpoint, organization, or KMS policychange only the proven policy layer
configuration exists but behavior failsroute, DNS, security, listener, health, logs, object versiondata path or applicationtest the next boundary and change one control
bill is higher than expectedhours, bytes, requests, AZs, addresses, retentioncost model or retained resourcestop optional work and reconcile the ledger
cleanup is blockeddependency inventory and owning servicedeletion order or immutable stateremove owned dependants in reviewed reverse order
resource has no recovery pointinventory versus selection tags/conditions and plan scopecoverage gapcorrect selection and alert on unprotected inventory; do not fabricate historical coverage
backup/copy job failsjob status message, service role, vault policy, KMS and support matrixpermission/encryption/unsupported pathrepair the named dependency and rerun within RPO bounds
recovery point exists but restore is deniedrestore role, destination account policy, SCP and KMS key staterecovery authorizationuse tested break-glass path; never weaken organization-wide controls blindly
restored infrastructure is unusabledependency manifest, network/secret/config and app integrity checksapplication recoveryreconstruct versioned dependencies and validate above resource status
retention cannot be shortenedvault lock mode, min/max retention and cooling-off stateintentional immutabilityescalate to governance/legal/cost owners; compliance lock is not bypassable
dashboard is green but a resource is absentindependent inventory and assignment evidenceobservability blind spotadd coverage reconciliation and Audit Manager control

Do not troubleshoot by attaching administrator access, opening administration ports to the internet, disabling encryption, retrying uncontrolled creation, deleting unknown resources, or weakening retention.

Cost, cleanup, and retained state

No AWS resource is created. Close CloudShell and remove or redact downloaded evidence.

Cleanup evidence requires terminal state and an after-inventory. Search related ENIs, public IPv4 addresses, EBS volumes and snapshots, load balancers, target groups, Auto Scaling instances, endpoints, logs, S3 versions and delete markers, backup recovery points, and global IAM roles when they apply. Billing data can lag, so schedule a later review.

Architecture and certification decisions

  • Certification coverage: SAA-C03; SOA-C03; SAP-C02; DOP-C02.
  • Exam mapping: SAA D1-D4.
  • Explain service scope, failure boundary, consistency, recovery, security, operations, and price rather than matching a keyword.
  • Treat availability and durability, encryption and authorization, routing and filtering, health and lifecycle, backup and replication, and discount and capacity as separate concepts.
  • Do not reproduce protected certification questions.

Knowledge check

  1. Does a completed backup prove recovery?

Expected direction: No. Perform and validate a restore.

  1. What adds administrative isolation?

Expected direction: A supported cross-account copy or logically air-gapped vault design.

  1. Can compliance Vault Lock be removed after grace time?

Expected direction: No. Treat it as an irreversible governance action.

  1. What must cross-account backup also consider?

Expected direction: Organizations, vault policy, IAM, KMS, service support, and restore procedure.

Completion gate and assessment

AreaPointsPassing evidence
Requirement and model15Correct scope, terminology, and final outcome
Console evidence15Current path and interpreted fields
CLI or API evidence15Scoped command, expected result, and limitations
Behavior or decision exercise20Reproducible result or defensible architecture reasoning
Troubleshooting15Original symptom, hypothesis, one change, retest, rollback
Security and cost10Least privilege, data protection, current price dimensions
Cleanup and handoff10Terminal-state proof or approved retained-state record

Pass at 80 out of 100 with no critical safety failure. A missing practical artifact, unexplained output, unsafe access, destructive action outside the owned scope, unplanned billed resource, or false cleanup claim requires remediation and a changed retest.

Official sources

Advertisement