Lesson 297 · AWS Learning Path

AWS 297: Cross-account and cross-Region backup isolation and ransomware considerations

· Published · 16 min read

Labelled process diagram for AWS 297: Protected workload and policy to Independent encrypted backup copy to Isolated recovery access and clean room to Validated restore and retained audit evidence, with decision...

Why this lesson matters

A backup stored beside production under the same administrator, organization path, Region, and encryption-key authority can disappear with production. A copied and immutable recovery point improves survival, but it can still contain malware, encrypted files, poisoned configuration, or compromised identities.

Ransomware-resilient recovery therefore needs several independent properties: known coverage, separate ownership, geographic diversity where required, deletion resistance, usable encryption keys, restricted restore access, clean recovery infrastructure, corruption-aware point selection, validation, and practiced business recovery.

This lesson teaches how AWS Backup, AWS Organizations, AWS KMS, Vault Lock, logically air-gapped vaults, AWS RAM, multi-party approval, restore testing, and GuardDuty malware scanning can contribute. No single feature creates an “air gap” for the whole organization. The architect must prove each trust boundary against an explicit threat.

What you will be able to do

By the end, you can:

  • distinguish availability, durability, immutability, isolation, confidentiality, and recoverability;
  • model compromised workload, administrator, organization, Region, KMS, pipeline, and backup-service paths;
  • design cross-account and cross-Region copies without assuming universal feature support;
  • explain destination-vault ownership, policies, KMS re-encryption, and organization requirements;
  • choose governance or compliance Vault Lock safely;
  • explain logically air-gapped vault access, RAM sharing, and multi-party approval boundaries;
  • identify a clean recovery point using timeline, scanning, and application evidence;
  • restore into a clean room without reconnecting contamination to production;
  • prove restoration and business recovery with scheduled tests; and
  • build cost, retention, quota, legal-hold, and evidence governance.

Before you start

  • This is a no-create lesson. Do not create, lock, share, copy, restore, scan, delete, or place a legal hold on a recovery point.
  • A compliance-mode Vault Lock becomes unchangeable after its grace period. An incorrect indefinite retention can create permanent cost and data-governance consequences. Never use it as a learning experiment.
  • Real backup inventories and threat gaps are highly sensitive. Use supplied identifiers and redact account IDs, vaults, keys, plans, resources, findings, and recovery locations.
  • Verify current feature availability by resource type and Region. AWS Backup support does not mean every resource supports cross-account copy, cross-Region copy, cold tier, restore testing, malware scanning, or every combination.
  • Legal, privacy, records, security, continuity, finance, and application owners must approve retention and recovery design.

1. Define what each protection property means

PropertyMeaningWhat it does not prove
AvailabilityBackup control/data can be reached when neededData is clean or immutable
DurabilityStored data is engineered against media lossAccount authority cannot delete it
ImmutabilityRecovery point cannot be shortened/deleted before lifecyclePoint is malware-free or useful
IsolationCompromise paths are separated across identities/accounts/organizations/RegionsNo shared dependency remains
ConfidentialityUnauthorized parties cannot read backup dataAuthorized compromise cannot destroy it
IntegrityData is unchanged and internally validApplication can start with all dependencies
RecoverabilityData and service can be restored and accepted in timeFuture restores will pass without repeated tests

Use these words precisely. “Encrypted” is not “isolated.” “Cross-Region” is not “cross-account.” “Locked” is not “clean.” “Restore job completed” is not “business recovered.”

2. Build the ransomware threat model

Analyze at least these adversary positions:

  1. Workload credentials compromised: attacker can encrypt/delete application data and call APIs available to the workload.
  2. Application administrator compromised: attacker can change backup selections, tags, schedules, or source retention.
  3. Backup operator compromised: attacker can alter plans, vault policy, copies, or restores within granted permissions.
  4. Source-account administrator or root compromised: attacker has broad authority but should not control the destination account and immutable points.
  5. Organizations management/delegated administrator compromised: attacker can change policies, trusted access, accounts, or recovery governance.
  6. KMS authority compromised or key unavailable: attacker disables/deletes a customer-managed key, or recovery cannot decrypt.
  7. CI/CD or infrastructure-code compromise: malicious settings are redeployed into production and recovery.
  8. Region unavailable: source data, APIs, keys, network, and operators in one Region are inaccessible.
  9. Malware/corruption already inside backups: the protection system faithfully stores bad state.
  10. Insider or collusion: one authorized person abuses destructive or restore privileges.

For each, record capability, reachable control plane, expected detection, prevention, immutable copy, recovery identity, clean-room path, residual risk, and test evidence. Do not claim protection against compromise of the same authority that can rewrite every protection layer.

3. Understand the AWS Backup control model

Plans, rules, selections, jobs, vaults, and points

A backup plan contains rules such as schedule, start/completion windows, lifecycle, destination vault, and copy actions. Resource assignments select resources through explicit identifiers, tags, or supported organization policies. A backup job creates a recovery point in a vault. Copy jobs create separate recovery points in destination vaults.

Confirm all of these:

  • expected resources selected, including newly created and untagged exceptions;
  • backup and copy job completion within RPO window;
  • retention/lifecycle on each resulting point;
  • destination account, Region, vault, key, and lock;
  • protected resource's application consistency and restore metadata;
  • notifications, EventBridge events, CloudWatch/CloudTrail evidence, Audit Manager controls, and ownership; and
  • failed, expired, partial, or continuously protected behavior by resource type.

An AWS Organizations backup policy can help apply plans at scale. It is governance configuration, not a stored copy. Test effective policies and resource selection in member accounts. Protect management and delegated-administrator privileges because central control is also a concentration of risk.

Independent copies

The first cross-account copy is generally full; later copies can be incremental where the service supports it. The destination recovery point is owned and encrypted in the destination context. Copy completion can lag source backup completion, so compute the real protected point age at the isolated destination.

Use at least two dimensions only where the threat and BIA justify them:

  • same account, cross-Region for Regional failure but not account compromise;
  • cross-account, same Region for account isolation but not Regional failure; or
  • cross-account and cross-Region for both dimensions, with higher cost and complexity.

Do not create one destination account with the same federated administrator, permission sets, pipelines, root-email process, and KMS administrators as production, then call it isolated. Separate routine access, organizational placement/SCPs, credentials, approval, logging, alerting, and recovery duties.

4. Cross-account copy requirements and ownership

AWS Backup cross-account copy requires source and destination accounts in the same AWS Organization for the documented workflow. The destination uses a non-default vault, destination vault access policy, and a destination encryption key that supports the resource behavior. Source-side policies and roles must permit the copy.

Protect the destination account from leaving the organization unexpectedly. AWS notes that a destination account that leaves can retain copied backups; an SCP can deny organizations:LeaveOrganization where appropriate. Understand that organization-level compromise can still affect governance, which is one reason newer multi-party recovery designs may use a separate recovery organization.

Destination operators should be able to monitor and recover without gaining routine production administration. Production operators should be able to initiate expected copies without deleting destination recovery points or weakening locks. Keep break-glass restore and key access unavailable to everyday sessions.

Copy shutdown is not always instantaneous because of eventual consistency. Monitor jobs and destination points after disabling a copy path.

5. Make KMS part of the recovery architecture

Cross-account and cross-Region copies are encrypted using the destination vault/resource rules. Behavior differs between resources fully managed by AWS Backup and those whose snapshots remain managed by the originating service.

AWS managed KMS key policies cannot be edited to share cross-account. For resource types that are not fully managed by AWS Backup, cross-account copy generally needs a customer-managed key or a supported service-specific method. The destination key policy, grants, IAM role, vault policy, and source key access must all align.

For every data class prove:

  • source encryption type and key owner;
  • whether the resource is fully managed by AWS Backup;
  • supported destination vault/key choice;
  • source decrypt and copy permissions;
  • destination encrypt/decrypt and restore permissions;
  • key policy and grant behavior;
  • key deletion/disable controls and alerts;
  • multi-Region or separate Regional-key decision; and
  • emergency recovery access independent of the failed identity path.

Immutability of a recovery point does not help if its only usable KMS key is scheduled for deletion or inaccessible during an organization compromise. Protect KMS administration with separate roles, MFA/emergency process, SCP and key-policy guardrails, CloudTrail alerts, and tested recovery use. Do not give backup job roles broad key administration.

6. Use Vault Lock with deliberate retention

Vault Lock enforces write-once, read-many controls for recovery-point retention.

Governance mode

Governance mode denies early deletion/lifecycle alteration to ordinary actors, but sufficiently privileged users can remove the lock or manage the vault. It is useful for controlled governance and rehearsal, but its threat model does not include compromise of every privileged lock administrator.

Compliance mode

Compliance mode includes a configurable grace period. During grace, authorized users can correct or remove the lock. After it expires, the lock configuration cannot be changed or removed by the customer, account/data owner, root user, or AWS while protected points remain under their lifecycle. This is powerful and intentionally unforgiving.

Minimum and maximum retention guard future backup/copy jobs: jobs whose lifecycle is outside the permitted range fail. They do not rewrite recovery points that already existed before lock activation. The recovery point's own lifecycle remains important.

Before compliance lock:

  1. Approve legal, privacy, security, finance, and business retention requirements.
  2. Remove accidental indefinite retention unless explicitly required.
  3. Test backup and copy rules against minimum/maximum limits.
  4. Test expiry, restore, clean-room recovery, legal hold, monitoring, and cost forecasts.
  5. Prove KMS, destination account, quota, and emergency-access behavior.
  6. Run governance mode or an equivalent pilot long enough to expose mistakes.
  7. Use the grace period for formal validation and sign-off.
  8. Record who accepts the irreversible consequence.

Lock configuration errors can block future jobs or preserve sensitive data and cost for years. “More immutable” is not automatically “more compliant.”

7. Understand logically air-gapped vaults

A logically air-gapped vault provides additional isolation features, is encrypted with an AWS-owned key in the documented model, and is locked in compliance mode. It can be shared for recovery through AWS Resource Access Manager according to current feature support. Recovery points are still online AWS resources, so “logical air gap” does not mean physically offline media.

Use it to reduce shared control and key-management paths, not as a marketing checkbox. Validate:

  • resource and Region feature availability;
  • retention settings and compliance lock consequence;
  • vault-owner, backup-copy, sharing, restore, and KMS responsibilities;
  • AWS RAM resource share and permission scope;
  • whether the recovery account can create the required backup access/restore path;
  • revocation behavior and what happens to existing access points;
  • CloudTrail, alerts, inventory, and cost visibility; and
  • restore into an isolated account with no routine production trust.

Multi-party approval

AWS Backup supports multi-party approval workflows for restore access to logically air-gapped vaults. A separate recovery organization and approval team can reduce the chance that one compromised organization administrator grants restore access alone. Requester, administrator, and approver duties should be held by different people.

The workflow integrates AWS Organizations, Account Management, IAM, AWS RAM, AWS Backup, and the Multi-party Approval service. Current documentation places approval-team resources and related commands in a specific control Region, so verify the design and service availability. Multi-party approval guards access approval; it does not scan data, build the clean room, validate the restore, or decide which point is safe.

8. Detect a clean recovery point

Ransomware discovery time is not infection time. Build an incident timeline from endpoint/security findings, CloudTrail, identity events, configuration changes, file/object changes, database/audit logs, backup/copy jobs, malware scans, and business anomalies.

Classify points:

  • known compromised;
  • suspected contamination window;
  • pre-compromise candidate;
  • scanned with stated engine/date/scope/result;
  • application-consistency tested;
  • business-accepted clean point; or
  • unknown.

AWS Backup integrates with GuardDuty Malware Protection for supported resources and Regions. Automatic scans can follow backups and on-demand scans can analyze existing points. Before restore, use the current guidance for a full scan with the latest detection model where supported. A “no threat found” result is not absolute proof: signature/model coverage, encrypted archives, unsupported file types, application backdoors, malicious IAM/configuration, and data corruption can remain.

Combine scanning with timeline analysis, threat hunting, integrity checks, vulnerability/configuration review, application startup, transaction reconciliation, credential/key rotation, and business validation. Preserve forensic evidence separately from the recovery copy.

9. Restore into a clean room

Never restore a suspected point directly over production. Build or pre-stage an isolated clean-room account and network with:

  • recovery identity independent of compromised federation;
  • restricted ingress/egress and no automatic peering to production;
  • known-clean infrastructure code, images, packages, and tools;
  • separate keys, secrets, certificates, DNS names, and service roles;
  • GuardDuty/security tooling, logs, packet/flow evidence, and time synchronization;
  • no outbound email, payments, jobs, webhooks, or destructive automation;
  • patched analysis hosts and least-privilege examiner access;
  • quarantined data transfer and explicit promotion gates; and
  • sufficient quotas, IP space, compute, storage, and licenses.

Recovery sequence:

  1. Establish incident authority and preserve evidence.
  2. Secure independent credentials and verify the destination inventory.
  3. Select candidate points outside the suspected window.
  4. Grant time-bounded restore access through the approved control path.
  5. Restore into quarantine with no production side effects.
  6. Scan, inspect, patch, rotate credentials, and remove persistence.
  7. Validate filesystem/database integrity and application consistency.
  8. Reconcile critical business records and determine effective RPC/RPO.
  9. Rebuild application/control configuration from known-clean sources when safer than restoring it.
  10. Obtain security and business acceptance before controlled promotion.

Restoring data while redeploying compute/configuration from clean artifacts is often safer than restoring entire compromised servers. Decide per resource.

10. Prove restore capability continuously

AWS Backup restore testing plans can select eligible recovery points, start restores on a schedule, infer required metadata, and record completion time. Optional validation workflows can submit validation results. The service cleans up test resources after the validation window according to its behavior, but operators must verify deletion status and unexpected retained dependencies.

A restore job reaching COMPLETED proves resource creation, not business usability. Validation must test:

  • decryption and emergency identity;
  • isolation and no prohibited egress;
  • boot/mount/open and engine recovery;
  • schema, checksums/control totals, object counts, and sampled values;
  • point age and effective RPC;
  • application startup and dependency compatibility;
  • synthetic critical transaction and resulting durable state;
  • recovery time from request through business acceptance;
  • malware/security findings and clean configuration; and
  • cleanup, retained evidence, and cost.

Random point selection can test general recoverability. Known clean candidates and oldest/edge-of-retention points test other risks. Define a rotation so all critical resource types, accounts, Regions, keys, tiers, and recovery teams receive coverage.

11. Read-only inspection

Only in authorized accounts. Start narrow and redact names:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text

aws backup list-backup-vaults \
  --query 'BackupVaultList[].{Name:BackupVaultName,Points:NumberOfRecoveryPoints,Locked:Locked,Min:MinRetentionDays,Max:MaxRetentionDays}'

aws backup list-backup-plans \
  --query 'BackupPlansList[].{Id:BackupPlanId,Name:BackupPlanName,Version:VersionId}'

aws backup list-restore-testing-plans \
  --query 'RestoreTestingPlans[].{Name:RestoreTestingPlanName,Schedule:ScheduleExpression,StartWindow:StartWindowHours}'

For one explicitly owned vault, inspect its access policy, lock configuration, recovery-point inventory, copy jobs, restore jobs, tags, notifications, and key metadata with read-only operations. Verify the caller account and Region before every command; a vault name can exist in more than one Region.

Organizations backup policies are visible only with appropriate organization permission. Do not grant broad organization visibility merely to complete a lesson. Use supplied effective-policy evidence when access is unavailable.

12. Three-data-class architecture workshop

Design protection for fictional Cedar Health Logistics:

  • Class A orders: RPO 5 minutes, seven-year records retention, cross-Region recovery, customer-managed encryption, 8 TB database plus 40 TB documents;
  • Class B application/configuration: RPO 24 hours, rebuild preferred, 90-day retention, repositories and pipelines may be compromised;
  • Class C analytics: RPO 24 hours, 30-day retention, reproducible from clean Class A data;
  • 18 production accounts in one AWS Organization;
  • one backup destination account in a security OU;
  • one separate recovery organization/account available for multi-party restore design;
  • suspected privileged-administrator and CI/CD compromise; and
  • infection may have existed for 12 days before detection.

Produce:

  1. Business objectives, data ownership, legal/privacy retention, and deletion requirements.
  2. Ten-scenario ransomware/insider/Region/KMS threat matrix.
  3. Source, same-account, cross-account, cross-Region, and recovery-organization trust diagram.
  4. Resource-by-resource AWS Backup feature matrix from current documentation.
  5. Plan/rule/selection/copy schedule meeting RPO and copy-lag budgets.
  6. Destination-vault policy, organization/SCP, delegated-admin, and routine-access boundaries.
  7. Source/destination KMS key, grant, deletion, emergency decrypt, and audit design.
  8. Governance versus compliance Vault Lock decision with retention simulations.
  9. Logically air-gapped vault, RAM, access-point, and multi-party approval decision.
  10. Detection, alerting, Audit Manager, copy-failure, policy-change, lock, KMS, and organization-leave evidence.
  11. Twelve-day contamination timeline and clean-point selection procedure.
  12. Malware-scan applicability and residual-risk analysis.
  13. Clean-room network, identity, toolchain, data, side-effect, and promotion gates.
  14. Scheduled restore-test matrix and application/business validation.
  15. Cost, quotas, legal holds, cleanup, key/vault decommission, and quarterly exercise plan.

Inject these supplied failures: untagged database omitted, copy job outside the 5-minute objective, source encrypted under an unshareable AWS managed key, destination lock retention mismatch, destination account allowed to leave the organization, compromised federation reaches recovery, key scheduled for deletion, latest points contain malware, restore completes but sequence values are wrong, and test resources fail cleanup.

13. Cost, quota, retention, and deletion governance

Model backup storage by class and lifecycle, cross-Region/account transfer, copy requests/jobs, warm/cold tier where supported, vault lock retention, logically air-gapped storage, restore testing resources, malware scans, indexing/search if used, KMS, CloudTrail/logging, Audit Manager, clean-room baseline, security tools, support, staff, and exercises. Ransomware response may restore several candidate points simultaneously.

Long retention under compliance lock is an enforceable cost, not merely a forecast. Model data growth, incremental/full behavior by service, legal holds, minimum/maximum lock bounds, duplicated Regions/accounts, and deletion after expiry.

Check quotas for plans, selections, vaults, points, concurrent backup/copy/restore/index/scan jobs, legal holds, KMS, EC2/EBS/RDS and clean-room resources. Queueing copy jobs can make the isolated point older than the RPO even when source backup is timely.

Deletion governance must reconcile privacy deletion, legal hold, regulatory minimum retention, backup lifecycle, and immutable-vault constraints before data enters the vault. Document how expired recovery points, keys, vaults, accounts, RAM shares, recovery organizations, and metadata are retired. Never schedule key deletion as a shortcut for normal lifecycle management.

Diagnose an isolation design

SymptomHidden weaknessCorrection
Cross-account copy failsResource/key type, vault policy, organization, role, or feature unsupportedTrace current feature matrix and both KMS/policy sides
Copies exist but attacker deletes themShared federated admin or removable governance lockSeparate routine authority; evaluate compliance lock/approval controls
Compliance vault rejects new jobsLifecycle outside min/max boundsValidate plan rules during grace before immutable lock
Immutable backup cannot restoreKMS, role, metadata, quota, network, or application dependency missingExercise full clean-room recovery and preserve emergency access
Malware scan is clean but service is unsafeUnsupported content, malicious config/identity, or integrity defectCombine full scan, hunt, rebuild, rotate, reconcile, and business test
Latest restore repeats ransomwareInfection preceded detection and replication preserved itTimeline analysis and earlier candidate testing
Destination account is “isolated”Same identity, pipeline, org admins, and keys control bothRedesign independent authority and recovery path
Restore testing says successOnly resource creation was validatedAdd application, data, security, business, and cleanup validation

Knowledge check

  1. Does cross-Region copy protect against source-account compromise?

Not by itself; the copy can remain under the same compromised account authority.

  1. What changes in a cross-account copy?

A destination-owned recovery point is created and encrypted according to destination vault/resource key behavior.

  1. How do governance and compliance Vault Lock differ?

Authorized users can remove governance lock; compliance lock becomes unchangeable after its grace period.

  1. Why can compliance lock be dangerous when misconfigured?

Retention cannot be shortened later, jobs can fail against bounds, and data/cost can persist for years or indefinitely.

  1. Is a logically air-gapped vault physically offline?

No. It provides documented logical isolation and immutable/access controls as an online AWS service.

  1. What does multi-party approval protect?

The approval path for restore access; it does not prove backup cleanliness or recovery correctness.

  1. Why is “no malware found” insufficient?

Scan scope and detection are limited, while malicious configuration, identity, hidden content, and corruption can remain.

  1. What proves recoverability?

Independent access, decryption, clean-room restore, integrity/application/business validation, measured objectives, and cleanup.

Lesson acceptance

You may continue when your submission contains:

  • explicit protection properties and ten-scenario threat model;
  • current resource/Region feature compatibility evidence;
  • cross-account/cross-Region ownership, policy, organization, and copy-lag design;
  • complete source/destination KMS and emergency decrypt path;
  • safe Vault Lock mode, grace, retention, legal, privacy, and cost decisions;
  • justified logically air-gapped, RAM, and multi-party approval design;
  • contamination timeline, scan limits, and clean-point selection;
  • isolated clean-room and controlled promotion architecture;
  • scheduled full recovery with data/application/security/business proof;
  • monitoring, quota, legal-hold, cleanup, and deletion governance; and
  • recurring exercises whose findings are owned and retested.

Official sources

Advertisement