Lesson 113 · AWS Learning Path

AWS 113: S3 replication and Multi-Region Access Points

· Published · 15 min read

One Region contains three isolated Availability Zones while smaller edge nodes serve nearby users

The real problem

A team recognizes the name S3 replication and Multi-Region Access Points but has not connected the feature to a real requirement, identity boundary, network or data path, failure mode, price dimension, and cleanup owner. A plausible configuration could still fail the workload.

Final outcome

The learner will produce a requirement-led artifact for S3 replication and Multi-Region Access Points, inspect the matching AWS control plane in the Management Console, run a matching CloudShell or AWS CLI query, interpret the output, diagnose one failure, defend one architecture choice, and prove cleanup or approved retained state.

The practical outcome is not a command transcript. It must show what was expected, what happened, what the result proves, what it does not prove, and which evidence would change the decision.

Learning objectives

By the end of this lesson, the learner can:

  • explain replication prerequisites;
  • explain new versus existing objects;
  • explain srr and crr;
  • explain delete behavior;
  • explain replication is asynchronous;
  • connect control-plane state to the real data, network, identity, or application behavior;
  • identify cost and cleanup ownership before any optional mutation;
  • troubleshoot from evidence without opening broad access or adding broad permissions.

Relationship model

Requirement
   |
   v
Identity and policy -> AWS configuration -> network or data path -> workload behavior
        |                    |                      |                    |
        +--------------------+----------------------+--------------------+
                                      |
                                      v
                         monitoring, cost, recovery, cleanup

Use this model to separate an AWS object that exists from a result that actually works. Every arrow is a verification boundary.

Prerequisites, permissions, Region, and safety

  • Learning baseline: This sequence assumes practical Linux knowledge but no prior cloud-computing or AWS knowledge. Cloud, networking, security, data, automation, and architecture concepts must come from completed earlier lessons. If a prerequisite checkpoint is incomplete, return to its linked lesson before continuing.
  • Confirm a non-root caller with aws sts get-caller-identity and keep the account number private.
  • Use ap-south-1 unless this lesson explicitly names a second Region.
  • Confirm the intended profile and Region with aws configure list before interpreting an empty result.
  • Use read-only List, Get, and Describe permissions for the named services. Design exercises run locally and require no resource-creation permission.
  • This is a no-create lesson. Console and CLI work is read-only, and every design artifact is created locally.
  • Never publish account IDs, public addresses, ARNs containing private account data, session IDs, presigned URLs, object data, credentials, or KMS material.
  • Do not use root, world-open SSH or RDP, disabled TLS verification, unowned resources, or irreversible retention controls in a training exercise.

Core model

ConceptWhat the learner must understand
Replication prerequisitesS3 live replication requires versioning on source and destination, a replication configuration, service permissions, destination authorization, and encryption permissions where applicable.
New versus existing objectsA live rule normally applies to eligible objects written after the rule is active. S3 Batch Replication is the controlled option for existing, failed, or previously replicated objects.
SRR and CRRSame-Region Replication supports same-Region compliance, log aggregation, or account separation. Cross-Region Replication supports geographic copies and Regional recovery designs.
Delete behaviorDelete-marker replication is configurable and has limitations. Deleting a specific source version does not become a replicated destructive delete of that destination version.
Replication is asynchronousReplication status and metrics must be monitored. Replication Time Control offers a predictable target for eligible new objects at additional cost; it is not an application failover mechanism.
Multi-Region Access PointsMRAP provides a global endpoint and routes requests to associated buckets. It does not copy data by itself; replication must keep required objects present in the routed buckets.

How it works

A complete recovery design includes write direction, conflict ownership, version selection, replica encryption, delete semantics, lag monitoring, failover control, failback, DNS or endpoint use, application consistency, account isolation, and tested restore.

Read the result in layers:

  1. Scope: account, Region, VPC, bucket, AZ, endpoint, principal, object version, or resource ARN.
  2. Control plane: the requested configuration exists and reached an expected state.
  3. Behavior: the request, connection, health check, replication, restore, or application result meets the requirement.
  4. Operations: monitoring, failure owner, cost, retention, rollback, and cleanup are known.

Control-plane success is necessary but not sufficient. A resource can be available while policy, routing, DNS, health, data, or application behavior remains wrong.

Architecture decision table

RequirementPreferred directionWhy
Second Region copy for recoveryCRR to a versioned destinationThe rule creates asynchronous copies across Regions.
Central immutable copy in another accountCross-account replication with destination ownership and restrictive policiesAccount separation reduces one administrative blast radius.
Copy objects that existed before the ruleS3 Batch ReplicationLive replication does not automatically backfill all prior objects.
Global application endpointMRAP plus replication and failover planRouting and data synchronization are separate requirements.

Professional questions normally contain several valid services. State the requirement that selects one option, why the nearest alternative fails it, and what changed requirement would reverse the choice.

Replication data path and prerequisites

source PUT creates source version
        |
        | rule filter + replication eligibility
        v
S3 assumes replication IAM role
        |
        +-- read source version, tags, retention, encryption context
        +-- decrypt with source KMS key when required
        |
        v
write replica to destination bucket/version
        |
        +-- destination bucket policy / owner override
        +-- encrypt with destination KMS key when configured
        +-- destination storage class / Object Lock behavior
        v
replication status + metrics + failure/threshold events

Both source and destination general purpose buckets must have versioning enabled for live replication. S3 needs an assumable IAM role with narrowly scoped source-read and destination-replication permissions; cross-account designs also need destination resource policy and ownership decisions. SSE-KMS/DSSE-KMS requires explicit replication configuration plus KMS decrypt/encrypt/data-key permissions and key policies in the relevant accounts/Regions. A healthy bucket and a valid-looking rule do not prove an object is eligible or replicated.

Rule anatomy and object eligibility

A production rule has ID, priority, enabled state, filter, delete-marker choice, destination bucket, destination storage class where required, IAM role, and optional ownership, encryption, metrics, RTC, and replica-modification controls. Multiple rules can target different datasets/destinations, so priority and overlapping filters must be reviewed.

Live rules handle eligible new versions created after the rule is active. Existing versions, versions missed before activation, failed objects, and already replicated replicas require S3 Batch Replication according to the use case. Do not “touch” millions of objects merely to force a new version without an approved data and cost impact.

Common eligibility boundaries include:

  • rule prefix/tag filter and rule status;
  • source version creation time relative to rule activation;
  • encryption type and whether KMS-encrypted replication was enabled;
  • object ownership, source/destination permissions, and destination versioning;
  • destination Region/class/key compatibility;
  • Object Lock configuration and required permissions;
  • objects created by another replication rule, which ordinary live replication does not simply chain as expected;
  • lifecycle actions and delete-marker origin.

Inspect each source version's replication status (PENDING, COMPLETED, FAILED, or replica-related status where returned), not only whether a key with the same name exists at destination. Compare version metadata, checksum/size, encryption, retention, tags, and destination ownership as required.

What replication does - and does not - copy

Replication preserves important object-version information and can copy metadata, tags, retention, and supported encryption state when configured. It can select another destination storage class and destination owner. It does not copy bucket configuration such as bucket policy, lifecycle, notifications, BPA, logging, or versioning setup; manage each bucket explicitly through IaC.

Deleting a specific source version is not replicated as a permanent deletion of the corresponding replica. Delete-marker replication is configurable for eligible deletes, but delete markers created by lifecycle expiration have different behavior and should not be assumed to replicate. This deliberate asymmetry can protect destination history, but it also means source and destination inventories diverge by design.

Replication is not backup isolation by itself. A broad role may delete destination versions, compromised credentials can write malicious new versions, two-way replication can propagate changes, and KMS/key/account dependencies can share a blast radius. Combine account separation, least privilege, versioning, Object Lock/backup where required, monitoring, and tested restore.

SRR, CRR, Batch Replication, and RTC

CapabilitySelect whenImportant limitation/cost
Same-Region ReplicationDifferent ownership/account, log aggregation, compliance copy, or processing copy must remain in one RegionDoes not address a Regional failure; adds requests and destination storage
Cross-Region ReplicationGeographic copy, data locality, lower read latency near compute, or Regional recoveryAsynchronous; adds inter-Region replication transfer, requests, storage, KMS and monitoring costs
S3 Batch ReplicationExisting, previously failed, or already replicated objects need an on-demand copyRequires inventory/manifest and Batch Operations job permissions; completion has a separate job timeline and RTC does not apply
S3 Replication Time ControlEligible new objects need the documented 15-minute SLA and threshold visibilityExtra replication/metrics cost, quotas and exclusions; it does not make an application transaction synchronous

Current AWS documentation presents inconsistent percentages in different RTC pages while consistently describing an SLA-backed 15-minute target for eligible new objects. A production requirement must cite the current S3 RTC SLA/legal service commitment - not a memorized percentage from a lesson - and record exclusions such as request-rate or replication-transfer quota conditions.

Monitor pending operations, bytes pending, replication latency, failed operations, threshold-missed/late events, KMS throttling, destination errors, and application-level missing-object behavior. Alarm ownership and replay procedure are required.

Multi-Region Access Points: routing is not copying

An S3 Multi-Region Access Point (MRAP) supplies a global endpoint across associated Regional buckets and routes requests through the AWS network. By default routing is based on AWS routing considerations such as proximity/health, not whether the chosen bucket contains the requested key. If the routed bucket lacks an object, the request can return 404. Configure and monitor replication separately.

MRAP failover controls can shift traffic between active and passive Regions. A complete design needs:

  1. associated buckets and MRAP/access policies;
  2. two-way replication for write continuity where the application may write after failover;
  3. RTC/metrics where the RPO requires it;
  4. replica modification sync when supported metadata changes must return;
  5. explicit active/passive routing state and change authority;
  6. application endpoint/SigV4A support, retries, idempotency, and consistency handling;
  7. failover entry criteria, replication-lag check, data-loss acceptance, and stakeholder approval;
  8. failback conflict reconciliation and verification.

MRAP is not Route 53 DNS failover, not CloudFront caching, not replication, and not a database multi-writer coordinator. It solves an S3 request-routing problem.

Recovery design: calculate RPO and RTO honestly

If a source write is acknowledged at 12:00:00 and replication completes at 12:07:00, a source Region loss at 12:03 can lose that version from the destination recovery view. The observed replication lag contributes to RPO. RTC changes the commitment/visibility for eligible objects; it does not make RPO zero.

RTO includes detection, decision approval, confirmation of destination completeness, endpoint/routing change, application credential/policy/KMS readiness, cache behavior, and validation - not merely “the replica exists.” Test with a known object manifest and application read/write operations.

Worked design 1: immutable cross-account recovery vault

Replicate from production to a security-owned destination account/Region with destination ownership, tightly scoped role and key policies, Object Lock where required, separate administrators, replication metrics, and a restore role that production cannot normally assume. Test recovery into a clean location. This reduces one-account administrator blast radius but still needs KMS and Organizations dependency analysis.

Worked design 2: active/passive global object application

Use two Regional versioned buckets, two-way CRR, MRAP, replication metrics/RTC according to RPO, deterministic immutable keys, and failover controls. Before failover, inspect lag and failed operations; after failover, test reads and writes; before failback, prove reverse replication and resolve concurrent-write policy.

Worked design 3: migration of a historical bucket

Enable live replication first for new changes, generate an S3 Inventory/manifest for existing versions, run Batch Replication with completion reports, reconcile counts/bytes/checksums and failures, then perform a final delta/control check. A live rule alone leaves the historical dataset behind.

Failure diagnosis

SymptomEvidenceLikely causesSafe next action
Source status stays PENDINGversion status, age, pending bytes/latencybacklog, request/transfer quota, large object, KMS throttleinspect metrics/quotas and wait with a bounded RPO threshold
Source status becomes FAILEDfailure event/reason and exact versionrole/key/destination/versioning/ownership/class issuecorrect only named dependency; use Batch Replication for failed versions if required
Destination key exists but wrong version/datasource and destination version IDs, size/checksum/metadataold copy, noneligible version, independent write, lagcompare exact version and replication status; do not trust key existence
MRAP GET returns 404routed Region/bucket, object inventories, replicationobject absent from routed bucketrepair replication/data placement; routing cannot manufacture the object
Failover writes do not return to original Regiontwo-way rules and replica modification configurationone-way replication or unsupported change pathstop failback, reconcile, then establish/test reverse flow
Encrypted objects fail onlyactual key ARNs, key state/policies, replication configKMS opt-in/permission/context/quotamake least-privilege key/config correction and replay failed objects safely

Required verification matrix

  1. Upload a post-rule object and two versions; prove each expected replica status and destination object behavior.
  2. Show that a pre-rule object is not automatically backfilled, then design or run an approved Batch Replication job and reconcile its completion report.
  3. Produce one intentional dependency failure using a supplied broken policy/KMS scenario; diagnose from failure reason without broad permissions.
  4. Exercise a delete marker and exact-version delete; document which effect replicates and which does not.
  5. For MRAP, verify routing/data separately and conduct a controlled failover/failback tabletop or lab with explicit stop conditions.
  6. Calculate source requests, replication requests, destination storage, inter-Region transfer, KMS, RTC/metrics, inventory/Batch Operations, and retained-version costs.

AWS Management Console guided practice

Before opening a service page, write the expected account, Region, starting state, and evidence. Do not choose Create, Save, Purchase, Lock, or Delete unless the lesson explicitly authorizes the live track.

  1. Open an instructor-supplied source bucket, Management, Replication rules and inspect status, priority, filter, destination, IAM role, encryption, delete-marker setting, and metrics.
  2. Open both buckets with Show versions and compare source version IDs, destination version IDs, timestamps, encryption, and replication status.
  3. Open Multi-Region Access Points and supplied routing evidence; explain why associated buckets still need replication for consistent reads.

For each step, capture the field name and value in text. A screenshot may support the record but does not replace the explanation. Console labels can evolve, so use the service search and current documentation if a navigation label differs.

CloudShell and AWS CLI practice

CloudShell is the default browser-based command environment taught in AWS 028. AWS 029 and AWS 030 cover local CLI installation and authentication. This lesson therefore does not assume that an unconfigured local shell is ready.

Start every session with:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account portion of the ARN before sharing. Then perform the topic query:

Inspect replication configuration and per-version replication status on owned buckets.

NW_SOURCE_BUCKET="replace-with-owned-source-bucket-name"
NW_REPLICA_BUCKET="replace-with-owned-replica-bucket-name"
aws s3api get-bucket-replication --bucket "$NW_SOURCE_BUCKET"
aws s3api head-object --bucket "$NW_SOURCE_BUCKET" --key documents/policy.txt
aws s3api head-object --bucket "$NW_REPLICA_BUCKET" --key documents/policy.txt --region ap-southeast-1

Expected interpretation:

A destination object proves that version arrived. It does not prove zero lag, every historical version, correct delete behavior, application failover, or restore permission.

Replace every replace-with-... sample value before running its command, and use only an explicitly owned resource. Explain each option first. These queries are read-only; a successful response does not authorize a later create or delete operation.

Practical work

Create p06-replication-plan.md: source nw-p06-source-{unique-suffix} in ap-south-1, destination nw-p06-replica-{unique-suffix} in ap-southeast-1, versioning enabled, one-way replication for documents/, SSE-S3, service role nw-p06-s3-replication, delete-marker replication enabled for the lab, no RTC or MRAP creation, status evidence, restore test, and version-aware cleanup. Record current cross-Region transfer and storage charges.

The evidence package must contain:

  • the problem and final requirement in the learner's own words;
  • caller type and Region with private identifiers redacted;
  • exact planned values, ownership, and cost class;
  • one Console observation and matching CLI or API evidence;
  • one behavior result or supplied data-plane record;
  • one denied, failed, or counterexample result and evidence-led diagnosis;
  • one architecture choice plus the rejected alternative;
  • cleanup proof or explicit retained-state owner, expiry, and next lesson.

Verification standard

Use expected state before observed state. Record timestamps in UTC and preserve the original failure before changing anything. A passing submission answers all four questions:

  1. What exact requirement was tested?
  2. Which evidence proves the AWS configuration?
  3. Which evidence proves the workload behavior?
  4. What remains unproven or requires later monitoring?

If AWS returns no rows, verify account, Region, permission, filters, pagination, resource type, and deletion state before concluding that nothing exists.

Common failures and troubleshooting

SymptomEvidence firstLikely boundarySmallest safe response
object appears missingcaller, Region, filters, pagination, tagsscope or read permissionalign scope before creating a duplicate
state remains pending or unavailableservice state, events, dependencies, quotasdependency or capacitycorrect the named dependency and wait with a bound
AccessDeniedprincipal, action, resource, explicit-deny contextidentity, resource, endpoint, organization, or KMS policychange only the proven policy layer
configuration exists but behavior failsroute, DNS, security, listener, health, logs, object versiondata path or applicationtest the next boundary and change one control
bill is higher than expectedhours, bytes, requests, AZs, addresses, retentioncost model or retained resourcestop optional work and reconcile the ledger
cleanup is blockeddependency inventory and owning servicedeletion order or immutable stateremove owned dependants in reviewed reverse order

Do not troubleshoot by attaching administrator access, opening administration ports to the internet, disabling encryption, retrying uncontrolled creation, deleting unknown resources, or weakening retention.

Cost, cleanup, and retained state

No AWS resource is created. Close CloudShell and remove or redact downloaded evidence.

Cleanup evidence requires terminal state and an after-inventory. Search related ENIs, public IPv4 addresses, EBS volumes and snapshots, load balancers, target groups, Auto Scaling instances, endpoints, logs, S3 versions and delete markers, backup recovery points, and global IAM roles when they apply. Billing data can lag, so schedule a later review.

Architecture and certification decisions

  • Certification coverage: SAA-C03; SOA-C03; SAP-C02; DOP-C02.
  • Exam mapping: SAA D1-D4.
  • Explain service scope, failure boundary, consistency, recovery, security, operations, and price rather than matching a keyword.
  • Treat availability and durability, encryption and authorization, routing and filtering, health and lifecycle, backup and replication, and discount and capacity as separate concepts.
  • Do not reproduce protected certification questions.

Knowledge check

  1. Does enabling CRR copy every old object automatically?

Expected direction: No. Use Batch Replication for existing objects.

  1. Does MRAP replicate data?

Expected direction: No. It routes requests; configure replication separately.

  1. Must both replication buckets use versioning?

Expected direction: Yes.

  1. Is replication synchronous?

Expected direction: No. Monitor status and lag according to the requirement.

Completion gate and assessment

AreaPointsPassing evidence
Requirement and model15Correct scope, terminology, and final outcome
Console evidence15Current path and interpreted fields
CLI or API evidence15Scoped command, expected result, and limitations
Behavior or decision exercise20Reproducible result or defensible architecture reasoning
Troubleshooting15Original symptom, hypothesis, one change, retest, rollback
Security and cost10Least privilege, data protection, current price dimensions
Cleanup and handoff10Terminal-state proof or approved retained-state record

Pass at 80 out of 100 with no critical safety failure. A missing practical artifact, unexplained output, unsafe access, destructive action outside the owned scope, unplanned billed resource, or false cleanup claim requires remediation and a changed retest.

Official sources

Advertisement