Lesson 129 · AWS Learning Path

AWS 129: Aurora Serverless v2 and Aurora Global Database

· Published · 14 min read

Labelled process diagram for AWS 129: Variable or global demand to Serverless v2 capacity or global endpoint to Aurora regional cluster to Scaling, lag, and recovery evidence, with decision, proof and rejection evidence.

Why this lesson matters

Use current Aurora Serverless v2 scaling and Global Database deliberately, without teaching retired Serverless v1 creation.

This lesson joins two separate Aurora capabilities that solve different problems. Aurora Serverless v2 changes how regional database compute capacity is provisioned and billed. Aurora Global Database changes the cross-Region replication and recovery topology. Neither capability automatically provides the other.

What you will be able to do

By the end, you can:

  • explain aurora serverless v2 and aurora global database in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeUse current Aurora Serverless v2 scaling and Global Database deliberately, without teaching retired Serverless v1 creation.
Scope and boundaryServerless v2 scales compatible Aurora compute within configured ACU bounds. Global Database uses one primary Region and read-only secondary clusters with managed cross-Region replication.
Evidence of successEvidence shows min and max ACUs, actual capacity, replica roles, replication lag, write-forwarding choice, switchover or failover steps, and application endpoints.
Cost modelACU time, I/O, storage, backup, cross-Region transfer, secondary compute, monitoring, and Global Database features can charge.
Safe rejection ruleDo not create or recommend Aurora Serverless v1 for a new learner, and do not promise synchronous cross-Region writes or zero data loss.

How the request flows

+-----------------------------+
|  Variable or global demand  |
+-----------------------------+
              |
              v
+---------------------------------------------+
|  Serverless v2 capacity or global endpoint  |
+---------------------------------------------+
                      |
                      v
+---------------------------+
|  Aurora regional cluster  |
+---------------------------+
             |
             v
+---------------------------------------+
|  Scaling, lag, and recovery evidence  |
+---------------------------------------+

For Aurora Serverless v2 and Aurora Global Database, the important boundary is this: Serverless v2 scales compatible Aurora compute within configured ACU bounds. Global Database uses one primary Region and read-only secondary clusters with managed cross-Region replication. Evidence shows min and max ACUs, actual capacity, replica roles, replication lag, write-forwarding choice, switchover or failover steps, and application endpoints. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.

Architecture decision table

SituationDirectionReason
Requirement matchesUse Serverless v2 for variable compatible workloads that benefit from fine-grained managed capacity, and Global Database for low-latency global reads or cross-Region recovery.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchDo not create or recommend Aurora Serverless v1 for a new learner, and do not promise synchronous cross-Region writes or zero data loss.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

Part 1: Aurora Serverless v2 mental model

Aurora Serverless v2 is an Aurora DB instance class (db.serverless) whose CPU and memory capacity can scale in Aurora Capacity Units (ACUs) within a cluster-configured minimum and maximum. Storage remains Aurora cluster storage and continues to incur storage, backup, I/O/configuration, and related charges. “Serverless” means AWS manages compute-capacity adjustment; it does not mean there are no instances, connections, engine versions, subnets, security groups, parameters, maintenance, backups, or cost while active.

Current supported configurations allow half-ACU increments up to 256 ACUs, subject to engine and platform version. Supported versions can use a minimum of 0 ACUs for automatic pause; older versions/configurations may require at least 0.5 ACU. Always query supported engine versions and current capacity documentation in the target Region.

incoming connections and SQL
          |
          v
Aurora endpoint -> db.serverless instance
                       |
                       | scales inside MinCapacity..MaxCapacity
                       v
             CPU + memory + connection/cache capacity
                       |
                       v
             shared Aurora cluster volume

Capacity is not an application request count. Aurora observes resource pressure and adjusts compute. Scaling is not instantaneous, and the configured maximum is a hard workload ceiling. Too-low minimums can produce cold cache, slower scale response, memory pressure, parameter incompatibility, or connection limits. Too-high minimums preserve performance but reduce savings.

Capacity range, parameters, and monitoring

One ACU represents an approximate bundle of memory with corresponding CPU/network resources; use AWS's current definition rather than translating ACUs directly into a fixed vCPU count. Aurora automatically adjusts certain parameters based on maximum capacity. After changing maximum ACUs, inspect pending-reboot and engine-specific parameter behavior; some changes require instance reboot to use the new derived values fully.

Monitor at least:

  • ServerlessDatabaseCapacity for current ACUs;
  • ACUUtilization relative to the configured maximum;
  • CPU, free memory, connections, query load/waits, latency, deadlocks, and storage behavior;
  • scale events, failover events, and application p95/p99 latency;
  • ServerlessV2Usage and billing dimensions where available;
  • missing metrics/log gaps while an instance is paused.

If capacity remains at maximum and latency rises, raising the maximum may help only when compute/memory is the bottleneck. Query plans, locks, connection storms, hot rows, storage, external calls, or downstream limits require different corrections.

Promotion tiers influence serverless reader scaling. Readers in promotion tiers 0 or 1 track the writer's capacity behavior more closely so they are prepared for failover. Higher-tier readers can scale more independently and, where supported/configured, pause independently. Design failover readiness and savings together.

Automatic pause to zero

On supported engine versions, setting cluster minimum capacity to 0 enables auto-pause capability. Configure an idle timeout. An instance pauses only after user-initiated connections are absent and no blocking condition applies. Open sessions - even idle application pools - can prevent pause. Logical/binlog replication, zero-ETL integrations, maintenance, Babelfish activity, provisioned members in a hybrid cluster, or promotion-tier relationships can change eligibility.

A connection attempt can wake a paused instance even when credentials are invalid. Writer endpoint access resumes the writer and related low-promotion-tier readers; reader endpoint access may resume writer first and the chosen reader. Resume adds connection latency, so zero-ACU auto-pause is for workloads whose SLO tolerates it. It is not the default architecture for latency-critical production traffic.

While paused, instance capacity charges stop, but storage, backup, snapshots, monitoring destinations, I/O-related storage configuration, global resources, and other cluster charges can continue. Most instance metrics/log production also pauses, so “no metrics” can be expected paused behavior rather than an outage.

Serverless v2 architecture choices

WorkloadDirectionWhy
Variable production traffic with continuous availabilitySet a tested nonzero minimum and maximum with headroomPreserves warm capacity/cache while allowing growth
Development database idle overnightSupported 0-ACU auto-pause with a tolerable wake delaySaves compute while retaining cluster data/configuration
Predictable constant loadCompare provisioned Reserved DB Instances with Serverless v2 ACU costAutomatic scaling may add no value and can cost more
Sudden unbounded connection burstAdd application backpressure and RDS Proxy where suitableScaling compute does not instantly create unlimited engine sessions
Reader analytics with independent variabilityHigher promotion-tier serverless reader/custom endpointAllows different scaling but weakens immediate failover readiness if undersized/paused

Part 2: Aurora Global Database mental model

An Aurora global database contains one primary Region with a read/write cluster and one or more secondary Regions with read-only clusters. Dedicated Aurora storage-based cross-Region replication is asynchronous. Secondary clusters can provide low-latency regional reads and are recovery candidates, but they do not imply zero RPO.

writers -> primary Region cluster (writer + readers)
                     |
                     | asynchronous storage replication
          +----------+-----------+
          v                      v
secondary Region A          secondary Region B
read-only cluster           read-only cluster
regional readers            regional readers

Each regional cluster needs its own instances/capacity, subnets, SGs, KMS key, parameters, secrets/authentication reachability, logs, alarms, quotas, and application path. Global replication does not copy all surrounding infrastructure.

Endpoints and application routing

Applications need separate write and read routing decisions. Global writer endpoint capabilities can follow the primary Region in supported configurations, while regional cluster/reader endpoints remain scoped to a cluster. Do not hard-code an instance endpoint for a failover-capable application.

Global readers can observe lag. For read-your-write business operations, route the read to the primary/writer path or use an explicit consistency design. Cache and DNS behavior, connection pools, transactions, and ORM failover handling must be tested.

Global write forwarding can allow SQL issued to a secondary cluster to be forwarded to the primary under supported Aurora/version modes. It does not make the secondary a synchronous local writer. It adds cross-Region latency, consistency-mode choices, transaction/statement restrictions, and failure dependencies. Use it only after SQL compatibility and latency tests.

Planned switchover versus unplanned failover

OperationUseData expectation and risk
Managed switchoverPlanned maintenance, regional rotation, or migration while primary is healthyWaits for synchronization and changes roles while preserving topology; target zero data loss under documented prerequisites
Managed failoverPrimary Region is unavailable or cannot complete a normal switchoverPromotes a selected secondary; asynchronous lag can become data loss and split-brain/conflict controls matter
Detach/promote workflowLegacy/manual recovery or topology change where managed operation is unsuitableMore manual endpoint, topology, replication, and rebuild responsibility

Before either operation, inspect global status, target cluster health/capacity, replication lag, pending maintenance, KMS/secrets, quotas, application readiness, and current writes. Freeze or fence writers when the runbook requires it. During failover, record the last confirmed commit and accepted RPO. After promotion, prove writes/reads, reconcile missing/duplicate business operations, redirect clients, and re-establish secondary topology before declaring recovery complete.

Failback is not “switch DNS back.” The old primary can contain divergent data. Determine authoritative history, rebuild/rejoin according to supported workflow, wait for replication, test, and perform a planned switchover.

Cost and capacity model

regional serverless cost = ACU-seconds/hours per active instance
                         + regional cluster storage configuration
                         + backup/snapshot storage
                         + monitoring/logging/proxy/secrets

global additions = compute in every secondary Region
                 + replicated storage in every Region
                 + cross-Region replicated write I/O/data transfer dimensions
                 + Global Database / global write-forwarding related usage
                 + backup, logs, alarms, KMS, and network paths per Region

Auto-pause one instance does not stop the global database bill. Secondary instances maintained for recovery are intentionally paid capacity. Compare the required RPO/RTO and global-read latency against snapshot copy/restore, regional read replicas, DMS, or application-level data architectures.

Worked examples

Spiky SaaS API

Use Serverless v2 with a tested nonzero minimum, maximum based on load test plus headroom, RDS Proxy for supported connection pooling, and alarms on capacity saturation and query load. Auto-pause is rejected because wake latency violates the API SLO.

Developer preview environment

Use a supported engine with minimum 0 ACUs and a ten-minute idle timeout. Ensure CI closes sessions, accept first-connection wake delay, retain a daily deletion schedule for abandoned clusters, and remember storage/backup cost continues.

Global read-heavy catalog

Use Global Database with primary writes and regional reader endpoints. The app routes consistency-sensitive post-write reads to primary and stale-tolerant catalog browsing locally. Monitor lag and test a managed switchover quarterly.

Regional disaster during accepted writes

Compare last confirmed primary commit with secondary replication position and business event log, invoke managed failover under incident authority, fence old-primary writes, validate the promoted Region, reconcile ambiguous operations idempotently, then rebuild the old Region as secondary. Do not promise zero loss from asynchronous replication.

Required tests

  1. Drive a controlled load from minimum toward maximum ACUs; correlate ACUs, query latency, waits, connections, and scale time.
  2. Saturate the configured maximum in a bounded test and prove the alarm/runbook distinguishes capacity ceiling from slow SQL.
  3. For a 0-ACU lab, close every connection, prove pause, measure resume latency from writer and reader endpoints, and explain remaining charges/metric gaps.
  4. Introduce an open idle session that blocks pause; diagnose it without deleting the cluster.
  5. Measure Global Database replication lag with marker writes and local secondary reads; quantify observed RPO exposure.
  6. Conduct a planned managed switchover and an unplanned-failure tabletop or approved lab, recording endpoint, role, transaction, lag, and elapsed-time evidence.
  7. Test one unsupported/unsafe global write-forwarding SQL pattern from current engine documentation and preserve the expected rejection.
  8. Clean up in dependency order: global topology/secondary clusters, regional instances/clusters, snapshots/backups, proxy, secrets, logs/alarms, network resources, and retained KMS keys - only when owned and approved.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open RDS Databases and inspect a supplied Aurora Serverless v2 cluster's Capacity range.
  2. Open an Aurora global database capture and identify primary and secondary Regions, clusters, and lag.
  3. Open Modify only to locate min and max ACU controls, then cancel without saving.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws rds describe-db-clusters --query 'DBClusters[].{Id:DBClusterIdentifier,Mode:EngineMode,MinACU:ServerlessV2ScalingConfiguration.MinCapacity,MaxACU:ServerlessV2ScalingConfiguration.MaxCapacity,Global:GlobalWriteForwardingStatus}' --output table
aws rds describe-global-clusters --query 'GlobalClusters[].{Id:GlobalClusterIdentifier,Engine:Engine,Members:GlobalClusterMembers[].DBClusterArn}' --output json

Expected interpretation

ServerlessV2ScalingConfiguration proves configured bounds, not instant scaling or application performance. Global membership proves topology, not zero RPO or tested promotion.

Practical work

Write two ADRs: one for a spiky development workload using Serverless v2 ACU bounds, and one for a global read-heavy service with a primary and two secondary Regions. Include failover and cost tests.

Diagnose this topic from its own evidence

SymptomEvidence firstLikely causeSafe correction
Capacity stays at maximumACU/CPU/memory, DB load/waits, SQL plans, max settingReal compute ceiling or inefficient SQL/locksCorrect proven SQL/lock issue or raise tested maximum with cost approval
Instance never auto-pausesmin ACU, engine version, sessions, replication/integration, member types and tiersIneligible configuration or active connectionClose/fix owner connection or change design; do not force-stop production
First connection times out after pauseresume events, client timeout/retry, endpoint and TLS logsClient cannot tolerate wake delayIncrease timeout/retry or use nonzero minimum for the SLO
Secondary read lacks recent writemarker timestamp and replication-lag metricsExpected asynchronous lagRoute consistency-sensitive read to primary or redesign consistency contract
Forwarded write failsengine/version, forwarding mode, SQL/transaction restriction, primary healthUnsupported statement/mode or cross-Region dependencyUse supported transaction path; do not grant broader IAM
Switchover is blockedglobal/cluster state, lag, target health/capacity, pending operationsTopology not synchronized/healthyCorrect target and wait within change window; do not convert to forced failover casually
Failover succeeds but app still uses old Regionendpoint/DNS, connection pools, secrets, SG, deployment configApplication routing/stateFence old writer, refresh clients/config, validate transactions

Always preserve cluster events, global topology, lag, client timestamps, endpoint answers, and last confirmed transaction before changing roles.

Cost and cleanup

ACU time, I/O, storage, backup, cross-Region transfer, secondary compute, monitoring, and Global Database features can charge.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Use current Aurora Serverless v2 scaling and Global Database deliberately, without teaching retired Serverless v1 creation.

  1. Which scope or ownership boundary must be proved first?

Expected direction: Serverless v2 scales compatible Aurora compute within configured ACU bounds. Global Database uses one primary Region and read-only secondary clusters with managed cross-Region replication.

  1. What evidence is strong enough to accept the result?

Expected direction: Evidence shows min and max ACUs, actual capacity, replica roles, replication lag, write-forwarding choice, switchover or failover steps, and application endpoints.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Do not create or recommend Aurora Serverless v1 for a new learner, and do not promise synchronous cross-Region writes or zero data loss.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: ACU time, I/O, storage, backup, cross-Region transfer, secondary compute, monitoring, and Global Database features can charge.

Lesson acceptance

  • Explain ACUs, min/max range, scale ceiling, promotion-tier interaction, auto-pause eligibility, wake behavior, and continuing non-compute charges.
  • Demonstrate that Serverless v2 still has Aurora instances, networking, connections, parameters, maintenance, backup, and operational ownership.
  • Diagram primary and secondary regional clusters and identify asynchronous replication and independent regional dependencies.
  • Separate regional endpoints, global endpoint behavior, replication, global write forwarding, switchover, failover, and failback.
  • Provide measured scale, pause/resume, lag, planned switch, failure, recovery, and application reconnection evidence.
  • Calculate every regional instance/ACU, storage, I/O, backup, replication/transfer, KMS, logging, proxy, and retained-resource dimension.

The lesson fails if it teaches retired Serverless v1 creation, promises instantaneous scaling or zero-loss asynchronous recovery, ignores open connections that prevent pause, or treats endpoint role change as proof the application recovered.

Official sources

Advertisement