Lesson 128 · AWS Learning Path

AWS 128: Amazon Aurora architecture

· Published · 14 min read

Labelled process diagram for AWS 128: Application connection to Writer or reader endpoint to Aurora compute instance to Shared replicated cluster storage, with decision, proof and rejection evidence.

Why this lesson matters

Understand Aurora's shared cluster storage, writer and reader instances, endpoints, availability, backup, and compatibility boundaries.

Aurora is not simply “faster RDS.” It changes the relationship between database compute, distributed cluster storage, replicas, endpoints, failover, backup, and I/O pricing. The learner must understand that system before choosing it over standard RDS engines.

What you will be able to do

By the end, you can:

  • explain amazon aurora architecture in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeUnderstand Aurora's shared cluster storage, writer and reader instances, endpoints, availability, backup, and compatibility boundaries.
Scope and boundaryAn Aurora cluster separates compute instances from replicated cluster storage across AZs. Cluster, reader, instance, custom, and global endpoints serve different routing purposes.
Evidence of successThe design proves writer routing, read scaling, replica lag, failover priority, backup and restore, engine compatibility, and client retry behavior.
Cost modelInstance or capacity hours, storage, I/O under the selected configuration, backup, cross-Region replication, data transfer, monitoring, and features can charge.
Safe rejection ruleAvoid assuming complete upstream-engine compatibility or treating a reader endpoint as a guarantee of perfectly even application connections.

How the request flows

+--------------------------+
|  Application connection  |
+--------------------------+
             |
             v
+-----------------------------+
|  Writer or reader endpoint  |
+-----------------------------+
              |
              v
+---------------------------+
|  Aurora compute instance  |
+---------------------------+
             |
             v
+-------------------------------------+
|  Shared replicated cluster storage  |
+-------------------------------------+

For Amazon Aurora, the important boundary is this: An Aurora cluster separates compute instances from replicated cluster storage across AZs. Cluster, reader, instance, custom, and global endpoints serve different routing purposes. The design proves writer routing, read scaling, replica lag, failover priority, backup and restore, engine compatibility, and client retry behavior. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.

Architecture decision table

SituationDirectionReason
Requirement matchesUse Aurora when MySQL or PostgreSQL compatibility, managed cluster storage, replicas, and Aurora-specific features meet the workload.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid assuming complete upstream-engine compatibility or treating a reader endpoint as a guarantee of perfectly even application connections.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

From traditional database servers to Aurora

A traditional database server combines an engine process, memory/cache, transaction log, and attached storage. High availability commonly requires another server with replicated data and a failover mechanism. Read scaling adds replicas with their own replication paths. Operators must reason about server and storage recovery together.

Aurora separates MySQL- or PostgreSQL-compatible compute instances from a shared distributed cluster volume. Compute instances submit redo/storage changes to the Aurora storage subsystem; they do not each own an independent full database volume. The cluster volume spans three Availability Zones in one Region. AWS describes each 10-GiB storage segment as six-way replicated across those AZs and self-healing. A failed writer can therefore be replaced by promoting an existing Aurora Replica attached to the same cluster storage.

Compatibility is not identity. Aurora MySQL and Aurora PostgreSQL support substantial compatibility with selected upstream versions, protocols, drivers, and tools, but Aurora-specific storage, privileges, extensions, parameters, replication, version lifecycle, and feature support differ. A migration must test schema, SQL, extensions, collations, privileges, drivers, binlog/logical replication, performance, and operations.

Complete Aurora resource model

ResourceMeaning and architect decision
DB clusterRegional Aurora control boundary containing cluster storage, engine/version, backup, encryption, endpoints, and member metadata.
Cluster volumeDistributed SSD-backed storage shared by cluster instances across three AZs. It grows automatically within current limits. Choose Standard or I/O-Optimized configuration from measured economics.
Writer DB instanceMember currently accepting writes. A normal provisioned cluster has one writer; a reader can be promoted during failover.
Aurora ReplicaCompute attached to the same cluster volume, serving reads and acting as a failover candidate. Application-visible replica lag and stale-read requirements still matter.
Cluster endpointDNS endpoint that follows the writer. Existing connections break during failover; clients must reconnect after endpoint/DNS change.
Reader endpointDistributes new connections among available readers. It balances connections, not SQL statements, and long-lived pools can be uneven.
Instance endpointPins connections to one member. Useful for diagnosis or specialist routing, but it bypasses automatic role following.
Custom endpointRoutes to a chosen member subset for workloads such as analytics. It does not create capacity or enforce query safety.
DB subnet groupCandidate subnets across AZs for placement. It does not route packets or grant access.
VPC security groupStateful network control on database ENIs. Permit the engine port only from approved workload security groups or networks.
Cluster/DB parameter groupVersioned engine configuration at cluster or instance scope. Track dynamic versus reboot-required changes and pending-reboot.
Backup/snapshotRecovery artifact that creates a new cluster when restored. It does not overwrite the source in place.

Write, read, storage, and failover paths

application writer pool                       read-only application pool
        |                                                |
        v                                                v
cluster endpoint DNS                              reader endpoint DNS
        |                                                |
        v                                                +--> reader A (AZ-a)
writer instance (AZ-a)                                 +--> reader B (AZ-b)
        | redo/storage operations                        +--> reader C (AZ-c)
        v
distributed cluster volume: six copies of each segment across three AZs
        |
        +--> continuous backup / point-in-time history
        +--> snapshots, clone, export, and recovery workflows

During writer failure, Aurora detects the problem, chooses an eligible reader using failover priority and then capacity/availability considerations, promotes it, and updates endpoint DNS. Connections to the failed writer are not moved. Applications need bounded connection timeouts, DNS refresh, exponential backoff with jitter, idempotent transaction handling, and pool invalidation. A cluster without a reader must create/recover a writer and normally takes longer.

Promotion tiers express candidate preference. Spread candidates across AZs and keep at least one candidate large enough for the write workload. A tiny reader can promote successfully and then fail the application through overload.

The reader endpoint chooses an instance for a new connection. A pool that holds ten sessions for hours does not continuously rebalance as readers change. Reads that require immediate visibility after a write may need the writer endpoint or an explicit consistency pattern.

Availability, durability, backup, and disaster recovery

  • Six-way regional storage replication protects the cluster volume; it does not create a second-Region copy.
  • Readers in separate AZs supply compute failover. One writer-only cluster still has multi-AZ storage but weaker compute recovery.
  • Automated backups support point-in-time recovery inside retention. Restore creates a new cluster that needs network, parameter, secret, validation, and cutover work.
  • Manual snapshots persist until deletion and follow encryption, copy, sharing, Region, and account constraints.
  • Aurora cloning uses copy-on-write behavior for fast test environments, but later changes consume storage and clones retain shared dependencies.
  • Backtrack is supported only for eligible Aurora MySQL configurations. It rewinds a cluster and is not a universal backup replacement.
  • Aurora Global Database handles a different cross-Region topology; AWS 129 covers asynchronous replication, switchover, failover, write forwarding, and failback.

RPO is acceptable committed-data loss; RTO is time to validated application service. Multi-AZ failover, PITR, snapshot restore, clone, backtrack, and Global Database failover address different threats and have different RPO/RTO.

Storage configurations, performance, and price

Aurora Standard charges database I/O separately from compute and storage. Aurora I/O-Optimized changes that model toward higher compute/storage pricing without per-request Aurora I/O charges for supported I/O. Compare both with measured instance time, storage, read/write I/O, backups, transfer, monitoring, and features. “Mission critical” is not a cost calculation.

Performance depends on SQL plans, indexes, locks, transaction duration, connections, buffer cache, CPU, memory, network, storage operations, replicas, parameters, and client behavior. Inspect Database Insights/engine statistics, Enhanced Monitoring, CloudWatch, logs, slow-query facilities, and query plans. Aurora does not repair missing indexes or unsafe queries.

Relevant evidence includes connections, CPU, free memory, commit/storage latency, deadlocks, replica lag, cache behavior, volume bytes/I/O, failover events, and query waits/load. Exact metric availability varies by engine/version.

Security and connection design

Keep the cluster private. Application networks route to database subnets; security groups allow only the engine port from the workload security group. Public accessibility, subnet routing, NACLs, SGs, TLS, and database authentication are separate boundaries.

Use Secrets Manager rotation or IAM database authentication only when their engine/version/client constraints fit. IAM authentication still requires database users, IAM permission, TLS, token generation, and connection management. RDS Proxy can pool connections and improve supported failover paths, but transaction/session pinning can reduce multiplexing.

At-rest encryption is chosen at cluster creation and applies through service-defined storage, backup, snapshot, and replica paths. Design KMS key policy, Region/account, rotation, grants, and deletion safety. Use TLS and the current RDS CA bundle; test certificate rotation before expiry.

AWS does not own schema roles, grants, row-level controls, stored code, query safety, data classification, or migration credentials. Apply database least privilege in addition to IAM/network controls.

Automation, maintenance, and upgrades

Define subnet groups, SGs, parameter groups, cluster, instances, monitoring roles, alarms, backup policy, and secrets through reviewed IaC. Keep schema migration in a transactional/versioned migration process. Stack creation does not prove a successful transaction or failover.

Upgrade workflow:

  1. Discover current engine/version support and end-of-standard-support timeline.
  2. Review release notes, extension/parameter compatibility, drivers, and replication consumers.
  3. Restore or clone production-shaped data into an isolated test environment.
  4. Run schema migration, correctness, performance, failover, backup/restore, and rollback tests.
  5. Use blue/green deployment where supported or a documented snapshot/replication cutover.
  6. Inspect pending maintenance and PendingModifiedValues; deliberately choose immediate versus maintenance-window application.
  7. Validate clients and data after change. A snapshot can restore to a new cluster but does not instantly reverse an in-place upgrade.

Worked architecture decisions

Transactional web application

Requirement: PostgreSQL-compatible transactions, regional high availability, reporting reads, 15-minute RPO, and 30-minute RTO. Use Aurora PostgreSQL only after extension/SQL compatibility testing; place writer and reader in different AZs, separate cluster/reader endpoint pools, retain/test PITR, and alarm on failover and load. Global Database is unnecessary unless a Regional requirement appears.

Exact Oracle compatibility

Aurora is not valid simply because both systems are relational. Evaluate RDS for Oracle or self-managed EC2 against edition, options, licensing, OS access, HA, and backup. Aurora becomes possible only through an approved application/schema conversion.

Unpredictable I/O-heavy SaaS database

Benchmark Standard and I/O-Optimized with representative load. Compare compute, storage, I/O, backup, and monitoring cost while inspecting query waits and cache hits. I/O-Optimized can win economically without making inefficient SQL acceptable.

Analytics readers overload production

Add sized readers and a custom endpoint for analytics. Set promotion tiers so an unusually configured analytics reader is not the preferred writer, and enforce engine workload controls. Test connection distribution because endpoints balance new sessions, not queries.

Required practical evidence and failure tests

Use supplied evidence unless an hourly Aurora lab has explicit owner approval and a cleanup timer.

  1. Positive transaction: connect to the writer with TLS, commit an owned test row, read it back, and identify endpoint/member without exposing credentials.
  2. Read path: connect repeatedly through the reader endpoint, identify serving readers, and explain connection-level distribution and stale-read limits.
  3. Negative path: prove an unauthorized SG or database role is denied. Preserve timeout versus authentication error to identify the failed layer.
  4. Failover dependency: run only an approved failover; measure disconnect-to-reconnect, identify promotion/AZ, and verify the committed row. Console status alone is insufficient.
  5. Recovery: restore to points before and after a marker transaction in a new isolated cluster; prove inclusion/exclusion and measured RTO.
  6. Performance: diagnose a supplied lock/slow-query case from waits and plan evidence, apply one reversible correction, and repeat the same test.
  7. Cleanup: inventory cluster/instances, snapshots, automated backups, proxies, secrets, logs, ENIs, and KMS retained state; remove only owned resources and schedule billing confirmation.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open RDS Databases and expand an Aurora cluster in supplied evidence.
  2. Compare cluster writer and reader roles, endpoints, instance AZs, failover tiers, and monitoring.
  3. Open Backups and maintenance and identify recovery and patch controls separately.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws rds describe-db-clusters --query 'DBClusters[?contains(Engine, `aurora`)].{Id:DBClusterIdentifier,Engine:Engine,Status:Status,Writer:Endpoint,Reader:ReaderEndpoint,Encrypted:StorageEncrypted}' --output table
aws rds describe-db-clusters --query 'DBClusters[?contains(Engine, `aurora`)].DBClusterMembers[].{Instance:DBInstanceIdentifier,Writer:IsClusterWriter,Tier:PromotionTier}' --output table

Expected interpretation

Endpoints describe routing roles, and member data describes failover candidates. Neither proves query compatibility, application retry behavior, or recovery time.

Practical work

Draw an Aurora order cluster across three AZs with one writer, two readers, cluster and reader endpoints, connection pool, backups, failover order, and one promotion test.

Diagnose this topic from its own evidence

SymptomEvidence firstLikely boundarySmallest safe correction
Connection times outDNS resolution, route, NACL, SG references, endpoint/port, ENI stateNetwork pathCorrect only the failed route/filter; do not make the cluster public
Connection refused/authentication failsengine status/log, port, TLS, user/role, secret version, pg_hba-equivalent service behaviorListener or database identityTest current secret/user and TLS from an approved client
Writes fail on readerendpoint DNS/member role and SQL errorWrong endpoint/roleRoute writes to cluster endpoint; do not promote a reader as a shortcut
Reads are stale or uneventransaction timestamp, serving instance, replica lag, pool lifetimeConsistency or connection distributionUse writer for consistency-sensitive read or recycle/partition pools intentionally
Failover takes longer than expectedRDS events, promotion tier, instance capacity/AZ, DNS TTL/cache, retry logsCandidate or client recoveryCorrect candidate sizing/priority and client reconnect behavior, then repeat controlled test
CPU/load remains high after adding readerwriter versus reader query load, plans, locks, connectionsSQL/write bottleneckFix query/index/transaction or route eligible reads; replicas do not scale writes
Restore exists but app failssubnet/SG, parameter group, secret, endpoint, schema/data checksRecovery integrationRebuild missing configuration and validate before cutover
KMS-related start/read failurecluster key ARN/state, key policy/grants, CloudTrailEncryption dependencyRestore least-privilege key access/state through approved KMS procedure

Preserve the first database error, RDS event timeline, client timestamps, DNS answers, and relevant metrics before changing anything. Change one layer and repeat the same transaction.

Cost and cleanup

Instance or capacity hours, storage, I/O under the selected configuration, backup, cross-Region replication, data transfer, monitoring, and features can charge.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Understand Aurora's shared cluster storage, writer and reader instances, endpoints, availability, backup, and compatibility boundaries.

  1. Which scope or ownership boundary must be proved first?

Expected direction: An Aurora cluster separates compute instances from replicated cluster storage across AZs. Cluster, reader, instance, custom, and global endpoints serve different routing purposes.

  1. What evidence is strong enough to accept the result?

Expected direction: The design proves writer routing, read scaling, replica lag, failover priority, backup and restore, engine compatibility, and client retry behavior.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid assuming complete upstream-engine compatibility or treating a reader endpoint as a guarantee of perfectly even application connections.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Instance or capacity hours, storage, I/O under the selected configuration, backup, cross-Region replication, data transfer, monitoring, and features can charge.

Lesson acceptance

  • Explain compute/storage separation, six-way segment replication across three AZs, and why a reader still matters for compute failover.
  • Correctly use cluster, reader, instance, and custom endpoints and describe connection - not query - distribution.
  • Produce compatibility evidence instead of claiming Aurora is identical to upstream MySQL/PostgreSQL.
  • Show private network, SG, TLS, database identity, secret/IAM authentication, KMS, and least-privilege boundaries.
  • Defend Standard versus I/O-Optimized with workload measurements and full price dimensions.
  • Present positive transaction, negative authorization/network, failover, stale-read, and restore evidence with measured RPO/RTO.
  • Provide IaC/change, monitoring/alarm, upgrade/rollback, retained-snapshot, and cleanup ownership.

A submission fails for public database exposure, routine root/master-user use by the application, untested restore, unmeasured failover, hidden credentials, missing KMS dependency, or cleanup claimed without final RDS/snapshot/secret/log inventory.

Official sources

Advertisement