AWS 128: Amazon Aurora architecture
Why this lesson matters
Understand Aurora's shared cluster storage, writer and reader instances, endpoints, availability, backup, and compatibility boundaries.
Aurora is not simply “faster RDS.” It changes the relationship between database compute, distributed cluster storage, replicas, endpoints, failover, backup, and I/O pricing. The learner must understand that system before choosing it over standard RDS engines.
What you will be able to do
By the end, you can:
- explain amazon aurora architecture in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Understand Aurora's shared cluster storage, writer and reader instances, endpoints, availability, backup, and compatibility boundaries. |
| Scope and boundary | An Aurora cluster separates compute instances from replicated cluster storage across AZs. Cluster, reader, instance, custom, and global endpoints serve different routing purposes. |
| Evidence of success | The design proves writer routing, read scaling, replica lag, failover priority, backup and restore, engine compatibility, and client retry behavior. |
| Cost model | Instance or capacity hours, storage, I/O under the selected configuration, backup, cross-Region replication, data transfer, monitoring, and features can charge. |
| Safe rejection rule | Avoid assuming complete upstream-engine compatibility or treating a reader endpoint as a guarantee of perfectly even application connections. |
How the request flows
+--------------------------+
| Application connection |
+--------------------------+
|
v
+-----------------------------+
| Writer or reader endpoint |
+-----------------------------+
|
v
+---------------------------+
| Aurora compute instance |
+---------------------------+
|
v
+-------------------------------------+
| Shared replicated cluster storage |
+-------------------------------------+
For Amazon Aurora, the important boundary is this: An Aurora cluster separates compute instances from replicated cluster storage across AZs. Cluster, reader, instance, custom, and global endpoints serve different routing purposes. The design proves writer routing, read scaling, replica lag, failover priority, backup and restore, engine compatibility, and client retry behavior. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use Aurora when MySQL or PostgreSQL compatibility, managed cluster storage, replicas, and Aurora-specific features meet the workload. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid assuming complete upstream-engine compatibility or treating a reader endpoint as a guarantee of perfectly even application connections. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
From traditional database servers to Aurora
A traditional database server combines an engine process, memory/cache, transaction log, and attached storage. High availability commonly requires another server with replicated data and a failover mechanism. Read scaling adds replicas with their own replication paths. Operators must reason about server and storage recovery together.
Aurora separates MySQL- or PostgreSQL-compatible compute instances from a shared distributed cluster volume. Compute instances submit redo/storage changes to the Aurora storage subsystem; they do not each own an independent full database volume. The cluster volume spans three Availability Zones in one Region. AWS describes each 10-GiB storage segment as six-way replicated across those AZs and self-healing. A failed writer can therefore be replaced by promoting an existing Aurora Replica attached to the same cluster storage.
Compatibility is not identity. Aurora MySQL and Aurora PostgreSQL support substantial compatibility with selected upstream versions, protocols, drivers, and tools, but Aurora-specific storage, privileges, extensions, parameters, replication, version lifecycle, and feature support differ. A migration must test schema, SQL, extensions, collations, privileges, drivers, binlog/logical replication, performance, and operations.
Complete Aurora resource model
| Resource | Meaning and architect decision |
|---|---|
| DB cluster | Regional Aurora control boundary containing cluster storage, engine/version, backup, encryption, endpoints, and member metadata. |
| Cluster volume | Distributed SSD-backed storage shared by cluster instances across three AZs. It grows automatically within current limits. Choose Standard or I/O-Optimized configuration from measured economics. |
| Writer DB instance | Member currently accepting writes. A normal provisioned cluster has one writer; a reader can be promoted during failover. |
| Aurora Replica | Compute attached to the same cluster volume, serving reads and acting as a failover candidate. Application-visible replica lag and stale-read requirements still matter. |
| Cluster endpoint | DNS endpoint that follows the writer. Existing connections break during failover; clients must reconnect after endpoint/DNS change. |
| Reader endpoint | Distributes new connections among available readers. It balances connections, not SQL statements, and long-lived pools can be uneven. |
| Instance endpoint | Pins connections to one member. Useful for diagnosis or specialist routing, but it bypasses automatic role following. |
| Custom endpoint | Routes to a chosen member subset for workloads such as analytics. It does not create capacity or enforce query safety. |
| DB subnet group | Candidate subnets across AZs for placement. It does not route packets or grant access. |
| VPC security group | Stateful network control on database ENIs. Permit the engine port only from approved workload security groups or networks. |
| Cluster/DB parameter group | Versioned engine configuration at cluster or instance scope. Track dynamic versus reboot-required changes and pending-reboot. |
| Backup/snapshot | Recovery artifact that creates a new cluster when restored. It does not overwrite the source in place. |
Write, read, storage, and failover paths
application writer pool read-only application pool
| |
v v
cluster endpoint DNS reader endpoint DNS
| |
v +--> reader A (AZ-a)
writer instance (AZ-a) +--> reader B (AZ-b)
| redo/storage operations +--> reader C (AZ-c)
v
distributed cluster volume: six copies of each segment across three AZs
|
+--> continuous backup / point-in-time history
+--> snapshots, clone, export, and recovery workflows
During writer failure, Aurora detects the problem, chooses an eligible reader using failover priority and then capacity/availability considerations, promotes it, and updates endpoint DNS. Connections to the failed writer are not moved. Applications need bounded connection timeouts, DNS refresh, exponential backoff with jitter, idempotent transaction handling, and pool invalidation. A cluster without a reader must create/recover a writer and normally takes longer.
Promotion tiers express candidate preference. Spread candidates across AZs and keep at least one candidate large enough for the write workload. A tiny reader can promote successfully and then fail the application through overload.
The reader endpoint chooses an instance for a new connection. A pool that holds ten sessions for hours does not continuously rebalance as readers change. Reads that require immediate visibility after a write may need the writer endpoint or an explicit consistency pattern.
Availability, durability, backup, and disaster recovery
- Six-way regional storage replication protects the cluster volume; it does not create a second-Region copy.
- Readers in separate AZs supply compute failover. One writer-only cluster still has multi-AZ storage but weaker compute recovery.
- Automated backups support point-in-time recovery inside retention. Restore creates a new cluster that needs network, parameter, secret, validation, and cutover work.
- Manual snapshots persist until deletion and follow encryption, copy, sharing, Region, and account constraints.
- Aurora cloning uses copy-on-write behavior for fast test environments, but later changes consume storage and clones retain shared dependencies.
- Backtrack is supported only for eligible Aurora MySQL configurations. It rewinds a cluster and is not a universal backup replacement.
- Aurora Global Database handles a different cross-Region topology; AWS 129 covers asynchronous replication, switchover, failover, write forwarding, and failback.
RPO is acceptable committed-data loss; RTO is time to validated application service. Multi-AZ failover, PITR, snapshot restore, clone, backtrack, and Global Database failover address different threats and have different RPO/RTO.
Storage configurations, performance, and price
Aurora Standard charges database I/O separately from compute and storage. Aurora I/O-Optimized changes that model toward higher compute/storage pricing without per-request Aurora I/O charges for supported I/O. Compare both with measured instance time, storage, read/write I/O, backups, transfer, monitoring, and features. “Mission critical” is not a cost calculation.
Performance depends on SQL plans, indexes, locks, transaction duration, connections, buffer cache, CPU, memory, network, storage operations, replicas, parameters, and client behavior. Inspect Database Insights/engine statistics, Enhanced Monitoring, CloudWatch, logs, slow-query facilities, and query plans. Aurora does not repair missing indexes or unsafe queries.
Relevant evidence includes connections, CPU, free memory, commit/storage latency, deadlocks, replica lag, cache behavior, volume bytes/I/O, failover events, and query waits/load. Exact metric availability varies by engine/version.
Security and connection design
Keep the cluster private. Application networks route to database subnets; security groups allow only the engine port from the workload security group. Public accessibility, subnet routing, NACLs, SGs, TLS, and database authentication are separate boundaries.
Use Secrets Manager rotation or IAM database authentication only when their engine/version/client constraints fit. IAM authentication still requires database users, IAM permission, TLS, token generation, and connection management. RDS Proxy can pool connections and improve supported failover paths, but transaction/session pinning can reduce multiplexing.
At-rest encryption is chosen at cluster creation and applies through service-defined storage, backup, snapshot, and replica paths. Design KMS key policy, Region/account, rotation, grants, and deletion safety. Use TLS and the current RDS CA bundle; test certificate rotation before expiry.
AWS does not own schema roles, grants, row-level controls, stored code, query safety, data classification, or migration credentials. Apply database least privilege in addition to IAM/network controls.
Automation, maintenance, and upgrades
Define subnet groups, SGs, parameter groups, cluster, instances, monitoring roles, alarms, backup policy, and secrets through reviewed IaC. Keep schema migration in a transactional/versioned migration process. Stack creation does not prove a successful transaction or failover.
Upgrade workflow:
- Discover current engine/version support and end-of-standard-support timeline.
- Review release notes, extension/parameter compatibility, drivers, and replication consumers.
- Restore or clone production-shaped data into an isolated test environment.
- Run schema migration, correctness, performance, failover, backup/restore, and rollback tests.
- Use blue/green deployment where supported or a documented snapshot/replication cutover.
- Inspect pending maintenance and
PendingModifiedValues; deliberately choose immediate versus maintenance-window application. - Validate clients and data after change. A snapshot can restore to a new cluster but does not instantly reverse an in-place upgrade.
Worked architecture decisions
Transactional web application
Requirement: PostgreSQL-compatible transactions, regional high availability, reporting reads, 15-minute RPO, and 30-minute RTO. Use Aurora PostgreSQL only after extension/SQL compatibility testing; place writer and reader in different AZs, separate cluster/reader endpoint pools, retain/test PITR, and alarm on failover and load. Global Database is unnecessary unless a Regional requirement appears.
Exact Oracle compatibility
Aurora is not valid simply because both systems are relational. Evaluate RDS for Oracle or self-managed EC2 against edition, options, licensing, OS access, HA, and backup. Aurora becomes possible only through an approved application/schema conversion.
Unpredictable I/O-heavy SaaS database
Benchmark Standard and I/O-Optimized with representative load. Compare compute, storage, I/O, backup, and monitoring cost while inspecting query waits and cache hits. I/O-Optimized can win economically without making inefficient SQL acceptable.
Analytics readers overload production
Add sized readers and a custom endpoint for analytics. Set promotion tiers so an unusually configured analytics reader is not the preferred writer, and enforce engine workload controls. Test connection distribution because endpoints balance new sessions, not queries.
Required practical evidence and failure tests
Use supplied evidence unless an hourly Aurora lab has explicit owner approval and a cleanup timer.
- Positive transaction: connect to the writer with TLS, commit an owned test row, read it back, and identify endpoint/member without exposing credentials.
- Read path: connect repeatedly through the reader endpoint, identify serving readers, and explain connection-level distribution and stale-read limits.
- Negative path: prove an unauthorized SG or database role is denied. Preserve timeout versus authentication error to identify the failed layer.
- Failover dependency: run only an approved failover; measure disconnect-to-reconnect, identify promotion/AZ, and verify the committed row. Console status alone is insufficient.
- Recovery: restore to points before and after a marker transaction in a new isolated cluster; prove inclusion/exclusion and measured RTO.
- Performance: diagnose a supplied lock/slow-query case from waits and plan evidence, apply one reversible correction, and repeat the same test.
- Cleanup: inventory cluster/instances, snapshots, automated backups, proxies, secrets, logs, ENIs, and KMS retained state; remove only owned resources and schedule billing confirmation.
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Open RDS Databases and expand an Aurora cluster in supplied evidence.
- Compare cluster writer and reader roles, endpoints, instance AZs, failover tiers, and monitoring.
- Open Backups and maintenance and identify recovery and patch controls separately.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws rds describe-db-clusters --query 'DBClusters[?contains(Engine, `aurora`)].{Id:DBClusterIdentifier,Engine:Engine,Status:Status,Writer:Endpoint,Reader:ReaderEndpoint,Encrypted:StorageEncrypted}' --output table
aws rds describe-db-clusters --query 'DBClusters[?contains(Engine, `aurora`)].DBClusterMembers[].{Instance:DBInstanceIdentifier,Writer:IsClusterWriter,Tier:PromotionTier}' --output table
Expected interpretation
Endpoints describe routing roles, and member data describes failover candidates. Neither proves query compatibility, application retry behavior, or recovery time.
Practical work
Draw an Aurora order cluster across three AZs with one writer, two readers, cluster and reader endpoints, connection pool, backups, failover order, and one promotion test.
Diagnose this topic from its own evidence
| Symptom | Evidence first | Likely boundary | Smallest safe correction |
|---|---|---|---|
| Connection times out | DNS resolution, route, NACL, SG references, endpoint/port, ENI state | Network path | Correct only the failed route/filter; do not make the cluster public |
| Connection refused/authentication fails | engine status/log, port, TLS, user/role, secret version, pg_hba-equivalent service behavior | Listener or database identity | Test current secret/user and TLS from an approved client |
| Writes fail on reader | endpoint DNS/member role and SQL error | Wrong endpoint/role | Route writes to cluster endpoint; do not promote a reader as a shortcut |
| Reads are stale or uneven | transaction timestamp, serving instance, replica lag, pool lifetime | Consistency or connection distribution | Use writer for consistency-sensitive read or recycle/partition pools intentionally |
| Failover takes longer than expected | RDS events, promotion tier, instance capacity/AZ, DNS TTL/cache, retry logs | Candidate or client recovery | Correct candidate sizing/priority and client reconnect behavior, then repeat controlled test |
| CPU/load remains high after adding reader | writer versus reader query load, plans, locks, connections | SQL/write bottleneck | Fix query/index/transaction or route eligible reads; replicas do not scale writes |
| Restore exists but app fails | subnet/SG, parameter group, secret, endpoint, schema/data checks | Recovery integration | Rebuild missing configuration and validate before cutover |
| KMS-related start/read failure | cluster key ARN/state, key policy/grants, CloudTrail | Encryption dependency | Restore least-privilege key access/state through approved KMS procedure |
Preserve the first database error, RDS event timeline, client timestamps, DNS answers, and relevant metrics before changing anything. Change one layer and repeat the same transaction.
Cost and cleanup
Instance or capacity hours, storage, I/O under the selected configuration, backup, cross-Region replication, data transfer, monitoring, and features can charge.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Understand Aurora's shared cluster storage, writer and reader instances, endpoints, availability, backup, and compatibility boundaries.
- Which scope or ownership boundary must be proved first?
Expected direction: An Aurora cluster separates compute instances from replicated cluster storage across AZs. Cluster, reader, instance, custom, and global endpoints serve different routing purposes.
- What evidence is strong enough to accept the result?
Expected direction: The design proves writer routing, read scaling, replica lag, failover priority, backup and restore, engine compatibility, and client retry behavior.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid assuming complete upstream-engine compatibility or treating a reader endpoint as a guarantee of perfectly even application connections.
- Which cost dimensions and retained resources need an owner?
Expected direction: Instance or capacity hours, storage, I/O under the selected configuration, backup, cross-Region replication, data transfer, monitoring, and features can charge.
Lesson acceptance
- Explain compute/storage separation, six-way segment replication across three AZs, and why a reader still matters for compute failover.
- Correctly use cluster, reader, instance, and custom endpoints and describe connection - not query - distribution.
- Produce compatibility evidence instead of claiming Aurora is identical to upstream MySQL/PostgreSQL.
- Show private network, SG, TLS, database identity, secret/IAM authentication, KMS, and least-privilege boundaries.
- Defend Standard versus I/O-Optimized with workload measurements and full price dimensions.
- Present positive transaction, negative authorization/network, failover, stale-read, and restore evidence with measured RPO/RTO.
- Provide IaC/change, monitoring/alarm, upgrade/rollback, retained-snapshot, and cleanup ownership.
A submission fails for public database exposure, routine root/master-user use by the application, untested restore, unmeasured failover, hidden credentials, missing KMS dependency, or cleanup claimed without final RDS/snapshot/secret/log inventory.