AWS 137: Deploy, protect, monitor, back up, and restore a learning database
Why this lesson matters
Deploy an optional private learning database, prove protection and monitoring, restore it, and remove every paid dependency in one controlled session.
This is the integration lab for the relational-database lessons. You will treat recovery as a measured application outcome - not a checkbox - and prove the complete path from a private client through DNS, security groups and database authentication to an encrypted database, its backups, a separate restored instance and verified rows.
What you will be able to do
By the end, you can:
- explain deploy, protect, monitor, back up, and restore a learning database in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This lesson has an optional live path. Check current prices, obtain the account owner's approval, set a hard timer, use course tags, and complete the stated cleanup. The evidence path is a complete alternative.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Deploy an optional private learning database, prove protection and monitoring, restore it, and remove every paid dependency in one controlled session. |
| Scope and boundary | The live path uses one approved Region, two private subnets in different AZs, a DB subnet group, restricted SG, encrypted RDS instance, managed master password, backups, a snapshot or PITR restore, and exact cleanup. |
| Evidence of success | Passing evidence includes before inventory, timer, creation events, private endpoint, encryption, backup window, monitoring, test data, restore validation, and empty after inventory. |
| Cost model | RDS instance and restored-instance time, storage, backup beyond allowance, snapshots, Secrets Manager, data transfer, and public IPv4 dependencies can charge. |
| Safe rejection rule | Do not make the database public, paste a password into commands, attach broad administrator policy, or leave the source or restore running overnight. |
How the request flows
+------------------------------+
| Private client requirement |
+------------------------------+
|
v
+---------------------------------------------+
| Encrypted RDS source and automated backup |
+---------------------------------------------+
|
v
+-------------------------+
| New restored database |
+-------------------------+
|
v
+----------------------------------------------+
| Integrity result and empty after inventory |
+----------------------------------------------+
For Protected RDS learning database lab, the important boundary is this: The live path uses one approved Region, two private subnets in different AZs, a DB subnet group, restricted SG, encrypted RDS instance, managed master password, backups, a snapshot or PITR restore, and exact cleanup. Passing evidence includes before inventory, timer, creation events, private endpoint, encryption, backup window, monitoring, test data, restore validation, and empty after inventory. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use the live track only when the owner has approved price, IAM actions, exact identifiers, and same-session cleanup. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Do not make the database public, paste a password into commands, attach broad administrator policy, or leave the source or restore running overnight. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Open RDS Databases and choose Create database only after owner approval and current-price review.
- Select Standard create, a supported MySQL-compatible engine, a small approved class, managed master credentials, private access, encryption, one-day backup retention, and course tags.
- After availability, inspect Connectivity, Configuration, Monitoring, Logs and events, Maintenance, and Backups before using Actions to create or plan the restore.
- Delete the restored database first, then the source without a final snapshot only after the evidence package is complete; remove the managed secret and owned network dependencies.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws rds describe-db-instances --query 'DBInstances[?starts_with(DBInstanceIdentifier, `nw-p06-`)].{Id:DBInstanceIdentifier,Status:DBInstanceStatus,Public:PubliclyAccessible,Encrypted:StorageEncrypted,Retention:BackupRetentionPeriod}' --output table
aws rds describe-events --source-type db-instance --duration 1440 --output table
aws rds describe-db-snapshots --query 'DBSnapshots[?starts_with(DBSnapshotIdentifier, `nw-p06-`)].{Id:DBSnapshotIdentifier,State:Status}' --output table
Expected interpretation
The expected live result is a private encrypted source, a separately named restored database that passes an integrity query, and then an empty scoped inventory. The required no-create track uses supplied evidence for the same decisions.
Practical work
Complete p06-rds-lab-evidence.md using either the full no-create evidence pack or the optional live run. Record every command, timestamp, result, restore measurement, and deletion state.
Optional live lab runbook
The no-create evidence track is a complete way to pass. Use the live track only after an owner approves the current RDS price in the chosen Region. Start a two-hour timer before creating anything.
Phase 0: define success before creating anything
Create p06-rds-lab-evidence.md and write these values first:
| Field | Required decision |
|---|---|
| Owner and deletion deadline | Named person and UTC time |
| Region and Availability Zones | One Region; two distinct AZs for the subnet group |
| Engine/version/class | Current MySQL-compatible RDS option supported in the Region; never guess a version or class |
| Client | An already-owned EC2 instance or other private client whose security group may be referenced |
| RPO | Maximum acceptable data loss; for this lab, the snapshot timestamp is the explicit recovery point |
| RTO target | Your estimate from restore request to validated SQL result |
| Test record | Fake value such as recovery-marker-001; no personal or production data |
| Budget boundary | Estimated instance, storage, backup, Secrets Manager and transfer charge |
Record the starting inventory. If an identifier already exists, choose a new course-owned identifier; never adopt or delete the existing resource.
set -euo pipefail
export AWS_DEFAULT_REGION="ap-south-1"
export DB_ID="nw-p06-learning-db"
export SNAPSHOT_ID="nw-p06-before-restore"
export RESTORE_ID="nw-p06-restored"
aws sts get-caller-identity --query Arn --output text
aws rds describe-db-instances \
--query 'DBInstances[?starts_with(DBInstanceIdentifier, `nw-p06-`)].DBInstanceIdentifier' \
--output text
aws rds describe-db-snapshots --snapshot-type manual \
--query 'DBSnapshots[?starts_with(DBSnapshotIdentifier, `nw-p06-`)].DBSnapshotIdentifier' \
--output text
Console build values
- In VPC > Security groups, create
nw-p06-db-sgin an owned lab VPC. Add MySQL TCP 3306 only from an owned client security group. Do not use0.0.0.0/0. - In RDS > Subnet groups, create
nw-p06-db-subnetsfrom two private subnets in different AZs. - In RDS > Databases > Create database, use Standard create, MySQL, a current small approved class, 20 GiB gp3, managed master credentials, the lab VPC and subnet group, no public access,
nw-p06-db-sg, storage encryption, one-day automated backup retention, deletion protection off, and tagProject=NitWings-P06. - Wait for
Available. Record the endpoint without publishing it. Inspect the managed secret, logs, metrics, events, earliest and latest restore time, and the final settings. - Add only fake learning data. Take a manual snapshot named
nw-p06-before-restore, then restore tonw-p06-restored. Validate row count and one checksum or known query result. - Delete
nw-p06-restored, thennw-p06-learning-db. Delete the manual snapshot and managed secret after their dependencies are gone. Remove only the owned subnet group and security group.
Phase 1: prove protection and reachability
The database endpoint is DNS, not a permanent IP address. The request path is:
private client process
-> VPC DNS resolves the RDS endpoint
-> client's outbound rule and route/local VPC path
-> DB security group's inbound reference to the client SG on TCP 3306
-> TLS negotiation
-> database user authentication and grants
-> MySQL engine and encrypted RDS storage
These controls solve different problems. A security group does not authenticate a database user; a password does not open a network path; storage encryption does not force TLS in transit. Capture evidence for each layer separately. The DB subnet group must span at least two AZs, even when this cost-bounded lab uses a Single-AZ DB instance. Single-AZ is not a production availability design.
From the approved private client, retrieve the master credentials without printing them, connect with TLS using the current RDS CA bundle, and create only synthetic data:
CREATE DATABASE recovery_lab;
CREATE TABLE recovery_lab.markers (
marker_id BIGINT PRIMARY KEY AUTO_INCREMENT,
marker_text VARCHAR(100) NOT NULL,
created_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP
);
INSERT INTO recovery_lab.markers(marker_text) VALUES ('recovery-marker-001');
SELECT COUNT(*) AS row_count,
MIN(marker_text) AS first_marker,
MAX(marker_text) AS last_marker
FROM recovery_lab.markers;
Do not put the password on a command line, in shell history or in the evidence file. Record whether the connection used TLS and prove PubliclyAccessible=false, StorageEncrypted=true, the KMS key, backup retention and subnet/security-group attachments with describe-db-instances.
Phase 2: create the recovery point and measure restore
Flush/commit the test transaction and note its UTC timestamp. Create the manual snapshot, wait for available, and record its source, engine, encrypted state and snapshot creation time. Snapshot restore creates a new DB instance; it does not overwrite or rewind the source. The restored instance needs its own instance class, subnet group, security groups, parameter choices and operational owner.
restore_requested_at="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
aws rds create-db-snapshot \
--db-instance-identifier "$DB_ID" \
--db-snapshot-identifier "$SNAPSHOT_ID"
aws rds wait db-snapshot-available --db-snapshot-identifier "$SNAPSHOT_ID"
aws rds describe-db-snapshots --db-snapshot-identifier "$SNAPSHOT_ID" \
--query 'DBSnapshots[0].{Status:Status,Encrypted:Encrypted,Created:SnapshotCreateTime,Source:DBInstanceIdentifier}' \
--output table
Restore through the console using the approved worksheet values, name it $RESTORE_ID, keep it private and attach only the approved DB security group. A snapshot preserves data and many settings, but you must inspect every restored setting rather than assuming that networking, parameter groups, monitoring or deletion protection match the source. Wait for availability, connect from the private client, rerun the integrity query and record restore_validated_at plus the measured elapsed time.
For an advanced repeat, use automated point-in-time recovery to a time after the committed insert and before a later synthetic change. PITR also creates a new instance. Its achievable recovery point is bounded by EarliestRestorableTime and LatestRestorableTime; it is not an in-place undo button.
Phase 3: observe behavior, not just green status
Correlate at least four evidence sources:
- RDS events for creation, backup, snapshot, restore and deletion transitions;
- CloudWatch metrics such as
CPUUtilization,DatabaseConnections,FreeStorageSpace,FreeableMemory, read/write latency and IOPS; - database SQL evidence proving the expected schema and row marker;
- configuration evidence for private access, encryption, retention, subnet group, security groups and certificate authority.
A DB instance reporting available proves that the RDS control plane considers it available. It does not prove DNS resolution from the client, permitted network traffic, valid credentials, TLS, application schema, correct recovery point or acceptable query behavior.
Useful scoped variables and checks:
export DB_ID="nw-p06-learning-db"
export RESTORE_ID="nw-p06-restored"
aws rds describe-db-instances --db-instance-identifier "$DB_ID" --query 'DBInstances[0].{Status:DBInstanceStatus,Public:PubliclyAccessible,Encrypted:StorageEncrypted,Retention:BackupRetentionPeriod}' --output table
aws rds wait db-instance-available --db-instance-identifier "$DB_ID"
aws rds create-db-snapshot --db-instance-identifier "$DB_ID" --db-snapshot-identifier nw-p06-before-restore
aws rds wait db-snapshot-available --db-snapshot-identifier nw-p06-before-restore
Creation values that contain subnet IDs, security-group IDs, engine versions, classes, or prices must come from the approved worksheet. Do not paste passwords into CLI arguments or shell history. A successful restore API call is not the pass gate. The learner must connect through the approved private client and validate the data.
Diagnose this topic from its own evidence
Work from the first broken layer instead of changing several controls at once:
| Symptom | First evidence | Likely causes | Safe next action |
|---|---|---|---|
Create remains creating or fails | RDS events and DB instance status reason | unsupported class/version/storage combination, KMS or IAM failure, invalid subnet group | preserve the event text; correct the specific input rather than repeatedly recreating |
| Client cannot resolve endpoint | getent hosts ENDPOINT, VPC DNS attributes | copied endpoint error, DNS disabled, wrong resolver/context | compare endpoint from API and test DNS before changing security groups |
| Connection times out | SG references, subnet/route/NACL path, Reachability Analyzer where applicable | no permitted TCP path, wrong port/VPC, NACL response-path issue | trace client-to-DB path; never add 0.0.0.0/0 as a diagnostic shortcut |
| Immediate access denied | engine error and username, secret metadata, DB grants | wrong user/secret, rotated secret, host/grant or IAM-auth mismatch | verify identity and authentication method without exposing the secret |
| TLS verification fails | client error, RDS CA configured on instance, client CA bundle date | missing/old trust bundle or wrong hostname | update from the official trust-store source and retain hostname verification |
| Restore is available but data is absent | snapshot/PITR timestamp and SQL transaction time | wrong recovery point, uncommitted write, wrong database/schema/endpoint | compare the UTC timeline and query the restored endpoint explicitly |
| Cleanup reports dependency | DB/snapshot/secret/subnet-group/SG inventory | instance still deleting, snapshot retained, SG attached elsewhere | wait for terminal deletion and remove only tagged, recorded lab resources |
For one intentional negative test, remove the client-SG reference, prove the timeout, restore the rule and prove recovery. Do not weaken the rule to the internet. Record the expected failure, actual evidence and rollback.
Cost and cleanup
RDS instance and restored-instance time, storage, backup beyond allowance, snapshots, Secrets Manager, data transfer, and public IPv4 dependencies can charge.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Deploy an optional private learning database, prove protection and monitoring, restore it, and remove every paid dependency in one controlled session.
- Which scope or ownership boundary must be proved first?
Expected direction: The live path uses one approved Region, two private subnets in different AZs, a DB subnet group, restricted SG, encrypted RDS instance, managed master password, backups, a snapshot or PITR restore, and exact cleanup.
- What evidence is strong enough to accept the result?
Expected direction: Passing evidence includes before inventory, timer, creation events, private endpoint, encryption, backup window, monitoring, test data, restore validation, and empty after inventory.
- Which tempting design or shortcut must be rejected?
Expected direction: Do not make the database public, paste a password into commands, attach broad administrator policy, or leave the source or restore running overnight.
- Which cost dimensions and retained resources need an owner?
Expected direction: RDS instance and restored-instance time, storage, backup beyond allowance, snapshots, Secrets Manager, data transfer, and public IPv4 dependencies can charge.
Lesson acceptance
Pass only when the evidence package proves all of the following:
- the owner, Region, identifiers, timer, RPO, RTO target and price estimate existed before creation;
- the source and restore were private and encrypted, and ingress referenced only the approved client security group;
- credentials were managed and never exposed in commands or submitted evidence;
- the source SQL marker was committed, the snapshot or PITR point was identified in UTC, and the restored instance returned the expected integrity result;
- events and CloudWatch metrics were interpreted, including why
availablealone is insufficient; - one controlled network failure was diagnosed and reversed without opening public access;
- measured recovery time and the gap between target and result are explained;
- final scoped inventories show no
nw-p06-DB instance or manual snapshot, and the owned secret, subnet group and SG were removed or an approved retained dependency is documented.
The no-create alternative must reach the same conclusions from the supplied event, configuration, metric and SQL evidence. A diagram or screenshot without causal explanation does not pass.