AWS 201: Recover the SAA capstone from one compute, network, data, and identity failure
Why this lesson matters
Recover one controlled compute, network, data, and identity fault in the P09 stack without hiding the original evidence, then remove every capstone resource. This turns four abstract reliability claims into observed failure and recovery timelines.
What you will be able to do
By the end, you can:
- prove how an Auto Scaling group replaces a terminated instance while preserving desired capacity;
- isolate an unhealthy-target symptom to the application security group's ALB-source ingress rule;
- recover a deleted S3 object by removing its current delete marker without deleting a data version;
- prove that an IAM explicit deny overrides an allow, then remove only the injected policy;
- calculate detection, diagnosis, repair and total recovery time from UTC evidence;
- restore a known baseline between experiments and prove complete P09 cleanup.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This lesson has an optional live path. Check current prices, obtain the account owner's approval, set a hard timer, use course tags, and complete the stated cleanup. The evidence path is a complete alternative.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Recover one controlled compute, network, data, and identity fault without hiding the original evidence, then remove every capstone resource. |
| Scope and boundary | For SAA capstone recovery, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup. |
| Evidence of success | Success requires matching Console, command-line, behavior, monitoring, and owner evidence for SAA capstone recovery. One green status is not enough. |
| Cost model | Capstone compute, load balancing, public IPv4, NAT, databases, logs, and backups can cost continuously. Keep the timer active until deletion and delayed billing review. |
| Safe rejection rule | Avoid broad permissions, random retries, rebuilding before preserving evidence, or leaving paid state after the recovery session. |
How the request flows
+-----------------------------+
| Controlled capstone fault |
+-----------------------------+
|
v
+-----------------------------------+
| Layered evidence and hypothesis |
+-----------------------------------+
|
v
+---------------------------------+
| One repair and changed retest |
+---------------------------------+
|
v
+------------------------------------------+
| Recovered service and complete cleanup |
+------------------------------------------+
The learner operates only the P09 stack created in AWS200. Each experiment has a hypothesis, exact target, restore trap, expected signals and independent recovery test. Random changes, broad permissions, deleting/recreating the stack, or running multiple faults together invalidate the evidence.
Incident method for every card
- Prove steady state: HTTP response, two healthy AZ-separated targets, current object value, role policies and no stack drift from an earlier card.
- Record UTC start, fault owner, blast radius, abort condition and restore command.
- Apply one exact mutation and preserve its CloudTrail request ID/event.
- Observe user symptom, ALB/ASG or S3/IAM evidence and timestamps without guessing.
- State one falsifiable hypothesis and the evidence that supports/rejects it.
- Apply the smallest restoration; do not hide evidence by rebuilding.
- Repeat the same transaction and prove the baseline independently.
- Record detection time, diagnosis time, repair time, total recovery time and preventive improvement.
Download the exact P09 fault cards. They are part of this lesson's required artifact, not optional examples.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use a bounded failure to prove that the architecture and runbook recover within the stated objective. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid broad permissions, random retries, rebuilding before preserving evidence, or leaving paid state after the recovery session. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Open CloudFormation > Stacks > nw-p09-capstone > Outputs and copy the ALB URL, target-group ARN, ASG name, bucket name and instance-role name. Confirm the account and
ap-south-1first. - Open EC2 > Target Groups > Targets. Before each card, record both instance IDs, Availability Zones, health states and reason fields.
- Open EC2 > Auto Scaling Groups > Activity for compute replacement evidence, and Security Groups > nw-p09-app-sg > Inbound rules for the exact ALB security-group reference used by the network card.
- Open the P09 S3 bucket with Show versions enabled. For the data card, distinguish an object version from a delete marker; never permanently delete a data version.
- Open IAM > Roles > the stack output role > Permissions. Confirm the normal bucket-scoped policy, then observe the separately named
InjectedDenypolicy during the identity card. - Use CloudTrail > Event history to correlate the terminate, revoke/authorize, object delete and role-policy API events. Record UTC event time, principal, resource and request ID; redact account identifiers in submitted evidence.
- After every restore, revisit target health, the ALB response, S3 version state and role policies. A Console status alone is not an end-to-end recovery test.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
stack_name="nw-p09-capstone"
aws cloudformation describe-stacks --stack-name "$stack_name" \
--query 'Stacks[0].{Status:StackStatus,Outputs:Outputs}' --output json
aws cloudformation describe-stack-events --stack-name "$stack_name" --max-items 30 --output table
asg_name="$(aws cloudformation describe-stacks --stack-name "$stack_name" --query 'Stacks[0].Outputs[?OutputKey==`AutoScalingGroupName`].OutputValue' --output text)"
target_group="$(aws cloudformation describe-stacks --stack-name "$stack_name" --query 'Stacks[0].Outputs[?OutputKey==`TargetGroupArn`].OutputValue' --output text)"
aws autoscaling describe-scaling-activities --auto-scaling-group-name "$asg_name" --max-items 20 --output table
aws elbv2 describe-target-health --target-group-arn "$target_group" --output table
aws cloudwatch describe-alarms --alarm-name-prefix nw-p09 --output table
Expected interpretation
StackStatus proves orchestration state, not application health. ASG activities explain replacement ownership; target-health reason codes explain registration and health-check transitions. Only the repeated ALB transaction proves user-visible recovery. The bucket version list and role-policy readback separately prove data and identity restoration.
Practical work
Run the four supplied P09 fault cards in sequence. Preserve each symptom, locate the failed layer, make one repair, repeat the transaction, restore baseline and finally delete the stack with an empty dependency inventory. If live deployment was not approved, use instructor-provided command/event/response captures and write the same evidence timeline; do not claim live execution.
Capstone fault cards and final cleanup
Use only an instructor-approved or learner-owned capstone. Inject one fault at a time and restore the approved state before moving on.
| Fault | First evidence | Safe repair target | Passing retest |
|---|---|---|---|
| compute process stopped or unhealthy target | target reason, instance status, service log | restore the service or replace the unhealthy unit through its owner | new request succeeds through the load balancer |
| route or security rule removed | route table, SG rule, Flow Log or Reachability Analyzer path | restore the exact reviewed route or rule | same network test reaches the intended endpoint only |
| data item deleted or database unavailable | backup timestamp, recovery point, database event | restore to a new target and validate data before cutover | known record and checksum match |
| role or resource permission denied | caller, CloudTrail event, policy evaluation, explicit deny | restore the smallest missing allow or remove the injected deny | same principal performs only the intended action |
After the four retests, delete the CloudFormation stack. Reconcile ENIs, load balancers, target groups, public IPv4 addresses, NAT gateways, volumes, snapshots, log groups, backups, secrets, roles, policies, DNS records, and retained buckets by exact name and tag. Never delete an item that is not in the signed capstone ledger.
Diagnose this topic from its own evidence
- Compute: correlate terminated instance, ASG activity, target deregistration/registration, healthy count and continuous client results. Replacement alone does not prove availability.
- Network: distinguish route, NACL, SG, listener, target port and process. The P09 card changes only the app SG's source-SG ingress and restores that exact rule.
- Data: distinguish current object absence from permanent deletion. Inspect versions/delete marker and remove only the marker; preserve version IDs.
- Identity: evaluate explicit deny before allows, verify exact role/policy/resource, allow propagation time and remove only
InjectedDeny. - Cleanup: a deleted stack is insufficient if the versioned bucket blocked deletion or retained resources remain. Query every ledger type and delayed billing.
Cost and cleanup
Capstone compute, load balancing, public IPv4, NAT, databases, logs, and backups can cost continuously. Keep the timer active until deletion and delayed billing review.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Recover one controlled compute, network, data, and identity fault without hiding the original evidence, then remove every capstone resource.
- Which scope or ownership boundary must be proved first?
Expected direction: For SAA capstone recovery, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup.
- What evidence is strong enough to accept the result?
Expected direction: Success requires matching Console, command-line, behavior, monitoring, and owner evidence for SAA capstone recovery. One green status is not enough.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid broad permissions, random retries, rebuilding before preserving evidence, or leaving paid state after the recovery session.
- Which cost dimensions and retained resources need an owner?
Expected direction: Capstone compute, load balancing, public IPv4, NAT, databases, logs, and backups can cost continuously. Keep the timer active until deletion and delayed billing review.
Lesson acceptance
- Four before/failure/after evidence timelines identify exact actor, resource, signal and elapsed recovery.
- Each card changes one bounded variable and restores the previous baseline before the next.
- Compute recovery preserves desired capacity; network recovery restores SG reference rather than broad CIDR.
- Data recovery retains object versions; identity recovery proves explicit-deny precedence without widening access.
- Stack deletion completes after version cleanup, and exact P09 inventory shows zero owned residuals.