AWS 248: Complete the SOA operations capstone without a supplied runbook
The gate
This capstone removes procedural prompting. You receive business impact, scope, recovery objectives, and evidence access - not a diagnosis or runbook. You must establish incident control, form falsifiable hypotheses, request evidence, isolate interacting faults, propose and sequence controlled recovery, prove customer and security outcomes, then create the runbook that did not exist.
The default track is a complete no-create simulation. An optional live version requires an isolated instructor-owned account, approved fault manifest, stop conditions, budget alarm, exact rollback, and cleanup; never inject these faults into production or a shared learning account.
Required pack and information separation
Download the AWS248 independent capstone pack. It contains:
CANDIDATE_BRIEF.md- the only file opened at incident start;CAPSTONE_WORKBOOK.md- incident command, evidence, changes, validation, RCA, runbook, cleanup;EVIDENCE_RELEASES.md- five releases opened only after the stated gate;ASSESSOR_GUIDE.md- fault model, expected reasoning, scoring, and stop rules; keep sealed until assessment ends.
Information separation matters. Reading the assessor guide or all releases before writing the initial evidence plan converts the capstone into a reading exercise and invalidates the attempt.
The beginner-to-operator transition
Earlier lessons told you where to look. Here you choose where to look and explain why. A Linux-experienced learner should use familiar reasoning:
- process state is not dependency health;
- a listening socket is not a successful transaction;
- a route is not a permitted packet;
- an API
200is not customer recovery; - a backup file is not a tested restore;
- a configuration change is not correct merely because it deployed;
- correlation becomes causation only after competing hypotheses are tested.
Use AWS service evidence to extend that model across account, Region, identity, control plane, network, compute, application, data, and observability layers.
Incident lifecycle and roles
detect -> declare -> scope/stabilize -> hypothesize -> gather evidence
-> contain -> recover -> verify -> reconcile owner source
-> learn/automate -> retain evidence -> clean up
Assign one incident commander, one operations lead, one communications lead, and one scribe - even if one learner wears several hats. The commander owns priority and approvals, not every command. Record UTC timestamps and separate facts, hypotheses, decisions, actions, and outcomes.
Classify:
- event - observed occurrence/change;
- incident - disruption requiring response;
- problem - underlying cause or recurring condition;
- known error/workaround - understood problem with temporary handling;
- change - approved mutation with owner, risk, rollback, and evidence.
Do not start with root-cause perfection while customers are harmed. Stabilization can precede full RCA, but each containment action still needs a hypothesis, scope, approval, abort condition, and rollback.
First 15 minutes
Before Release 1, the candidate must:
- declare severity from business impact rather than alarm color;
- record start/detection/declare times and uncertainty;
- confirm authorized account, Region, application, and mutation boundary;
- name roles, update interval, and escalation path;
- state RTO, RPO, SLO/error-budget impact, and last known good;
- freeze nonessential changes without destroying evidence;
- draw the expected request path;
- write at least five falsifiable hypotheses across different layers;
- rank the next read-only evidence by discriminatory value;
- define conditions for containment, rollback, and emergency escalation.
“Restart everything” and “open access temporarily” are not evidence plans.
Evidence discipline
For every observation record source, query/filter, account, Region, resource, UTC window, value, interpretation, and limitations. Screenshots without scope are weak. Preserve relevant CloudTrail events, CloudFormation/Auto Scaling activity, target health, metrics/logs, IAM denial context, network/config state, and owner-source versions before mutation.
Use an evidence matrix:
| Hypothesis | Prediction if true | Evidence supporting | Evidence contradicting | Next discriminator | Status |
|---|
Actively seek contradictory evidence. Green EC2 status checks can reject host failure while leaving capacity, application startup, IAM, load balancer, and database hypotheses open.
Change protocol
Every mutation entry must include:
- change ID, requester, approver, operator, UTC;
- exact target and current state;
- hypothesis and expected observation;
- command/API or owner-source diff;
- blast radius, dependencies, and cost;
- rollback or forward-fix steps;
- abort/stop condition;
- result and independent verification.
Prefer changing source of truth and deploying a reviewed version. Emergency console changes must be reconciled immediately afterward. Do not mutate an unknown resource because its name resembles the capstone.
Recovery sequencing
Containment limits harm; recovery restores service; remediation removes causes; prevention reduces recurrence. They are not synonyms.
Sequence by dependency and risk:
- preserve evidence and stop harmful automation/change;
- restore safe capacity/path/dependency prerequisites;
- replace or reconfigure affected components through owner source;
- verify new transactions and negative security behavior;
- observe for recurrence over an explicit window;
- remove temporary containment and reconcile drift;
- only then close or downgrade the incident.
If a recovery mechanism behaves dangerously, stop it. AWS reliability guidance explicitly recommends an abort mechanism for risky automated recovery.
Four-dimensional validation
No incident closes on one green dashboard. Prove:
- Customer: new uncached transactions succeed at expected latency/error rate.
- System: capacity, target health, dependencies, queues, saturation, logs, and alarms recover.
- Security/data: unauthorized access still fails; encryption and integrity hold; no secret leaked.
- Operations: owner source matches runtime, automation is bounded, temporary access/containment is removed, and inventory/cost are understood.
Test replacement/restart behavior so recovery is not dependent on one warm instance or cached credential.
Write the missing runbook
After recovery, convert only proven steps into a human runbook. Include purpose, triggers, prerequisites, authorized role, inputs, read-only diagnosis, branch decisions, exact changes, approvals, timeouts/retries, verification, rollback/compensation, escalation, evidence, cost, and cleanup. Mark unresolved judgment as a playbook decision rather than pretending it is deterministic automation.
Test the runbook against:
- intended fault;
- already-healthy/idempotent state;
- wrong account/Region/tag/resource;
- missing permission;
- partial success and failed verification;
- cancellation and rollback failure;
- concurrent execution.
Only then propose Systems Manager Automation, using AWS245 safety controls.
Post-incident analysis
Write a blameless but accountable timeline. Separate trigger, primary cause, contributing conditions, detection gaps, response gaps, and latent organizational weaknesses. “Human error” is not a sufficient root cause; ask why review, testing, guardrails, observability, or ownership allowed the action to create impact.
Each prevention item needs an owner, due date, priority, acceptance test, and tracking ID. Balance immediate fixes with systemic work. Include what went well, what increased time to detect/recover, and which assumptions were disproved.
Scoring gate
The assessor awards 100 points:
- incident command and communication: 10;
- hypothesis/evidence quality: 20;
- diagnosis and fault interaction: 15;
- safe containment/recovery: 20;
- customer/security/data/operations validation: 15;
- runbook and automation design: 10;
- RCA, prevention, cost, evidence retention, cleanup: 10.
Pass requires 80/100, at least half in every category, and no critical safety violation. Automatic fail conditions include unauthorized mutation, fabricated evidence, secret exposure, broad public access or admin privilege as a fix, destructive action without verified target/rollback, bypassing governance, reading the sealed guide early, or declaring success without customer and security proof.
Hints reduce independence score. An assessor may stop the exercise for safety and still allow learning review; service restoration alone does not override a safety failure.
Cost and cleanup
The supplied track costs nothing. A live sandbox plan must include resource-hour and request estimates, log retention, NAT/endpoints/load balancer/compute/database/backup costs, a hard stop timer, owner tags, and post-cleanup billing review. Cleanup includes stacks and retained resources, instances, ENIs, volumes/snapshots, EIPs, load balancers/targets, DNS, endpoints/NAT, logs/alarms, parameters/secrets, IAM policies/roles, backups, and temporary evidence stores. Never delete shared evidence or audit logs outside the approved retention plan.
Acceptance evidence
Submit the complete workbook, original pre-evidence hypothesis plan, all evidence-release gate timestamps, communication updates, change log, validation matrix, final runbook and tests, post-incident analysis, prevention backlog, cost review, and no-create or empty-inventory proof. The assessor signs the information-separation and safety statements.