AWS 404: Detect, contain, remediate, verify, and document a production-shaped failure
Why this lesson matters
This capstone joins the controls from AWS392-AWS403 into one evidence-led response. The objective is not to guess a root cause quickly. It is to protect customers, preserve evidence, make reversible decisions under uncertainty, restore a verified service, communicate honestly and convert learning into owned improvements.
Scenario and operating rules
At 10:02 UTC a deployment completes for a fictional API behind an Application Load Balancer and Auto Scaling group. At 10:06 the checkout SLI falls, target 5xx rises and queue age grows. A Config change, an AWS Health notice and a dependency timeout also appear near the event. Only one is causal.
Use the supplied CloudWatch metrics/logs, trace samples, ALB access logs, deployment record, CloudTrail events, configuration snapshots, Health event, architecture, runbook and customer-impact data. Do not modify evidence. Record every hypothesis as supported, contradicted or unknown with a precise query/result. Time proximity is not causality.
Roles are incident commander, operations lead, communications lead, subject-matter experts and scribe. The commander owns priority and decisions, not every command. Use one UTC timeline and one decision log. Separate observation, interpretation, action and result. Credentials, personal data and callback tokens never enter the incident channel.
Response lifecycle
Detection proves a user-impact signal crossed a defined threshold and confirms the alarm itself is healthy. Triage determines scope, severity, affected journeys, Regions/accounts/tenants and whether security or data integrity may be involved. Declare the incident early enough to establish authority and communications.
Containment limits impact without destroying forensic evidence. Examples include stopping rollout, shifting traffic to a known-good target, disabling a feature, scaling bounded capacity, quarantining a compromised component or pausing consumers. Predict customer/data effects, define a success metric and prepare reversal before acting.
Diagnosis builds a timeline across metric event time, log event/ingestion time, trace IDs, deployment IDs, configuration versions and CloudTrail actor/request IDs. Compare healthy and failed requests, pre/post-deployment cohorts and affected/unaffected zones. Eliminate plausible alternatives explicitly. Capture uncertainty when evidence cannot distinguish causes.
Remediation changes the causal condition through an approved runbook. Use immutable versions, least privilege, idempotency and bounded waves. Recovery restores traffic or processing, but verification must prove user SLI, error/latency/saturation, security and data outcomes for a bake period. A green instance or successful command is insufficient.
Communication states confirmed impact, scope, mitigation, current risk and next update time. Never promote an untested hypothesis to fact. After recovery preserve evidence, reconcile queued/duplicated/lost work, monitor recurrence, conduct the review and track corrective actions to tested completion.
| Phase | Minimum evidence | Exit condition |
|---|---|---|
| Detect | User SLI, alarm state/history | Signal and alarm path validated |
| Triage | Impact, scope, severity, ownership | Incident declared and roles assigned |
| Contain | Predicted effect, approval, reversal | Impact bounded without worse harm |
| Diagnose | Timeline, comparisons, hypotheses | Causal claim survives contradiction tests |
| Remediate | Runbook/version/change identity | Causal condition changed safely |
| Verify | User, system, security, data and bake | Sustained customer outcome restored |
| Learn | Review, actions and tests | Improvements owned and verified |
Evidence commands
Run these against supplied identifiers or an authorized sandbox only:
aws cloudwatch get-metric-data --metric-data-queries file://queries.json \
--start-time 2026-09-23T09:45:00Z --end-time 2026-09-23T11:00:00Z
aws logs start-query --log-group-names /fictional/api --start-time START --end-time END \
--query-string 'fields @timestamp, @message | filter trace_id="TRACE"'
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=EventName,AttributeValue=UpdateService
aws deploy get-deployment --deployment-id DEPLOYMENT_ID
Record command, region/account, time window, query ID, result hash and interpretation. Missing telemetry is itself a finding but cannot prove an event did not occur.
Lab phases and injects
Phase 1 establishes baselines, ownership, severity and customer statement. Phase 2 evaluates five hypotheses and chooses a containment action. Phase 3 performs supplied runbook results, verifies recovery and reconciles work. Phase 4 writes communications, timeline and review. An assessor releases injects for a stale dashboard, failed rollback, unauthorized responder, duplicate queue processing and apparent recovery followed by regression.
Calculate time to detect, acknowledge, declare, contain, recover and verify from event timestamps. Distinguish service restoration from incident closure. Grade decision quality using information available at that time, not hindsight.
Failure matrix
Diagnose 24 failures: alarm no-data, wrong period, dashboard wrong Region, log ingestion delay, sampling hides errors, clock skew, trace broken at queue, deployment ID mismatch, Config snapshot stale, CloudTrail wrong account, Health coincidence, severity too low, no commander, shared credentials, destructive containment, rollback data incompatibility, retry duplicates orders, remediation unapproved, instance green but SLI red, bake too short, queue backlog ignored, customer update speculates, evidence overwritten, and corrective action untested.
Cost, cleanup and acceptance
No AWS resources are required. If adapted to a sandbox, budget telemetry ingestion/query/storage, notifications, compute overlap and recovery capacity, then remove test alarms, routes and data under retention policy.
Submit incident declaration, role log, impact statement, evidence index, five-hypothesis table, UTC timeline, decision/authorization log, containment and reversal plan, remediation proof, verification/bake report, three stakeholder updates, reconciliation, post-incident review and 24 diagnoses. Pass requires a defensible causal chain, safe actions, honest communication and customer-level recovery evidence.