Lesson 404 · AWS Learning Path

AWS 404: Detect, contain, remediate, verify, and document a production-shaped failure

· Published · 4 min read

Labelled process diagram for AWS 404: Healthy baseline and normally detected failure to Triage, roles, evidence and containment to Proven cause and controlled remediation to User verification, communication, review...

Why this lesson matters

This capstone joins the controls from AWS392-AWS403 into one evidence-led response. The objective is not to guess a root cause quickly. It is to protect customers, preserve evidence, make reversible decisions under uncertainty, restore a verified service, communicate honestly and convert learning into owned improvements.

Scenario and operating rules

At 10:02 UTC a deployment completes for a fictional API behind an Application Load Balancer and Auto Scaling group. At 10:06 the checkout SLI falls, target 5xx rises and queue age grows. A Config change, an AWS Health notice and a dependency timeout also appear near the event. Only one is causal.

Use the supplied CloudWatch metrics/logs, trace samples, ALB access logs, deployment record, CloudTrail events, configuration snapshots, Health event, architecture, runbook and customer-impact data. Do not modify evidence. Record every hypothesis as supported, contradicted or unknown with a precise query/result. Time proximity is not causality.

Roles are incident commander, operations lead, communications lead, subject-matter experts and scribe. The commander owns priority and decisions, not every command. Use one UTC timeline and one decision log. Separate observation, interpretation, action and result. Credentials, personal data and callback tokens never enter the incident channel.

Response lifecycle

Detection proves a user-impact signal crossed a defined threshold and confirms the alarm itself is healthy. Triage determines scope, severity, affected journeys, Regions/accounts/tenants and whether security or data integrity may be involved. Declare the incident early enough to establish authority and communications.

Containment limits impact without destroying forensic evidence. Examples include stopping rollout, shifting traffic to a known-good target, disabling a feature, scaling bounded capacity, quarantining a compromised component or pausing consumers. Predict customer/data effects, define a success metric and prepare reversal before acting.

Diagnosis builds a timeline across metric event time, log event/ingestion time, trace IDs, deployment IDs, configuration versions and CloudTrail actor/request IDs. Compare healthy and failed requests, pre/post-deployment cohorts and affected/unaffected zones. Eliminate plausible alternatives explicitly. Capture uncertainty when evidence cannot distinguish causes.

Remediation changes the causal condition through an approved runbook. Use immutable versions, least privilege, idempotency and bounded waves. Recovery restores traffic or processing, but verification must prove user SLI, error/latency/saturation, security and data outcomes for a bake period. A green instance or successful command is insufficient.

Communication states confirmed impact, scope, mitigation, current risk and next update time. Never promote an untested hypothesis to fact. After recovery preserve evidence, reconcile queued/duplicated/lost work, monitor recurrence, conduct the review and track corrective actions to tested completion.

PhaseMinimum evidenceExit condition
DetectUser SLI, alarm state/historySignal and alarm path validated
TriageImpact, scope, severity, ownershipIncident declared and roles assigned
ContainPredicted effect, approval, reversalImpact bounded without worse harm
DiagnoseTimeline, comparisons, hypothesesCausal claim survives contradiction tests
RemediateRunbook/version/change identityCausal condition changed safely
VerifyUser, system, security, data and bakeSustained customer outcome restored
LearnReview, actions and testsImprovements owned and verified

Evidence commands

Run these against supplied identifiers or an authorized sandbox only:

aws cloudwatch get-metric-data --metric-data-queries file://queries.json \
  --start-time 2026-09-23T09:45:00Z --end-time 2026-09-23T11:00:00Z
aws logs start-query --log-group-names /fictional/api --start-time START --end-time END \
  --query-string 'fields @timestamp, @message | filter trace_id="TRACE"'
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=EventName,AttributeValue=UpdateService
aws deploy get-deployment --deployment-id DEPLOYMENT_ID

Record command, region/account, time window, query ID, result hash and interpretation. Missing telemetry is itself a finding but cannot prove an event did not occur.

Lab phases and injects

Phase 1 establishes baselines, ownership, severity and customer statement. Phase 2 evaluates five hypotheses and chooses a containment action. Phase 3 performs supplied runbook results, verifies recovery and reconciles work. Phase 4 writes communications, timeline and review. An assessor releases injects for a stale dashboard, failed rollback, unauthorized responder, duplicate queue processing and apparent recovery followed by regression.

Calculate time to detect, acknowledge, declare, contain, recover and verify from event timestamps. Distinguish service restoration from incident closure. Grade decision quality using information available at that time, not hindsight.

Failure matrix

Diagnose 24 failures: alarm no-data, wrong period, dashboard wrong Region, log ingestion delay, sampling hides errors, clock skew, trace broken at queue, deployment ID mismatch, Config snapshot stale, CloudTrail wrong account, Health coincidence, severity too low, no commander, shared credentials, destructive containment, rollback data incompatibility, retry duplicates orders, remediation unapproved, instance green but SLI red, bake too short, queue backlog ignored, customer update speculates, evidence overwritten, and corrective action untested.

Cost, cleanup and acceptance

No AWS resources are required. If adapted to a sandbox, budget telemetry ingestion/query/storage, notifications, compute overlap and recovery capacity, then remove test alarms, routes and data under retention policy.

Submit incident declaration, role log, impact statement, evidence index, five-hypothesis table, UTC timeline, decision/authorization log, containment and reversal plan, remediation proof, verification/bake report, three stakeholder updates, reconciliation, post-incident review and 24 diagnoses. Pass requires a defensible causal chain, safe actions, honest communication and customer-level recovery evidence.

Official sources

Advertisement