AWS 421: Recover the final platform and preserve audit evidence
Incident objective
Recover the AWS420 platform from a compound failure while preserving a defensible evidence chain. The learner must separate coincidence from cause, choose safe containment under uncertainty, restore user and data outcomes, remove emergency access and turn findings into verified controls.
Scenario
At 14:00 UTC a release begins. At 14:08 checkout latency rises, at 14:10 a deployment alarm fires, at 14:12 queue age climbs, and at 14:14 the primary responder cannot federate. One task definition references the wrong secret stage, a security-group emergency edit broadens ingress, rollback uses a database-incompatible image, the notification target throttles, and an AWS Health event affects an unrelated service.
| Evidence packet | Questions to answer |
|---|---|
| Pipeline/change set/task/deployment histories | What changed, by whom and in what order? |
| SLI metrics, logs, traces and ALB/SQS evidence | Which user journeys and dependencies failed? |
| Secret versions/application auth failures | Did rotation/configuration contribute? |
| CloudTrail/Config/Security Hub | Was emergency access/change authorized and preserved? |
| Health event and account/resource pages | Relevant impact or temporal coincidence? |
| Backup/restore and reconciliation manifests | What data can be recovered without duplication/loss? |
Response phases
Declare severity, customer impact, commander, operations/communications/security leads, scribe and next-update time. Use one UTC timeline and decision log. Activate emergency access only under its criteria, with short session and immediate alerting. Never paste credentials or sensitive request data into the channel.
Triage scope by Region/AZ, endpoint, release cohort, resource and data state. Establish five hypotheses and record supporting, contradicting and missing evidence. Event proximity is not causality; Health envelope Region/account fields are not necessarily impact fields.
Contain by stopping traffic shift and preventing further incompatible deploys. Choose whether to route to known-good capacity, disable a feature, pause consumers or isolate a component. Before action state expected benefit, customer/data/security risk, authorization, reversal and success metric. Preserve failed tasks/logs/configuration where safe.
Recover with a data-compatible artifact or roll forward configuration; do not deploy the incompatible rollback merely because it is previous. Correct the secret stage/reference under source-controlled configuration, restore least-privilege security group state and redrive/reconcile queued work using idempotency keys. Restore notification delivery/DLQ processing without duplicating containment.
Verify sustained availability, latency, order correctness, queue freshness, authorization/security, data counts and alarms through a bake period. Distinguish technical restoration, data reconciliation and incident closure. Revoke emergency sessions/policies, reconcile IaC/drift and preserve originals.
Evidence custody and timeline
Record original S3 object/version/hash, CloudTrail event/request IDs, Config snapshots, log query IDs/results, trace IDs, deployment/artifact digests, secret stages without values, commands, analyst identity and collection time. Work on copies with access logs. Validate CloudTrail digest chain and state its coverage limits.
~~~bash aws codepipeline get-pipeline-execution --pipeline-name PIPELINE --pipeline-execution-id ID aws ecs describe-services --cluster PLATFORM --services order-api aws secretsmanager list-secret-version-ids --secret-id SECRET_ARN aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=AuthorizeSecurityGroupIngress aws configservice get-resource-config-history --resource-type AWS::EC2::SecurityGroup --resource-id SG_ID aws cloudtrail validate-logs --trail-arn TRAIL_ARN --start-time START --end-time END ~~~
Do not run mutation commands against real systems in this supplied-evidence lab. A production adaptation requires incident authority and command peer review.
Injects and calculations
Injects arrive when the responder assumes a single cause: identity provider remains unavailable, secret candidate works but cache is stale, rollback fails compatibility test, queue redrive duplicates one message, digest validation reports a temporary delivery gap and a second alarm appears during apparent recovery.
Calculate detection, acknowledgement, declaration, containment, restoration, reconciliation and verification intervals; queue drain time; actual RTO/RPO; duplicate/lost/late order counts; emergency-session exposure; notification loss/retry/DLQ reconciliation; and customer-update cadence. Mark measured versus estimated.
Build a recovery ledger for every order state before the event, accepted during degradation, queued, processed, duplicated, compensated, failed permanently and still unknown. Totals must reconcile with source-of-truth counts and idempotency records. Sample individual journeys from request through queue, database and customer result. If reconciliation cannot prove correctness, keep the incident open with explicit residual risk and owner rather than declaring recovery from aggregate HTTP health.
Diagnose 24 failures: wrong severity, no commander, clock zones mixed, Health event blamed, metric ingestion mistaken event time, trace link absent, release cohort ignored, emergency role too broad, access unlogged, containment destroys evidence, rollback assumed safe, secret value exposed, cache behavior ignored, SG fixed outside IaC, queue redrive non-idempotent, DLQ unmonitored, evidence object overwritten, digest gap called tampering without delivery check, instance health used as recovery, data reconciliation skipped, security verification skipped, emergency access retained, root cause called human error, and action closed untested.
Review and acceptance
Write communications at declaration, containment, restoration and closure. Produce a blameless review separating trigger, contributors and control failures. Corrective actions require owner, due date, priority, acceptance test and recurrence metric; update ORR/runbooks/pipeline controls.
Submit role/decision logs, evidence index/custody, hypothesis table, timeline, containment/recovery/verification, calculations, communications, audit reconstruction, post-incident review and all diagnoses. Pass requires evidence-led causality, safe data-compatible recovery, customer/data/security proof, removed emergency access and tested corrective actions.