AWS 403: AWS Health events, operational readiness, game days, and post-incident review
Why this lesson matters
Reliable operations begin before an alarm. Architects must discover AWS changes that affect their accounts, prove a workload is supportable, rehearse failure safely, and turn incidents into completed engineering changes. AWS Health, an operational readiness review, a game day and a post-incident review solve different parts of that lifecycle.
AWS Health event model
AWS Health publishes public service events and account-specific events. An account event can identify affected resources; a public event may be relevant even when no resource is listed. In EventBridge, source is aws.health, but the envelope account and region describe event delivery. For organizational events use detail.affectedAccount and detail.eventRegion for impact. Do not route on the wrong fields.
Event details include status, category, event type code, start/end times, actionability, personas and affected entities. Resource lists can arrive on multiple pages; reconcile event ARN/account plus page and totalPages before declaring inventory complete. Handle updates and duplicates idempotently. A backup event is not a second independent incident unless policy says so.
Organizational view aggregates member-account events in the management or delegated administrator account without disabling local rules. Keep local routing for account survival and central routing for governance. The console organizational view is available broadly, while Health API entitlement depends on the current support plan. Archive normalized events because console/API history is not an indefinite audit store.
| Evidence | Question it answers | Common mistake |
|---|---|---|
| Event ARN/type/status | What changed and is it current? | Treating an update as a new incident |
| Affected account/Region/entities | What is actually exposed? | Using delivery account/Region instead |
| Page/total pages | Is resource inventory complete? | Acting after only page one |
| Actionability/personas | Who must decide or act? | Routing every notice as critical |
| Central and local delivery logs | Did notification paths work? | Assuming console visibility proves delivery |
aws health describe-events --region us-east-1 \\
--filter eventStatusCodes=open,upcoming
aws health describe-events-for-organization --region us-east-1
aws events list-rules --event-bus-name default --region us-east-1
Readiness review
An operational readiness review is a release gate, not a questionnaire signed after launch. For each workload record owner/on-call/escalation, user journeys and SLOs, dependency and quota headroom, deployment/rollback, backup/restore, security response, observability, runbooks, capacity and cost, continuity, data classification/retention, Health routing and known exceptions. Every item needs evidence, result, risk owner, due date and blocking rule.
Reject readiness when monitoring observes components but not customer outcomes, restoration has never been tested, quotas lack demand math, one person holds critical knowledge, rollback is incompatible with data changes, or an exception has no expiry. Re-run relevant checks after architecture, dependency, traffic, team or incident changes. Post-incident actions should improve the reusable checklist.
Safe game days
Write a hypothesis and steady-state measures first. Define exact accounts/resources, blast radius, fault method, observers, communications, stop conditions, abort authority, recovery procedure and evidence clock. Begin with tabletop and isolated tests, then progress only when recovery is proven. Never improvise destructive experiments in production.
Exercise people and process as well as technology: missing primary responder, stale runbook, unavailable identity provider, noisy alarm, regional control-plane issue and communication delay. Capture UTC detection, acknowledgement, diagnosis, containment, recovery and verification times. A successful game day can reveal a failed control; it is not merely a system that stayed green.
Post-incident learning
Preserve a factual timeline before memory changes. Separate trigger, contributing conditions, control failures, impact and recovery from blame. Ask why defenses allowed customer impact and why detection/recovery behaved as observed. Distinguish immediate containment from durable prevention.
Each corrective action needs an owner, priority, due date, acceptance evidence and verification method. Track recurrence, repeated contributing factors, overdue actions, detection/acknowledgement/recovery time and runbook effectiveness. Closing a ticket without testing the change is administrative closure, not risk reduction. Share sanitized lessons with teams that own similar systems.
Workshop and failure analysis
Given twelve Health event envelopes, reconstruct affected accounts/Regions/resources, detect missing pages and duplicates, classify actionability/persona, and design central plus local EventBridge routes. Build an ORR for the incident workflow from AWS399-AWS402. Plan a game day with four injections and write a blameless review from the supplied timeline.
Test 20 failures: public/account confusion, wrong Region field, wrong account field, missing page, duplicate update, backup duplicate, expired history, central route outage, local route absent, API entitlement error, stale owner, unchecked quota, untested restore, unsafe blast radius, stop alarm shares failed dependency, no abort owner, timeline clock mismatch, root cause reduced to human error, action without owner, and recurrence never measured.
Cost and acceptance
Price EventBridge, notification targets, archive/storage, observability, experiments and engineering time. This lesson creates nothing. Submit event reconciliation, routing diagram, ORR with evidence, game-day charter, timeline, post-incident review, action tracker and all failure diagnoses. Pass requires correct Health fields, bounded experimentation, objective stop/recovery proof and verified corrective actions.