AWS 299: Architecture: design and tabletop-test an enterprise DR runbook
Why this lesson matters
An enterprise disaster-recovery plan is credible only when business owners, responders, platforms, data, security, networks, providers, and communications can make correct decisions together. A technically elegant secondary Region does not recover the business if no one can declare disaster, emergency identity fails, a replica is stale, a partner blocks the new source IP, or the runbook has no data-safe response to failed failback.
This P17 capstone integrates AWS293 through AWS298. You will not create resources. You will produce a complete architecture decision record (ADR), recovery runbook, evidence model, and tabletop package for one fictional enterprise service. Then another learner or reviewer will lead a blind scenario that challenges the happy path.
A tabletop proves reasoning, roles, access assumptions, communication, decision criteria, and document quality. It does not prove API behavior, capacity, restore duration, or real RPO/RTO. The final report must say exactly what was and was not proven and must convert gaps into funded, owned technical tests.
What you will be able to do
By the end, you can:
- produce an enterprise DR ADR that connects BIA objectives to a feasible technical strategy;
- write a runbook that a different authorized operator can follow under pressure;
- distinguish plan review, tabletop, simulation, recovery drill, game day, and live failover;
- design facilitator-only injects that expose hidden dependencies and decision ambiguity;
- assign incident, technical, business, security, vendor, and communication authority;
- measure declaration, detection, decision, recovery, RPO, RTO, and business acceptance clocks;
- evaluate one-writer data integrity, traffic, failback, and ransomware recovery branches;
- score evidence without rewarding confident unsupported answers;
- run a blameless retrospective and remediation program; and
- defend residual risk, cost, and retest requirements before an architecture board.
Before you start
- This capstone is a tabletop only. Participants must not run commands, start workflows, restore data, change DNS, contact real vendors, or invoke emergency access.
- Use the supplied fictional names and values. Real DR documents contain sensitive weaknesses, recovery locations, credentials, contacts, and provider relationships.
- The facilitator keeps the inject schedule private from players. Players receive only the initial scenario and evidence released during the exercise.
- “We would check” is not proof. Players must name the exact evidence, owner, decision threshold, and time budget.
- Exercise failure is useful. Score the plan and evidence, not personal confidence or job title.
1. Know what this capstone must produce
Submit one controlled package:
p17-enterprise-dr-adr.mdp17-recovery-runbook.mdp17-evidence-register.mdp17-role-and-contact-matrix.mdp17-tabletop-facilitator-guide.mdp17-tabletop-player-brief.mdp17-inject-deck.md, facilitator-only until exercise closep17-execution-log.mdp17-after-action-report.mdp17-remediation-register.md
Version every artifact. Record owner, approver, sensitivity, source, creation date, last exercise, next review, and change history. Store an accessible offline or recovery-account copy according to security policy. A link available only through the failed identity or network path is not a recovery artifact.
2. The supplied enterprise scenario
Aster Market runs a 24x7 ordering service:
- public API and web frontend behind Route 53 and Regional ALBs;
- ECS services in
ap-south-1, warm standby ineu-west-1; - Aurora-compatible primary database with asynchronous cross-Region secondary;
- S3 order documents replicated to the recovery Region;
- SQS order and fulfillment queues, ElastiCache sessions, and EventBridge schedules;
- workforce federation through an external identity provider plus emergency recovery roles;
- central KMS, Secrets Manager, ACM, CloudWatch, CloudTrail, GuardDuty, Security Hub, AWS Backup, and a separate backup account;
- Transit Gateway, Direct Connect plus VPN backup, Route 53 Resolver endpoints, interface endpoints, and centralized inspection;
- payment, carrier, email, fraud, and warehouse providers with IP allowlists; and
- Step Functions plus Systems Manager Automation recovery orchestration.
Business objectives:
- MTPD 4 hours;
- MBCO accepts priority orders at 30 percent peak throughput, without recommendation or email features;
- RTO 75 minutes to business-accepted MBCO;
- RPO 5 minutes for accepted orders and payment state, 30 minutes for documents, 24 hours for analytics;
- no acknowledged order may silently disappear;
- maximum 15 minutes to declare a disaster after confirmed Regional unavailability; and
- source must not accept writes after target write authority is granted.
The recovery environment is scaled to 40 percent normal peak. Last full recovery drill was nine months ago at half today's data volume. The latest tabletop was four months ago but did not include business or vendor participants.
3. Write the architecture decision record
The ADR must contain:
Decision and scope
- one-paragraph decision;
- business process, customers, minimum service, data classes, and scenarios;
- primary/recovery Regions and accounts;
- inclusions, exclusions, assumptions, constraints, and evidence confidence; and
- accountable business, service, data, security, and recovery owners.
Objectives and measured capability
Show MTPD, MBCO, RTO, dataset-level RPO, SLO, detection/declaration target, backlog/catch-up limits, and regulatory obligations. Compare each with current estimated capability and last measured RTC/RPC. Never relabel an estimate as a drill result.
Alternatives
Compare at least:
- backup and restore;
- pilot light;
- warm standby; and
- multi-Region active-active.
Assess failure coverage, data semantics, control-plane reliance, operator complexity, RTO/RPO, capacity, security, licensing, provider compatibility, testing burden, one-time and recurring cost, and residual risk. Explain why the selected warm-standby design wins and where it does not protect the business.
Complete architecture
Draw separate diagrams for:
- account/organization and trust boundaries;
- identity, emergency access, and approval paths;
- normal and recovery network/DNS/traffic paths;
- every authoritative dataset, replication direction, lag, backup, and write authority;
- application, queue/event, cache, scheduler, and provider dependencies;
- observability, security, incident command, and evidence paths; and
- failback and source decommission/isolation.
Decisions and consequences
Record strategy by component, traffic mechanism, data promotion/fencing, backup isolation, corruption response, orchestration boundary, manual approvals, failback method, exercise cadence, costs, known gaps, compensating controls, residual-risk owner, expiry, and revisit triggers.
4. Write an operator-ready recovery runbook
The runbook begins with a one-page emergency summary:
- purpose and scenario scope;
- current approved version and change record;
- incident bridge/out-of-band communication;
- command authority and alternates;
- disaster-declaration criteria;
- current RTO/RPO/MBCO and decision deadlines;
- stop/hold/fix-forward authority;
- emergency identity procedure reference;
- source-fencing invariant; and
- evidence and execution-log location.
Required phases
- Detect, classify, and establish command.
- Obtain emergency access and preserve evidence.
- Declare or reject disaster; communicate impact.
- Freeze changes and fence unsafe writers.
- Verify recovery account, Region, identity, KMS, secrets, quotas, network, DNS, security, observability, and provider readiness.
- Select a data/recovery point and record effective RPO.
- Recover/promote data, files, queues, and workflow state read-only where possible.
- Validate data integrity and one-writer preconditions.
- Authorize target writes; start applications, consumers, and schedulers in order.
- Run technical, security, integration, capacity, and business validation.
- Shift canary then approved traffic.
- Operate MBCO, reconcile backlog, and enter hypercare.
- Decide and execute failback only as a separate controlled recovery.
- Close incident, retain evidence, clean temporary access/resources, and update risk.
Every task row
| Field | Requirement |
|---|---|
| ID/phase/version | Stable reference to immutable automation or procedure |
| Planned/maximum duration | Supports critical-path and decision-deadline math |
| Preconditions/invariants | Exact evidence required before action |
| Action and parameters | Approved path, typed inputs, account/Region guard |
| Executor/verifier | Named role plus alternate |
| Expected output | State/value, not “success” |
| Validation/evidence | Independent query/test, timestamp, artifact |
| Idempotency | Prior-state check and correlation key |
| Failure decision | Retry, hold, compensate, fix forward, or abort |
| Compensation | Data-safe counterpart and owner |
| Actuals | Start/end, result, deviation, decision reference |
The runbook must survive interruption. A replacement commander should reconstruct current state from the execution log and evidence, not from memory or chat history.
5. Define roles and decision rights
Include primary and alternate for:
- incident commander;
- deputy/timekeeper;
- technical recovery lead;
- application, database/data, identity, network/DNS, platform, security, backup, and observability leads;
- business decision owner and business validator;
- continuity/risk and legal/privacy representative;
- communications/customer-support lead;
- external provider coordinator;
- scribe/evidence custodian; and
- tabletop facilitator/controllers/observers, who do not make player decisions.
Create a decision-rights table:
| Decision | Recommends | Decides | Evidence | Deadline | Alternate |
|---|---|---|---|---|---|
| Declare disaster | Incident/technical leads | Business/incident authority per policy | Scope, impact, source status | T+15m | Named deputy |
| Accept recovery point/data loss | Data and business leads | Data/business risk owner | Lag, timeline, reconciliation | Before promotion | Named alternate |
| Grant target writes | Data/security/application leads | Incident authority | Source fenced, target validated | Before writers | Named deputy |
| Shift traffic | Technical/business validators | Incident authority | Full gate packet | Before RTO deadline | Named deputy |
| Fail back | Recovery/data/business leads | Change/business authority | Stable source, synced data, test | Separate window | Named alternate |
Authority must remain valid when normal identity, phones, or leaders are unavailable. Test delegation and conflict resolution.
6. Understand tabletop limits and progression
| Exercise | What occurs | Strong evidence | Cannot prove alone |
|---|---|---|---|
| Document review | Experts inspect artifacts | Completeness and design issues | Human response or technology |
| Tabletop | Players discuss decisions against injects | Roles, decisions, dependencies, document usability | API, duration, capacity, actual recovery |
| Simulation | Tools or isolated systems mimic behavior | Integrations and workflow logic | Full production equivalence |
| Recovery drill | Real resources/data restored in isolation | Measured technical RTC/RPC and validation | Production traffic behavior unless included |
| Game day | Teams act in realistic/prod-like system | Technology, people, process under conditions | Every future scenario |
| Controlled failover | Production authority/traffic moves | Highest realism for tested path | Untested corruption or different failures |
The capstone tabletop should create the technical test backlog for claims it cannot prove. Never publish “RTO met” from a discussion. Publish “players produced a plausible 62-minute plan; measured RTO remains unproven.”
7. Prepare the tabletop
Exercise objectives
Test whether players can:
- detect and classify the event;
- establish command without normal federation;
- use the runbook and evidence rather than intuition;
- decide disaster within 15 minutes;
- preserve forensic evidence while meeting recovery objectives;
- select a safe recovery point and keep one writer;
- achieve MBCO despite optional-provider failure;
- communicate uncertainty and customer impact;
- respond to failed traffic change and unavailable experts; and
- plan data-safe failback.
Rules of play
- Exercise time advances only when the facilitator announces it.
- No real system actions or vendor/customer messages occur.
- Players can request evidence; controllers return only prepared artifacts or “not available.”
- Assumptions are written and scored as evidence gaps.
- An inject cannot be wished away by saying a specialist handles it.
- Controllers do not rescue the team with the intended answer.
- Safety phrase
STOP EXERCISEpauses the simulation for real-world risk or distress. - The discussion is blameless; artifacts and systems are evaluated.
Prerequisites
Approve scope, objectives, date/time/time zone, players, observers, confidentiality, no-action boundary, safety process, facilitator script, injects, evidence cards, scoring rubric, baseline architecture, incident log, communication templates, and retrospective schedule. Verify participants can access sanitized exercise documents.
8. Build facilitator injects
Each inject card contains:
- inject ID and simulated time;
- prerequisite or player action that releases it;
- information given to players;
- private purpose and expected decisions;
- evidence available on request;
- incorrect assumptions to challenge;
- escalation if players stall;
- clock impact; and
- scoring notes.
Do not script one correct dialogue. Let decisions change later injects. If players choose not to recover, test how they manage business continuity and risk.
9. Run the blind scenario
Release these injects progressively:
Inject 0, T+00: initial event
Customers in several networks cannot complete checkout in the primary Region. ALB 5xx and database connection failures rise. AWS Health information is inconclusive. The normal incident channel is available.
Test detection, classification, commander selection, baseline impact, evidence requests, and declaration clock.
Inject 1, T+05: normal identity fails
The external identity provider stops issuing sessions. Existing sessions expire in 20 minutes. One emergency-role owner is unreachable.
Test independent access, alternate authority, credential custody, CloudTrail, and out-of-band communication.
Inject 2, T+12: source status uncertain
Some primary API hosts answer, but database write status cannot be confirmed. A queued order may have been acknowledged after the last visible database commit.
Test source fencing, transaction ledger, duplicate/loss handling, and whether players avoid creating two writers.
Inject 3, T+20: stale recovery data
The cross-Region database reports 11 minutes lag, breaching the 5-minute RPO. An isolated backup point is 42 minutes old. Business impact reaches the hospital-priority threshold in 55 minutes.
Test data-loss decision rights, alternate reconstruction/replay, MBCO, and honest RPO reporting.
Inject 4, T+28: possible ransomware
Security reports privileged configuration changes began 36 hours earlier; malware status of recent recovery points is unknown. The newest immutable backup predates the changes by 39 hours.
Test containment, evidence preservation, clean-room branch, point selection, credential rotation, and whether speed improperly overrides safety.
The facilitator chooses either ransomware branch or infrastructure-failure branch for the remainder based on exercise goals. Players must explain how RTO changes.
Inject 5, T+37: recovery capacity and provider gap
Recovery ECS capacity supports 40 percent peak, but payment and carrier providers have not allowlisted recovery egress. The provider coordinator estimates 30 minutes and needs an approved caller.
Test minimum-service design, external dependency, capacity, authorization, and communications.
Inject 6, T+45: data promotion timed out
The orchestration task timed out. The API response was lost, but monitoring suggests promotion might have succeeded. A workflow redrive is available.
Test inspect-before-retry, idempotency, current-state evidence, write authority, and redrive hazard.
Inject 7, T+52: DNS/traffic error
A responder changed the Route 53 record early. Some resolvers send users to a target that remains read-only. Cached clients still use the primary.
Test stop/compensation, TTL/cache reasoning, communication, source/target safety, and validation gates.
Inject 8, T+61: business validator unavailable
The primary business approver loses connectivity. The alternate discovers that synthetic test orders send real email and could create carrier jobs.
Test alternate authority, safe validation data, side-effect controls, and RTO impact.
Inject 9, T+72: failed compensation
The attempt to restore the previous traffic route fails because emergency role permission is missing. Error rate is above the stop threshold, but target priority-order transactions succeed.
Test fix-forward versus rollback, permission escalation, objective decision criteria, and customer messaging.
Inject 10, post-recovery: failback conflict
Primary Region returns. Recovery has accepted 3,200 orders not present in the primary database. Management asks for immediate automatic failback before peak traffic.
Test separate failback change, reverse data synchronization, reconciliation, stable period, risk authority, and refusal of unsafe executive pressure.
10. Keep one authoritative exercise log
For every event record:
| Field | Meaning |
|---|---|
| Simulated/real time | Scenario clock and exercise clock |
| Event/inject | What became known |
| Observation vs assumption | Evidence quality |
| Decision | GO, HOLD, DECLARE, FIX FORWARD, ROLLBACK, DEGRADE, or ESCALATE |
| Decider and authority | Who made it and why authorized |
| Alternatives | Options considered and rejected |
| Evidence | Exact card, runbook section, metric, or missing item |
| Expected result/deadline | What happens next and by when |
| Actual simulated result | Controller response |
| Gap/action | Defect to track after exercise |
Start clocks for event, detection, command, declaration, recovery start, source fencing, data selected, target write authority, technical ready, business accepted, traffic shifted, MBCO, full service, and exercise close.
11. Score evidence, not performance theater
Score each 0-3:
0: missing or unsafe;1: stated but owner/evidence/threshold absent;2: complete on paper but untested or slow;3: clear, evidence-backed, exercised, and within target.
Domains:
- business objectives and MBCO;
- declaration and command authority;
- emergency identity and communication;
- dependency completeness;
- source fencing and data integrity;
- backup/corruption/ransomware branch;
- network/DNS/traffic control;
- application/provider/minimum service;
- observability/security/evidence;
- RTO/RPO measurement;
- compensation/fix-forward/failback;
- runbook usability and versioning;
- business/customer communication;
- cost, quota, licensing, and capacity; and
- remediation ownership and retest.
Any zero in emergency access, one-writer safety, data-loss authority, security containment, or business acceptance is a capstone fail regardless of total score. A tabletop cannot receive level 3 for measured technical recovery unless prior representative drill evidence is supplied and current.
12. Conduct a blameless hotwash and retrospective
Immediately ask:
- What surprised us?
- Where did the team wait or debate?
- Which evidence was unavailable, stale, or contradictory?
- Which runbook step was ambiguous or unsafe?
- Which dependency, person, provider, quota, key, or credential was hidden?
- Where did decisions exceed their deadline?
- Which automation could repeat or leave partial state?
- What would real customers have experienced?
- What claim remains unproven by a tabletop?
Within the agreed period, publish an after-action report containing objectives, scenario, participants, decisions, measured exercise times, evidence, strengths, findings, unsafe near misses, objective gaps, tabletop limitations, cost impact, and residual risk. Do not edit the original log to make the exercise look cleaner.
13. Convert findings into closed improvements
Each action requires ID, finding/evidence, risk, corrective change, owner, funding, due date, interim control, acceptance evidence, retest type, and closure approver. Categories include architecture, automation, runbook, access, observability, data, vendor, capacity/quota, communications, training, and business policy.
Use closure states:
- open;
- funded/planned;
- implemented but untested;
- technically retested;
- business accepted; or
- risk accepted until an expiry date.
“Documentation updated” does not close a failed recovery mechanism. A runbook defect can close through peer walkthrough; a database promotion defect needs a technical drill; a cross-team decision gap needs another tabletop; full RTO/RPO claims require an integrated exercise.
Schedule the next exercise before closing the current one. Rotate scenarios, recovery points, absent people, workload peaks, vendors, and facilitators while keeping core recovery paths few and repeatable.
14. Read-only supporting evidence
Use supplied evidence by default. If authorized, these inventory summaries may support the ADR but do not prove recovery:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws backup list-protected-resources \
--query 'Results[].{Type:ResourceType,LastBackup:LastBackupTime}'
aws drs describe-source-servers \
--query 'items[].{Lifecycle:lifeCycle.state,Replication:dataReplicationInfo.dataReplicationState,Lag:dataReplicationInfo.lagDuration}'
aws route53 list-health-checks \
--query 'HealthChecks[].{Id:Id,Type:HealthCheckConfig.Type,Disabled:HealthCheckConfig.Disabled}'
aws fis list-experiments \
--query 'experiments[].{Id:id,Status:state.status,Created:creationTime}'
Redact identifiers. Link each result to date, account, Region, owner, and interpretation. A green health check or backup timestamp cannot replace the capstone's business and dependency evidence.
15. Cost and governance
The ADR must show steady DR cost and exercise/recovery cost: secondary compute and database, replication, backup/copies/locks, storage, data transfer, network connectivity, DNS/accelerator, keys/secrets/certificates, observability/security, support, licenses, quotas/reservations, orchestration, vendor retainers, staff/on-call/training, drills, clean room, failback, and backlog recovery.
Cost must correspond to capability. A 40-percent warm standby needs measured scale-up and quota evidence. An unstaffed overnight response changes RTO. A provider without a recovery SLA is not made reliable by internal automation.
Governance approves BIA/objectives, architecture, residual risk, runbook versions, exercise cadence, evidence expiry, access review, vendor review, capacity/load tests, data restoration, and action closure. Review after every material architecture/provider/team change and incident, not only annually.
Diagnose a misleading capstone
| Symptom | Hidden failure | Required correction |
|---|---|---|
| Players know all injects | Discussion rehearses answers rather than decisions | Separate facilitator deck and blind release |
| Tabletop claims RTO met | No technology or data actually recovered | Report decision timeline only; schedule drill |
| Runbook says “DB team promotes” | No preconditions, authority, evidence, timeout, or retry safety | Add complete step contract |
| Business absent | MBCO, data loss, and acceptance lack authority | Include owner and alternate |
| DNS is changed first | Target data/write readiness is unproven | Gate traffic after fencing and validation |
| Ransomware branch uses latest point | Replicated compromise can return | Timeline, clean room, scan/hunt, older candidate |
| Findings close as “noted” | No implementation evidence or retest | Track owner, due date, acceptance, exercise |
| High score hides data risk | Critical unsafe domain averaged away | Apply mandatory fail domains |
Knowledge check
- What can a tabletop prove?
Decision, role, dependency, communication, and document behavior under a simulated scenario.
- What can it not prove?
Actual API behavior, capacity, restore duration, data integrity, or measured technical RPO/RTO.
- Why are facilitator injects hidden?
Players must respond to new information rather than recite a scripted happy path.
- What is the central data invariant?
Only one authorized writer set exists, with an explicit and evidenced transfer of authority.
- Why include unavailable personnel?
Recovery must work through alternates and delegated authority, not one expert.
- Why is immediate failback unsafe?
Recovery contains new writes and the original environment needs repair, synchronization, validation, and a separate change.
- What makes a finding closed?
Implemented change plus evidence from the appropriate retest and accountable acceptance.
- What is the final capstone standard?
A traceable ADR and usable runbook that survive blind injects, expose unproven claims, and produce funded retests.
Lesson acceptance
P17 passes only when the package includes:
- approved objectives, scope, evidence confidence, and complete ADR alternatives;
- full account/identity/network/data/application/security/provider/failback architecture;
- an operator-ready phased runbook with task contracts and deadlines;
- primary/alternate roles and explicit decision rights;
- facilitator/player separation and a controlled exercise plan;
- all 11 injects with branching evidence and scoring;
- one authoritative execution log and every required clock;
- honest tabletop limitations and measured-evidence gaps;
- mandatory safety passes for emergency access, one writer, data-loss authority, containment, and business acceptance;
- after-action report and blameless findings;
- funded owners, due dates, interim controls, acceptance evidence, and retests; and
- a scheduled progression from tabletop to representative technical recovery.