Lesson 299 · AWS Learning Path

AWS 299: Architecture: design and tabletop-test an enterprise DR runbook

· Published · 15 min read

Labelled process diagram for AWS 299: Business declaration and incident command to Dependency-aware technical recovery to Business validation and traffic decision to Failback, evidence, lessons, and funded actions...

Why this lesson matters

An enterprise disaster-recovery plan is credible only when business owners, responders, platforms, data, security, networks, providers, and communications can make correct decisions together. A technically elegant secondary Region does not recover the business if no one can declare disaster, emergency identity fails, a replica is stale, a partner blocks the new source IP, or the runbook has no data-safe response to failed failback.

This P17 capstone integrates AWS293 through AWS298. You will not create resources. You will produce a complete architecture decision record (ADR), recovery runbook, evidence model, and tabletop package for one fictional enterprise service. Then another learner or reviewer will lead a blind scenario that challenges the happy path.

A tabletop proves reasoning, roles, access assumptions, communication, decision criteria, and document quality. It does not prove API behavior, capacity, restore duration, or real RPO/RTO. The final report must say exactly what was and was not proven and must convert gaps into funded, owned technical tests.

What you will be able to do

By the end, you can:

  • produce an enterprise DR ADR that connects BIA objectives to a feasible technical strategy;
  • write a runbook that a different authorized operator can follow under pressure;
  • distinguish plan review, tabletop, simulation, recovery drill, game day, and live failover;
  • design facilitator-only injects that expose hidden dependencies and decision ambiguity;
  • assign incident, technical, business, security, vendor, and communication authority;
  • measure declaration, detection, decision, recovery, RPO, RTO, and business acceptance clocks;
  • evaluate one-writer data integrity, traffic, failback, and ransomware recovery branches;
  • score evidence without rewarding confident unsupported answers;
  • run a blameless retrospective and remediation program; and
  • defend residual risk, cost, and retest requirements before an architecture board.

Before you start

  • This capstone is a tabletop only. Participants must not run commands, start workflows, restore data, change DNS, contact real vendors, or invoke emergency access.
  • Use the supplied fictional names and values. Real DR documents contain sensitive weaknesses, recovery locations, credentials, contacts, and provider relationships.
  • The facilitator keeps the inject schedule private from players. Players receive only the initial scenario and evidence released during the exercise.
  • “We would check” is not proof. Players must name the exact evidence, owner, decision threshold, and time budget.
  • Exercise failure is useful. Score the plan and evidence, not personal confidence or job title.

1. Know what this capstone must produce

Submit one controlled package:

  1. p17-enterprise-dr-adr.md
  2. p17-recovery-runbook.md
  3. p17-evidence-register.md
  4. p17-role-and-contact-matrix.md
  5. p17-tabletop-facilitator-guide.md
  6. p17-tabletop-player-brief.md
  7. p17-inject-deck.md, facilitator-only until exercise close
  8. p17-execution-log.md
  9. p17-after-action-report.md
  10. p17-remediation-register.md

Version every artifact. Record owner, approver, sensitivity, source, creation date, last exercise, next review, and change history. Store an accessible offline or recovery-account copy according to security policy. A link available only through the failed identity or network path is not a recovery artifact.

2. The supplied enterprise scenario

Aster Market runs a 24x7 ordering service:

  • public API and web frontend behind Route 53 and Regional ALBs;
  • ECS services in ap-south-1, warm standby in eu-west-1;
  • Aurora-compatible primary database with asynchronous cross-Region secondary;
  • S3 order documents replicated to the recovery Region;
  • SQS order and fulfillment queues, ElastiCache sessions, and EventBridge schedules;
  • workforce federation through an external identity provider plus emergency recovery roles;
  • central KMS, Secrets Manager, ACM, CloudWatch, CloudTrail, GuardDuty, Security Hub, AWS Backup, and a separate backup account;
  • Transit Gateway, Direct Connect plus VPN backup, Route 53 Resolver endpoints, interface endpoints, and centralized inspection;
  • payment, carrier, email, fraud, and warehouse providers with IP allowlists; and
  • Step Functions plus Systems Manager Automation recovery orchestration.

Business objectives:

  • MTPD 4 hours;
  • MBCO accepts priority orders at 30 percent peak throughput, without recommendation or email features;
  • RTO 75 minutes to business-accepted MBCO;
  • RPO 5 minutes for accepted orders and payment state, 30 minutes for documents, 24 hours for analytics;
  • no acknowledged order may silently disappear;
  • maximum 15 minutes to declare a disaster after confirmed Regional unavailability; and
  • source must not accept writes after target write authority is granted.

The recovery environment is scaled to 40 percent normal peak. Last full recovery drill was nine months ago at half today's data volume. The latest tabletop was four months ago but did not include business or vendor participants.

3. Write the architecture decision record

The ADR must contain:

Decision and scope

  • one-paragraph decision;
  • business process, customers, minimum service, data classes, and scenarios;
  • primary/recovery Regions and accounts;
  • inclusions, exclusions, assumptions, constraints, and evidence confidence; and
  • accountable business, service, data, security, and recovery owners.

Objectives and measured capability

Show MTPD, MBCO, RTO, dataset-level RPO, SLO, detection/declaration target, backlog/catch-up limits, and regulatory obligations. Compare each with current estimated capability and last measured RTC/RPC. Never relabel an estimate as a drill result.

Alternatives

Compare at least:

  • backup and restore;
  • pilot light;
  • warm standby; and
  • multi-Region active-active.

Assess failure coverage, data semantics, control-plane reliance, operator complexity, RTO/RPO, capacity, security, licensing, provider compatibility, testing burden, one-time and recurring cost, and residual risk. Explain why the selected warm-standby design wins and where it does not protect the business.

Complete architecture

Draw separate diagrams for:

  • account/organization and trust boundaries;
  • identity, emergency access, and approval paths;
  • normal and recovery network/DNS/traffic paths;
  • every authoritative dataset, replication direction, lag, backup, and write authority;
  • application, queue/event, cache, scheduler, and provider dependencies;
  • observability, security, incident command, and evidence paths; and
  • failback and source decommission/isolation.

Decisions and consequences

Record strategy by component, traffic mechanism, data promotion/fencing, backup isolation, corruption response, orchestration boundary, manual approvals, failback method, exercise cadence, costs, known gaps, compensating controls, residual-risk owner, expiry, and revisit triggers.

4. Write an operator-ready recovery runbook

The runbook begins with a one-page emergency summary:

  • purpose and scenario scope;
  • current approved version and change record;
  • incident bridge/out-of-band communication;
  • command authority and alternates;
  • disaster-declaration criteria;
  • current RTO/RPO/MBCO and decision deadlines;
  • stop/hold/fix-forward authority;
  • emergency identity procedure reference;
  • source-fencing invariant; and
  • evidence and execution-log location.

Required phases

  1. Detect, classify, and establish command.
  2. Obtain emergency access and preserve evidence.
  3. Declare or reject disaster; communicate impact.
  4. Freeze changes and fence unsafe writers.
  5. Verify recovery account, Region, identity, KMS, secrets, quotas, network, DNS, security, observability, and provider readiness.
  6. Select a data/recovery point and record effective RPO.
  7. Recover/promote data, files, queues, and workflow state read-only where possible.
  8. Validate data integrity and one-writer preconditions.
  9. Authorize target writes; start applications, consumers, and schedulers in order.
  10. Run technical, security, integration, capacity, and business validation.
  11. Shift canary then approved traffic.
  12. Operate MBCO, reconcile backlog, and enter hypercare.
  13. Decide and execute failback only as a separate controlled recovery.
  14. Close incident, retain evidence, clean temporary access/resources, and update risk.

Every task row

FieldRequirement
ID/phase/versionStable reference to immutable automation or procedure
Planned/maximum durationSupports critical-path and decision-deadline math
Preconditions/invariantsExact evidence required before action
Action and parametersApproved path, typed inputs, account/Region guard
Executor/verifierNamed role plus alternate
Expected outputState/value, not “success”
Validation/evidenceIndependent query/test, timestamp, artifact
IdempotencyPrior-state check and correlation key
Failure decisionRetry, hold, compensate, fix forward, or abort
CompensationData-safe counterpart and owner
ActualsStart/end, result, deviation, decision reference

The runbook must survive interruption. A replacement commander should reconstruct current state from the execution log and evidence, not from memory or chat history.

5. Define roles and decision rights

Include primary and alternate for:

  • incident commander;
  • deputy/timekeeper;
  • technical recovery lead;
  • application, database/data, identity, network/DNS, platform, security, backup, and observability leads;
  • business decision owner and business validator;
  • continuity/risk and legal/privacy representative;
  • communications/customer-support lead;
  • external provider coordinator;
  • scribe/evidence custodian; and
  • tabletop facilitator/controllers/observers, who do not make player decisions.

Create a decision-rights table:

DecisionRecommendsDecidesEvidenceDeadlineAlternate
Declare disasterIncident/technical leadsBusiness/incident authority per policyScope, impact, source statusT+15mNamed deputy
Accept recovery point/data lossData and business leadsData/business risk ownerLag, timeline, reconciliationBefore promotionNamed alternate
Grant target writesData/security/application leadsIncident authoritySource fenced, target validatedBefore writersNamed deputy
Shift trafficTechnical/business validatorsIncident authorityFull gate packetBefore RTO deadlineNamed deputy
Fail backRecovery/data/business leadsChange/business authorityStable source, synced data, testSeparate windowNamed alternate

Authority must remain valid when normal identity, phones, or leaders are unavailable. Test delegation and conflict resolution.

6. Understand tabletop limits and progression

ExerciseWhat occursStrong evidenceCannot prove alone
Document reviewExperts inspect artifactsCompleteness and design issuesHuman response or technology
TabletopPlayers discuss decisions against injectsRoles, decisions, dependencies, document usabilityAPI, duration, capacity, actual recovery
SimulationTools or isolated systems mimic behaviorIntegrations and workflow logicFull production equivalence
Recovery drillReal resources/data restored in isolationMeasured technical RTC/RPC and validationProduction traffic behavior unless included
Game dayTeams act in realistic/prod-like systemTechnology, people, process under conditionsEvery future scenario
Controlled failoverProduction authority/traffic movesHighest realism for tested pathUntested corruption or different failures

The capstone tabletop should create the technical test backlog for claims it cannot prove. Never publish “RTO met” from a discussion. Publish “players produced a plausible 62-minute plan; measured RTO remains unproven.”

7. Prepare the tabletop

Exercise objectives

Test whether players can:

  • detect and classify the event;
  • establish command without normal federation;
  • use the runbook and evidence rather than intuition;
  • decide disaster within 15 minutes;
  • preserve forensic evidence while meeting recovery objectives;
  • select a safe recovery point and keep one writer;
  • achieve MBCO despite optional-provider failure;
  • communicate uncertainty and customer impact;
  • respond to failed traffic change and unavailable experts; and
  • plan data-safe failback.

Rules of play

  • Exercise time advances only when the facilitator announces it.
  • No real system actions or vendor/customer messages occur.
  • Players can request evidence; controllers return only prepared artifacts or “not available.”
  • Assumptions are written and scored as evidence gaps.
  • An inject cannot be wished away by saying a specialist handles it.
  • Controllers do not rescue the team with the intended answer.
  • Safety phrase STOP EXERCISE pauses the simulation for real-world risk or distress.
  • The discussion is blameless; artifacts and systems are evaluated.

Prerequisites

Approve scope, objectives, date/time/time zone, players, observers, confidentiality, no-action boundary, safety process, facilitator script, injects, evidence cards, scoring rubric, baseline architecture, incident log, communication templates, and retrospective schedule. Verify participants can access sanitized exercise documents.

8. Build facilitator injects

Each inject card contains:

  • inject ID and simulated time;
  • prerequisite or player action that releases it;
  • information given to players;
  • private purpose and expected decisions;
  • evidence available on request;
  • incorrect assumptions to challenge;
  • escalation if players stall;
  • clock impact; and
  • scoring notes.

Do not script one correct dialogue. Let decisions change later injects. If players choose not to recover, test how they manage business continuity and risk.

9. Run the blind scenario

Release these injects progressively:

Inject 0, T+00: initial event

Customers in several networks cannot complete checkout in the primary Region. ALB 5xx and database connection failures rise. AWS Health information is inconclusive. The normal incident channel is available.

Test detection, classification, commander selection, baseline impact, evidence requests, and declaration clock.

Inject 1, T+05: normal identity fails

The external identity provider stops issuing sessions. Existing sessions expire in 20 minutes. One emergency-role owner is unreachable.

Test independent access, alternate authority, credential custody, CloudTrail, and out-of-band communication.

Inject 2, T+12: source status uncertain

Some primary API hosts answer, but database write status cannot be confirmed. A queued order may have been acknowledged after the last visible database commit.

Test source fencing, transaction ledger, duplicate/loss handling, and whether players avoid creating two writers.

Inject 3, T+20: stale recovery data

The cross-Region database reports 11 minutes lag, breaching the 5-minute RPO. An isolated backup point is 42 minutes old. Business impact reaches the hospital-priority threshold in 55 minutes.

Test data-loss decision rights, alternate reconstruction/replay, MBCO, and honest RPO reporting.

Inject 4, T+28: possible ransomware

Security reports privileged configuration changes began 36 hours earlier; malware status of recent recovery points is unknown. The newest immutable backup predates the changes by 39 hours.

Test containment, evidence preservation, clean-room branch, point selection, credential rotation, and whether speed improperly overrides safety.

The facilitator chooses either ransomware branch or infrastructure-failure branch for the remainder based on exercise goals. Players must explain how RTO changes.

Inject 5, T+37: recovery capacity and provider gap

Recovery ECS capacity supports 40 percent peak, but payment and carrier providers have not allowlisted recovery egress. The provider coordinator estimates 30 minutes and needs an approved caller.

Test minimum-service design, external dependency, capacity, authorization, and communications.

Inject 6, T+45: data promotion timed out

The orchestration task timed out. The API response was lost, but monitoring suggests promotion might have succeeded. A workflow redrive is available.

Test inspect-before-retry, idempotency, current-state evidence, write authority, and redrive hazard.

Inject 7, T+52: DNS/traffic error

A responder changed the Route 53 record early. Some resolvers send users to a target that remains read-only. Cached clients still use the primary.

Test stop/compensation, TTL/cache reasoning, communication, source/target safety, and validation gates.

Inject 8, T+61: business validator unavailable

The primary business approver loses connectivity. The alternate discovers that synthetic test orders send real email and could create carrier jobs.

Test alternate authority, safe validation data, side-effect controls, and RTO impact.

Inject 9, T+72: failed compensation

The attempt to restore the previous traffic route fails because emergency role permission is missing. Error rate is above the stop threshold, but target priority-order transactions succeed.

Test fix-forward versus rollback, permission escalation, objective decision criteria, and customer messaging.

Inject 10, post-recovery: failback conflict

Primary Region returns. Recovery has accepted 3,200 orders not present in the primary database. Management asks for immediate automatic failback before peak traffic.

Test separate failback change, reverse data synchronization, reconciliation, stable period, risk authority, and refusal of unsafe executive pressure.

10. Keep one authoritative exercise log

For every event record:

FieldMeaning
Simulated/real timeScenario clock and exercise clock
Event/injectWhat became known
Observation vs assumptionEvidence quality
DecisionGO, HOLD, DECLARE, FIX FORWARD, ROLLBACK, DEGRADE, or ESCALATE
Decider and authorityWho made it and why authorized
AlternativesOptions considered and rejected
EvidenceExact card, runbook section, metric, or missing item
Expected result/deadlineWhat happens next and by when
Actual simulated resultController response
Gap/actionDefect to track after exercise

Start clocks for event, detection, command, declaration, recovery start, source fencing, data selected, target write authority, technical ready, business accepted, traffic shifted, MBCO, full service, and exercise close.

11. Score evidence, not performance theater

Score each 0-3:

  • 0: missing or unsafe;
  • 1: stated but owner/evidence/threshold absent;
  • 2: complete on paper but untested or slow;
  • 3: clear, evidence-backed, exercised, and within target.

Domains:

  1. business objectives and MBCO;
  2. declaration and command authority;
  3. emergency identity and communication;
  4. dependency completeness;
  5. source fencing and data integrity;
  6. backup/corruption/ransomware branch;
  7. network/DNS/traffic control;
  8. application/provider/minimum service;
  9. observability/security/evidence;
  10. RTO/RPO measurement;
  11. compensation/fix-forward/failback;
  12. runbook usability and versioning;
  13. business/customer communication;
  14. cost, quota, licensing, and capacity; and
  15. remediation ownership and retest.

Any zero in emergency access, one-writer safety, data-loss authority, security containment, or business acceptance is a capstone fail regardless of total score. A tabletop cannot receive level 3 for measured technical recovery unless prior representative drill evidence is supplied and current.

12. Conduct a blameless hotwash and retrospective

Immediately ask:

  • What surprised us?
  • Where did the team wait or debate?
  • Which evidence was unavailable, stale, or contradictory?
  • Which runbook step was ambiguous or unsafe?
  • Which dependency, person, provider, quota, key, or credential was hidden?
  • Where did decisions exceed their deadline?
  • Which automation could repeat or leave partial state?
  • What would real customers have experienced?
  • What claim remains unproven by a tabletop?

Within the agreed period, publish an after-action report containing objectives, scenario, participants, decisions, measured exercise times, evidence, strengths, findings, unsafe near misses, objective gaps, tabletop limitations, cost impact, and residual risk. Do not edit the original log to make the exercise look cleaner.

13. Convert findings into closed improvements

Each action requires ID, finding/evidence, risk, corrective change, owner, funding, due date, interim control, acceptance evidence, retest type, and closure approver. Categories include architecture, automation, runbook, access, observability, data, vendor, capacity/quota, communications, training, and business policy.

Use closure states:

  • open;
  • funded/planned;
  • implemented but untested;
  • technically retested;
  • business accepted; or
  • risk accepted until an expiry date.

“Documentation updated” does not close a failed recovery mechanism. A runbook defect can close through peer walkthrough; a database promotion defect needs a technical drill; a cross-team decision gap needs another tabletop; full RTO/RPO claims require an integrated exercise.

Schedule the next exercise before closing the current one. Rotate scenarios, recovery points, absent people, workload peaks, vendors, and facilitators while keeping core recovery paths few and repeatable.

14. Read-only supporting evidence

Use supplied evidence by default. If authorized, these inventory summaries may support the ADR but do not prove recovery:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text

aws backup list-protected-resources \
  --query 'Results[].{Type:ResourceType,LastBackup:LastBackupTime}'

aws drs describe-source-servers \
  --query 'items[].{Lifecycle:lifeCycle.state,Replication:dataReplicationInfo.dataReplicationState,Lag:dataReplicationInfo.lagDuration}'

aws route53 list-health-checks \
  --query 'HealthChecks[].{Id:Id,Type:HealthCheckConfig.Type,Disabled:HealthCheckConfig.Disabled}'

aws fis list-experiments \
  --query 'experiments[].{Id:id,Status:state.status,Created:creationTime}'

Redact identifiers. Link each result to date, account, Region, owner, and interpretation. A green health check or backup timestamp cannot replace the capstone's business and dependency evidence.

15. Cost and governance

The ADR must show steady DR cost and exercise/recovery cost: secondary compute and database, replication, backup/copies/locks, storage, data transfer, network connectivity, DNS/accelerator, keys/secrets/certificates, observability/security, support, licenses, quotas/reservations, orchestration, vendor retainers, staff/on-call/training, drills, clean room, failback, and backlog recovery.

Cost must correspond to capability. A 40-percent warm standby needs measured scale-up and quota evidence. An unstaffed overnight response changes RTO. A provider without a recovery SLA is not made reliable by internal automation.

Governance approves BIA/objectives, architecture, residual risk, runbook versions, exercise cadence, evidence expiry, access review, vendor review, capacity/load tests, data restoration, and action closure. Review after every material architecture/provider/team change and incident, not only annually.

Diagnose a misleading capstone

SymptomHidden failureRequired correction
Players know all injectsDiscussion rehearses answers rather than decisionsSeparate facilitator deck and blind release
Tabletop claims RTO metNo technology or data actually recoveredReport decision timeline only; schedule drill
Runbook says “DB team promotes”No preconditions, authority, evidence, timeout, or retry safetyAdd complete step contract
Business absentMBCO, data loss, and acceptance lack authorityInclude owner and alternate
DNS is changed firstTarget data/write readiness is unprovenGate traffic after fencing and validation
Ransomware branch uses latest pointReplicated compromise can returnTimeline, clean room, scan/hunt, older candidate
Findings close as “noted”No implementation evidence or retestTrack owner, due date, acceptance, exercise
High score hides data riskCritical unsafe domain averaged awayApply mandatory fail domains

Knowledge check

  1. What can a tabletop prove?

Decision, role, dependency, communication, and document behavior under a simulated scenario.

  1. What can it not prove?

Actual API behavior, capacity, restore duration, data integrity, or measured technical RPO/RTO.

  1. Why are facilitator injects hidden?

Players must respond to new information rather than recite a scripted happy path.

  1. What is the central data invariant?

Only one authorized writer set exists, with an explicit and evidenced transfer of authority.

  1. Why include unavailable personnel?

Recovery must work through alternates and delegated authority, not one expert.

  1. Why is immediate failback unsafe?

Recovery contains new writes and the original environment needs repair, synchronization, validation, and a separate change.

  1. What makes a finding closed?

Implemented change plus evidence from the appropriate retest and accountable acceptance.

  1. What is the final capstone standard?

A traceable ADR and usable runbook that survive blind injects, expose unproven claims, and produce funded retests.

Lesson acceptance

P17 passes only when the package includes:

  • approved objectives, scope, evidence confidence, and complete ADR alternatives;
  • full account/identity/network/data/application/security/provider/failback architecture;
  • an operator-ready phased runbook with task contracts and deadlines;
  • primary/alternate roles and explicit decision rights;
  • facilitator/player separation and a controlled exercise plan;
  • all 11 injects with branching evidence and scoring;
  • one authoritative execution log and every required clock;
  • honest tabletop limitations and measured-evidence gaps;
  • mandatory safety passes for emergency access, one writer, data-loss authority, containment, and business acceptance;
  • after-action report and blameless findings;
  • funded owners, due dates, interim controls, acceptance evidence, and retests; and
  • a scheduled progression from tabletop to representative technical recovery.

Official sources

Advertisement