Lesson 296 · AWS Learning Path

AWS 296: Business impact analysis, dependency recovery order, and recovery evidence

· Published · 15 min read

Labelled process diagram for AWS 296: Business process and impact timeline to Dependencies and objective setting to Ordered technical and business recovery to Measured evidence and approved gaps, with decision, proof...

Why this lesson matters

Disaster recovery is not a competition to give every application the smallest RTO. Recovery investment should preserve the business processes whose disruption creates the greatest safety, legal, financial, customer, operational, or reputational impact over time.

A business impact analysis (BIA) begins with business services and obligations, not EC2 instances. It establishes how impact changes as disruption continues, the minimum acceptable service, tolerable data loss, deadlines, manual workarounds, and accountable owners. Architects then map those outcomes to technical dependencies and measured recovery capability.

This lesson closes a common gap: a “Tier 0” application cannot recover in 30 minutes if identity, network, keys, data, external providers, and the people with authority take four hours. A backup record proves protection activity, not restorable business service. The recovery order must follow a dependency graph, shared-resource capacity, and one coherent evidence model.

What you will be able to do

By the end, you can:

  • facilitate a BIA without asking stakeholders to invent arbitrary technical numbers;
  • distinguish maximum tolerable disruption, minimum service, RTO, RPO, SLO, and recovery capability;
  • model nonlinear impact by time, season, customer, geography, and scenario;
  • map business processes to applications, data, infrastructure, people, facilities, and third parties;
  • calculate a dependency-aware critical recovery path;
  • build recovery waves without making every service “highest priority”;
  • reconcile objectives with AWS Resilience Hub estimates and measured drill evidence;
  • expose impossible objectives, hidden single points, resource contention, and manual bottlenecks;
  • define acceptance evidence for technical and business recovery; and
  • produce a funded, owned improvement and retest plan.

Before you start

  • Use the supplied fictional organization. A real BIA can contain confidential revenue, safety, legal, customer, supplier, staffing, and risk information.
  • Business owners set acceptable impact and minimum service. Technology teams explain capabilities, dependencies, costs, and risks; they do not silently assign business tolerance.
  • This is a no-create lesson. Do not change backup, replication, DNS, Resilience Hub, or recovery configurations.
  • Existing-account inspection requires authorization. Redact account IDs, resource names, business values, recovery locations, contacts, and sensitive gaps.
  • Document currency, time zone, peak periods, data date, assumptions, confidence, owner, and review date.

1. Begin with a business process and outcome

An application name is not a business outcome. Start with a verb and result, for example:

  • accept and authorize a customer order;
  • dispatch urgent medical supplies;
  • calculate and submit payroll;
  • receive a regulatory report;
  • allow staff to authenticate and communicate; or
  • settle completed orders with finance.

Define customers, channels, operating hours, geography, products, transaction volume, legal deadlines, safety implications, peak periods, upstream inputs, downstream obligations, and process owner. One process can use many applications; one application can support several processes with different criticality.

Scope the disruption scenarios. Loss of one instance, one Availability Zone, one Region, an AWS account, identity provider, network, operator access, source code, data integrity, supplier, or staff creates different impacts and recovery paths. A BIA that assumes only “application unavailable” can miss the dominant risk.

2. Build an impact-over-time curve

Interview the people who own the outcome, finance, legal/compliance, safety, customer support, operations, and major dependencies. Ask what happens after 15 minutes, 1 hour, 4 hours, 8 hours, 24 hours, 3 days, and one week. Adjust intervals to the service.

Score or quantify each category with traceable evidence:

ImpactEvidence to collect
Safety and welfarePeople exposed, severity, emergency workaround, statutory duty
Legal/regulatoryFiling deadline, breach threshold, notification, contractual penalty
FinancialLost margin, delayed cash, penalties, overtime, spoilage, recovery cost
CustomerUsers affected, abandonment, SLA credit, churn risk, support contacts
OperationalBacklog growth, blocked teams/sites, capacity to catch up, manual work
ReputationPublic visibility, strategic partner impact, evidence rather than guesses
DataLost transactions, reconstruction ability, integrity and privacy impact

Impact is often nonlinear. A payroll outage on an ordinary morning may have little immediate effect, then cross an irreversible bank-submission deadline. A warehouse backlog can be absorbed for two hours, then exceed next-shift capacity. Show these cliffs rather than averaging them away.

Separate cash loss, delayed revenue, risk exposure, and qualitative harm. Do not multiply maximum revenue by downtime unless every sale is truly lost and unrecoverable. Record formulas and confidence ranges.

3. Use recovery terms precisely

Maximum tolerable period of disruption

The maximum tolerable period of disruption (MTPD), sometimes expressed through related organizational terminology such as maximum acceptable outage, is the point beyond which impact becomes unacceptable to the organization. Confirm the exact vocabulary used by the organization's continuity standard.

MTPD is not the technical target. Recovery needs margin before the unacceptable point for detection, decision, restoration, validation, backlog, and uncertainty.

Minimum business continuity objective

The minimum business continuity objective (MBCO) describes the smallest service level that must be restored during disruption. It can be a percentage of transaction volume, specific customer tier, geography, product, channel, or manual service.

Examples include “accept priority hospital orders at 25 percent normal throughput” or “pay all employees but defer reporting.” “Application up” is not an MBCO.

Recovery time objective

RTO is the maximum acceptable delay from interruption until the required service is restored. State the exact endpoint: infrastructure ready, application available, minimum business service accepted, or full service. This course uses accepted minimum business service unless explicitly stated.

RTO must fit inside MTPD with time for backlog and contingency:

detection + declaration + technical recovery + validation
+ business restart + contingency < MTPD

Recovery point objective

RPO is the maximum acceptable amount of data loss expressed as time since the last usable recovery point. Define it for each authoritative dataset and event stream. Orders, audit events, uploads, identity changes, and analytics can have different tolerances.

RPO does not mean replication interval alone. Corruption can replicate, snapshots can fail validation, and a multi-system process can recover to inconsistent times. Define reconstruction, replay, reconciliation, duplicate handling, and the business acceptance of loss.

Recovery capability

AWS guidance distinguishes objectives from tested capability. Use:

  • RTC: measured recovery time capability from representative drills;
  • RPC: measured recovery point capability from usable recovered data.

Some tools provide estimated workload RTO/RPO based on configuration. Label these estimates. A passing estimate is not a measured end-to-end recovery, and a backup success is not RPC.

4. Define minimum service and manual workarounds

A degraded service can reduce recovery cost and time. Specify which functions remain, throughput, user groups, geography, hours, security controls, data captured, maximum duration, support model, and transition back to normal.

Evaluate manual workarounds honestly:

  • How many trained people are available during nights, weekends, or regional disruption?
  • What is secure transaction capacity per hour?
  • How are identity, authorization, privacy, audit, and segregation of duties maintained?
  • Where is data stored, backed up, reconciled, and imported later?
  • How fast does backlog grow and how long does catch-up take?
  • Which errors or duplicates are likely?

A spreadsheet on one person's laptop is not a continuity strategy. A workaround has its own dependencies and recovery limit.

5. Create the service-to-technology map

For each business process map:

business outcome
  -> user/channel/facility
  -> edge/DNS/network
  -> identity and authorization
  -> application/API/job
  -> queue/event/workflow
  -> database/object/file/cache
  -> keys/secrets/certificates
  -> observability/operations/support
  -> external provider and downstream consumer

Record direction, protocol, data, transaction criticality, owner, operating window, fallback, RTO/RPO contribution, and evidence source. Include control-plane dependencies needed to recover, such as organization/account access, IAM federation, infrastructure code, artifact registry, pipeline, service quotas, support access, and recovery-region configuration.

Discovery telemetry shows observed calls, not complete business dependency. Supplement it with interviews, code/configuration, scheduled jobs, contracts, DNS, firewall rules, data lineage, incident history, and drills. Mark each relationship confirmed, inferred, stale, or unknown.

6. Convert the graph into recovery order

Model each recoverable capability as a node and prerequisites as directed edges. If Order API requires identity, secrets, database, and payment connectivity, those prerequisites must be ready or an approved degraded mode must bypass them.

For each node record:

  • preparation and recovery duration from evidence;
  • validation duration;
  • earliest start after disaster declaration;
  • prerequisites;
  • people/team and scarce resources;
  • maximum parallelism;
  • data point and write-authority requirement;
  • recovery and business acceptance owner; and
  • failure/rollback path.

The critical path is the longest dependency chain to the MBCO, not the sum of every component. Independent branches can recover in parallel, but shared people, network bandwidth, restore throughput, API quotas, IP space, licenses, and vendor availability can serialize them.

Example:

emergency access (10m)
  -> recovery network/DNS (20m)
     -> KMS/secrets (10m)
        -> database restore and validation (55m)
           -> application start (15m)
              -> business transaction validation (15m)

critical-path capability = 125 minutes before contingency

If the business RTO is 90 minutes, changing a tier label does not fix the gap. Shorten work, pre-provision a dependency, automate, run branches in parallel, select a stronger recovery pattern, redefine an acceptable MBCO, or obtain explicit risk acceptance.

7. Recover foundations without creating one universal order

Common early capabilities include:

  • incident command and out-of-band communication;
  • emergency identity, credentials, MFA, and authorization;
  • recovery accounts/organization controls and quotas;
  • network, routing, firewalls, endpoints, IP address management, and connectivity;
  • DNS and time synchronization;
  • KMS keys, certificates, secrets, and configuration;
  • artifact/source/IaC repositories and deployment automation;
  • security monitoring, logs, metrics, traces, incident evidence, and backup catalog;
  • authoritative data and messaging; and
  • external providers, facilities, operators, and business validators.

Do not blindly recover every foundational platform before every business service. A self-contained emergency function might recover with a narrower set. Conversely, restoring shared identity first can block everything if its own dependency is unavailable. Produce scenario-specific orders and test bootstrap loops.

Break circular dependencies with prepositioned recovery credentials, offline runbooks, replicated artifacts, controlled static configuration, alternate communications, or a carefully limited emergency mode. Protect and exercise those mechanisms.

8. Build recovery tiers and waves responsibly

Use no more tiers than the organization can fund and operate. A tier must define objective ranges, minimum service, standard patterns, test frequency, evidence age, ownership, and escalation. Do not call every executive-visible application Tier 0.

Within a tier, create dependency-aware waves:

  1. command, access, and recovery-control foundations;
  2. shared technical prerequisites and authoritative data;
  3. minimum business services by impact deadline;
  4. supporting operations and integrations;
  5. deferred analytics, reporting, optimization, and convenience functions; and
  6. full-service catch-up and normal-state reconciliation.

Priorities change by scenario and time. During ransomware, clean identity, forensic containment, and uncontaminated data points precede rapid restoration. During a Regional infrastructure loss, prebuilt standby services can start immediately. During a supplier outage, restoring internal compute may not restore the process.

9. Build an evidence hierarchy

Classify evidence strength:

LevelEvidenceWhat it proves
0Owner statement or design intentRequirement or belief only
1Configuration inventory or tool estimateA capability appears configured
2Component testOne mechanism worked in a bounded context
3Integrated recovery drillDependencies recovered together in a realistic environment
4Business-accepted game dayMBCO, RPO, RTO, people, and process were measured
5Repeated representative evidenceCapability is reproducible across change and time

Every evidence record needs scope, scenario, environment, timestamp, data volume, load, versions, participants, result, exceptions, artifact location, expiry, and approver. Evidence older than a major architecture, team, provider, or data-volume change must be revalidated.

Use these statuses:

  • objective met with current representative evidence;
  • estimated to meet but not drill-proven;
  • objective breached by measured capability;
  • untested or evidence expired;
  • blocked by dependency; or
  • risk accepted until a dated remediation.

Avoid a single green score that hides one breached critical path.

10. Reconcile AWS evidence correctly

AWS Backup inventory can show protected resources, plans, vaults, jobs, and recovery points. It does not prove that a recovery point is uncorrupted, that restore permissions work in a disaster, or that the application is usable.

DRS can show source-server replication and lag. It does not prove multi-server consistency, dependencies, launch settings, or business validation.

Route 53 and Global Accelerator can show traffic-health and steering configuration. They do not prove the target's data authority and capacity.

AWS Resilience Hub policies express target RTO/RPO by disruption type and assessments estimate whether application components meet them from discovered configuration. Recommendations can identify alarms, SOPs, and tests. An assessment does not change the application and is not an actual recovery. Reconcile its scope and estimated results with BIA objectives and measured drills.

Use a matrix:

ServiceBIA RTO/RPOEstimated capabilityLast measured RTC/RPCDependency critical pathGap and owner
Order acceptance90m / 5m75m / 5m128m / 11mIdentity -> network -> DB -> APIBreached; platform owner

The strictest number is not automatically correct. Investigate scope, scenario, endpoint, data definition, and evidence date before reconciling.

11. Read-only AWS evidence

Only in an authorized account, use narrow summaries and redact output:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text

aws backup list-protected-resources \
  --query 'Results[].{Type:ResourceType,LastBackup:LastBackupTime}'

aws backup list-backup-jobs \
  --by-created-after 2026-09-01T00:00:00Z \
  --query 'BackupJobs[].{Type:ResourceType,State:State,Created:CreationDate,Completed:CompletionDate}'

aws drs describe-source-servers \
  --query 'items[].{Lifecycle:lifeCycle.state,Replication:dataReplicationInfo.dataReplicationState,Lag:dataReplicationInfo.lagDuration}'

aws resiliencehub list-apps \
  --query 'appSummaries[].{Name:name,Status:status,AssessmentSchedule:assessmentSchedule}'

Dates are illustrative; use the approved evidence window. Listing many resources may reveal sensitive scope and still proves little. Follow an owned service through exact recovery-point, policy, assessment, drill, and business evidence instead of equating count with readiness.

12. Five-service BIA workshop

Use fictional Northstar Supply Cooperative:

  1. Emergency ordering: hospitals place priority orders 24x7; manual phone mode supports 20 percent volume for 4 hours.
  2. Warehouse dispatch: creates pick/ship instructions; backlog above 12,000 orders misses carrier collection.
  3. Identity and workforce access: supports ordering, warehouse, support, and recovery operators; emergency identities are limited.
  4. Payroll: weekly bank file due Friday 14:00; outage impact is low until the submission deadline.
  5. Analytics/reporting: daily dashboards and monthly regulator extract; operational systems do not require it for transactions.

Shared capabilities: Route 53, Transit Gateway/DX/VPN, IAM Identity Center and directory, KMS, Secrets Manager, certificate services, CI/CD and artifact registry, PostgreSQL orders, S3 documents, SQS events, monitoring/security account, external payment, carrier, bank, email/SMS, and an on-premises warehouse controller.

Supplied evidence includes backup jobs, DRS lag, a Resilience Hub estimate, a six-month-old game day, incident timelines, transaction peaks, staffing, contract deadlines, and vendor recovery claims of mixed confidence.

Produce:

  1. Process charters, owners, customers, operating periods, and scenario scope.
  2. Impact curves across at least seven time bands and all impact categories.
  3. Evidence-backed MTPD and MBCO for each process.
  4. Dataset-specific RPO plus reconstruction/reconciliation rules.
  5. RTO decomposition from detection through business restart and backlog.
  6. Manual-workaround capacity, security, duration, and catch-up analysis.
  7. Service-to-technology and external-provider dependency graph.
  8. Bootstrap-loop and hidden-control-plane analysis.
  9. Node durations, resources, validation, and recovery owners.
  10. Scenario-specific critical paths and parallel recovery schedule.
  11. Tier and wave assignment with a reason no service is overclassified.
  12. Objective/estimate/measured-capability matrix with evidence levels and expiry.
  13. Resource-contention simulation for two simultaneous recoveries.
  14. Gap options showing cost, risk reduction, revised capability, and owner.
  15. Executive decision record plus quarterly evidence-refresh plan.

One expected finding is that payroll can have low immediate impact but a hard deadline, while identity is a dependency for high-impact services yet requires an independent bootstrap path. Analytics should not be recovered before emergency ordering merely because its infrastructure is easier to restore.

13. Fund the gap rather than hiding it

For every objective breach, present options:

  • reduce detection/declaration delay;
  • automate and parallelize proven steps;
  • pre-provision identity, network, data, or capacity;
  • change from backup/restore to pilot light, warm standby, or active-active;
  • simplify dependencies or add degraded operation;
  • improve backup frequency, replication, immutability, or reconciliation;
  • reserve quotas, licenses, people, bandwidth, and vendor support;
  • change the MBCO or objective with accountable business approval; or
  • accept the risk for a defined period with trigger and expiry.

Compare one-time and recurring cost, scenarios covered, residual risk, operational complexity, test cost, implementation lead time, and lock-in. Do not claim an objective is met because a project to meet it was approved.

14. Governance and review cadence

The business process owner approves impact, MTPD, MBCO, RTO, RPO, and residual business risk. Application/data/platform/network/security owners attest dependencies and capability. Continuity and risk teams govern method. Finance validates cost. Executives resolve priority and funding conflicts.

Review at least on a defined periodic cadence and after acquisition, new product/geography, regulatory change, architecture migration, major incident, provider change, significant volume growth, owner change, or failed exercise. Version the BIA and preserve prior approved objectives; do not rewrite history to match current capability.

Track percentage of services with current owner-approved BIA, confirmed dependency map, funded strategy, unexpired drill evidence, objectives met, overdue high-risk gaps, and successful business validation. Counting backups or documents is not enough.

Diagnose a misleading BIA

SymptomCauseCorrection
Every application is Tier 0Political priority replaced impact evidenceStart with processes, time curves, deadlines, and MBCO
RTO equals MTPDNo margin for detection, validation, backlog, or uncertaintyDecompose the timeline and set a safer objective
All data has one RPODataset authority and reconstruction differAssign dataset-specific loss and consistency rules
Recovery order is an application listIdentity, network, keys, data, people, and vendors are hiddenBuild a directed dependency graph and critical path
Backup dashboard is greenProtection activity was mistaken for recoveryRestore, reconcile, integrate, and obtain business acceptance
Resilience estimate says compliantConfiguration estimate was mistaken for measured RTC/RPCRun representative drills and reconcile scope
Parallel plan misses RTOSame operator, bandwidth, quota, or license is double-bookedResource-level schedule and contention test
Manual workaround has no limitCapacity, security, backlog, and fatigue were omittedMeasure throughput and maximum safe duration

Knowledge check

  1. Why begin with a business process instead of an application?

Processes deliver outcomes and often span several applications, people, data stores, and providers.

  1. How does MTPD differ from RTO?

MTPD is the unacceptable-disruption boundary; RTO is an earlier restoration target with needed margin.

  1. What does MBCO define?

The minimum acceptable service level during disruption, including users, functions, volume, and duration.

  1. Why can replication lag differ from effective RPO?

The newest blocks might be inconsistent or corrupt and recovery can require an older usable point.

  1. What are RTC and RPC?

Measured capabilities for recovery time and recovery point, compared with business objectives.

  1. How is recovery order determined?

By prerequisites, scenario, critical path, shared resources, data authority, and business deadlines.

  1. What does a Resilience Hub assessment prove?

An estimated configuration-based comparison to policy, not an actual end-to-end recovery.

  1. Who accepts an objective breach?

The accountable business/risk authority, informed by technical evidence and a dated remediation or acceptance.

Lesson acceptance

You may continue when your submission contains:

  • process-based, scenario-specific impact analysis with owner evidence;
  • nonlinear time curves, MTPD, MBCO, RTO, and dataset-level RPO;
  • secure manual-workaround and backlog analysis;
  • complete technical, human, facility, control-plane, and provider dependencies;
  • scenario-specific directed recovery graph, critical path, parallelism, and contention;
  • defensible tiers and waves;
  • an objective versus estimate versus measured RTC/RPC matrix;
  • dated evidence levels, exceptions, expiry, and approvers;
  • funded options or explicit residual-risk acceptance for every breach; and
  • governance, change triggers, metrics, and recurring business-accepted drills.

Official sources

Advertisement