Lesson 298 · AWS Learning Path

AWS 298: Application, database, identity, network, and DNS recovery orchestration

· Published · 15 min read

Labelled process diagram for AWS 298: Declared disaster and selected recovery point to Ordered dependency recovery to Application and data validation gate to Traffic switch, monitoring, or safe rollback, with...

Why this lesson matters

Recovery procedures fail when technically correct steps execute in the wrong order. Starting an application before identity, secrets, routes, data, or queues are ready produces secondary failures. Moving DNS before target data is authoritative can lose or duplicate transactions. Blindly retrying a database promotion or traffic change can make recovery worse.

Orchestration coordinates people, automated actions, dependencies, evidence, timeouts, approvals, and compensation as one durable state machine. It does not mean putting every API call into one large script. Deterministic, reversible steps are good automation candidates. Disaster declaration, data-loss acceptance, write-authority transfer, destructive cleanup, and customer traffic often need an explicit accountable decision.

This lesson builds a recovery control plane that can stop safely, resume from evidence, and explain what happened. AWS Step Functions and Systems Manager Automation are useful implementation options, but the design begins with step contracts and failure semantics.

What you will be able to do

By the end, you can:

  • distinguish runbook, automation, orchestration, and incident command;
  • model recovery as states with prerequisites, invariants, outputs, and evidence;
  • bootstrap recovery when normal identity, network, and observability are unavailable;
  • sequence identity, KMS, secrets, network, DNS, data, queues, cache, application, and external providers;
  • design idempotent actions and reject unsafe retries;
  • use Step Functions branches, waits, callbacks, retries, catches, timeouts, and redrive safely;
  • use Systems Manager Automation as bounded, versioned operational runbooks;
  • place manual approvals at irreversible or business-risk gates;
  • design compensation, fix-forward, and data-safe failback; and
  • produce a five-tier tabletop-tested state machine with complete evidence.

Before you start

  • This is a no-create lesson. Do not execute state machines, Automations, restores, promotions, failovers, route changes, or DNS updates.
  • Use supplied fictional identifiers. Recovery workflow definitions reveal sensitive accounts, Regions, endpoints, identities, providers, and security controls.
  • Never embed passwords, access keys, task tokens, private endpoints, or full resource ARNs in coursework or workflow definitions.
  • Verify current API behavior, quotas, workflow duration, integration support, and recovery-service semantics before a real implementation.
  • A successful control-plane execution is not proof that the business service recovered.

1. Separate four responsibilities

Artifact/capabilityResponsibility
Human runbookExplain decision context, authority, communication, manual alternatives, and interpretation
Automation runbookExecute one bounded, tested operational procedure with defined inputs/outputs
OrchestratorMaintain overall state, dependencies, branching, parallelism, waits, evidence, and failure handling
Incident commandDeclare the event, own risk decisions, approve material deviations, and communicate impact

Do not force incident judgment into code merely to claim full automation. Do not make operators manually coordinate 100 deterministic API calls when automation can reduce time and transcription errors.

A good orchestrator calls small versioned workers. A database-promotion worker, network-readiness worker, or validation worker should be testable separately. The parent workflow owns order and policy; the child owns one domain action.

2. Protect the recovery control plane itself

The orchestrator must work during the failure it manages. Identify dependencies on:

  • the failed Region and its service endpoints;
  • normal identity provider, MFA, email, chat, and corporate DNS;
  • primary network, Direct Connect, VPN, NAT, or egress;
  • source Git repository, pipeline, artifact registry, and package sources;
  • KMS keys, secrets, certificates, and parameter stores;
  • CloudWatch/EventBridge/SNS and log destinations;
  • AWS Organizations management or delegated administration;
  • quotas and service control policies; and
  • one operator's laptop or knowledge.

Preposition versioned workflow definitions, runbooks, known-clean artifacts, emergency identities, contact lists, and configuration in an approved recovery account/Region. Provide out-of-band communication and evidence capture. Keep emergency access narrow, MFA-protected, monitored, regularly tested, and unavailable to routine sessions.

Avoid a bootstrap loop: do not require the failed identity service to obtain credentials that recover identity, or the failed DNS path to resolve the endpoint that repairs DNS. Document a minimal root-of-recovery and how it is secured.

3. Define a contract for every step

Use this schema before choosing a tool:

FieldRequired meaning
Step ID/versionStable reference and immutable implementation version
ObjectiveOne observable result, not a vague activity
PreconditionsEvidence that must already be true
InvariantsConditions that must remain true, such as one write authority
InputsTyped values, allowed scope, source, validation, and sensitivity
ActionExact API, Automation, function, or human procedure
Idempotency key/checkHow repeated execution detects prior success
Timeout/heartbeatMaximum wait and liveness signal
Retry policyOnly retryable errors, delay, backoff, jitter, and maximum attempts
OutputTyped state consumed by later steps
ValidationIndependent query/test and acceptance threshold
EvidenceTimestamp, caller, request/execution ID, before/after, logs
Failure pathRetry, compensate, hold, fix forward, manual intervention, or abort
Owner/approverExecutor, verifier, and decision authority

Example objective: “Recovery database endpoint accepts read-only validation and reports the selected recovery point.” Weak objective: “Check DB.”

4. Model recovery as explicit states

A practical top-level state machine is:

READY
 -> INCIDENT_SUSPECTED
 -> DISASTER_DECLARED
 -> SOURCE_FENCED
 -> RECOVERY_FOUNDATION_READY
 -> DATA_RECOVERED_READ_ONLY
 -> DATA_VALIDATED
 -> TARGET_WRITE_AUTHORIZED
 -> APPLICATION_READY
 -> BUSINESS_VALIDATED
 -> TRAFFIC_SHIFTED
 -> HYPERCARE
 -> RECOVERY_ACCEPTED
 -> FAILBACK_PREPARING
 -> NORMAL_OPERATION_RESTORED

Add terminal states such as ABORTED_SAFE, MANUAL_HOLD, FIX_FORWARD_REQUIRED, and COMPENSATION_FAILED. Every transition needs an event, authority, evidence, and time limit.

Persist a correlation ID across Step Functions execution, Automation execution, API calls, logs, change record, incident, data validation, traffic decision, and business acceptance. Persist checkpoints outside ephemeral worker memory so a replacement operator can determine the last proven state.

5. Determine the dependency order

Recovery order is scenario-specific, but examine these layers:

Command, identity, and authorization

Establish incident authority, out-of-band communication, emergency credentials, account/Region access, least-privilege roles, break-glass logging, and approval principals. Normal workforce federation can recover later if emergency identity safely bootstraps the process.

Keys, secrets, certificates, and configuration

Verify KMS keys are enabled and policies/grants work in the recovery context. Replicate or recover secrets without copying compromised values blindly. Issue or activate certificates and configuration with known-clean provenance. Test rotation and revocation paths.

Network and DNS foundations

Validate VPC/subnets, IP capacity, routes, Transit Gateway, VPN/Direct Connect alternatives, endpoints, NAT/egress, security groups, NACLs, inspection, resolver rules, hosted zones, and time synchronization. Prove forward and return flows. Do not move client DNS yet.

Authoritative data

Select and validate recovery points, restore or promote databases, object/file stores, and event logs, reconcile consistency, and keep the target read-only until one write authority is approved. Record effective RPO and any loss/duplication boundary.

Messaging, workflows, and caches

Decide whether queues/streams are recovered, replayed, drained, or recreated. Preserve ordering, offsets, visibility, dead-letter messages, idempotency, and poison-message controls. Rebuild caches from authoritative data unless cache state has explicit durable meaning. Keep schedulers and consumers disabled until dependencies and write ownership are ready.

Applications and external providers

Deploy known-clean immutable artifacts, attach workload identities, retrieve secrets, configure endpoints, start in dependency order, and test internal paths. Verify payment, identity, carrier, bank, email, license, allowlist, certificate, quota, and support dependencies. A vendor phone call is a workflow task with an owner and timeout.

Observability and security

Telemetry should be ready before production traffic. Validate alarms, logs, traces, audit trails, dashboards, on-call routes, security findings, backup, and incident evidence. Do not wait until the end to discover that recovery is invisible.

Traffic and business acceptance

Run synthetic and business validation, authorize writes, then shift controlled traffic through Route 53, Global Accelerator, CloudFront, load balancers, or other approved control. Observe cache and existing-connection behavior. Increase only after thresholds pass.

6. Design for idempotency before retry

An idempotent step can receive the same intent more than once without producing an incorrect additional effect. Use an operation/correlation ID and “inspect before act” pattern:

read current state
if desired state is already proven:
    return prior evidence
if state conflicts with expected precondition:
    stop for investigation
perform one conditional change
record resulting resource/version/request ID
independently validate

Examples:

  • safe: ensure an existing route has the exact approved target using a conditional comparison;
  • conditionally safe: start an already stopped instance after confirming ownership/state;
  • unsafe blind retry: promote a different database, rotate a secret again, submit a payment, replay a batch, or shift traffic repeatedly.

Classify errors as transient, throttling, dependency-not-ready, validation-failed, authorization, conflict, permanent unsupported state, or unknown. Retry only classes with documented safe behavior. Use bounded exponential backoff and jitter, but respect the RTO. A retry storm during recovery can consume the remaining control plane or database.

7. Use compensation, not imaginary rollback

Many recovery actions cannot be rolled back to their exact prior state. Use three terms:

  • rollback: restore the previous state when it remains authoritative and safe;
  • compensation: perform a new action that semantically counteracts an earlier action;
  • fix forward: remain on the new state and repair it because reversal is riskier.

For every mutating step, define its compensation before approval. Examples include restoring a previous route, disabling a newly started consumer, returning a traffic dial, or revoking temporary access. Database promotion after accepted writes needs data replication/replay or fail-forward, not merely a DNS reversal.

Run compensations in reverse dependency order only where dependencies support it. Compensation can fail; preserve a COMPENSATION_FAILED state, stop automatic progression, page the right owner, and protect evidence.

8. Use Step Functions as a durable coordinator

AWS Step Functions can express tasks, choices, waits, parallel branches, maps, service integrations, nested workflows, retries, catches, timeouts, and callback waits. Standard Workflows are commonly suited to long-running, auditable recovery coordination; verify current duration, execution, pricing, and delivery semantics against the design.

Retry and catch

Task, Parallel, and Map states can define ordered retriers. Catchers route supported failures to explicit handling. Never use States.ALL with generic retries for every error. Match known transient errors; route validation, authorization, data-integrity, and unknown failures to hold or incident decisions.

Redriving a failed execution can reset retry counts for rerun states. That can repeat side effects unless workers check durable idempotency records. Before redrive, inspect execution history, current infrastructure/data state, and which actions completed outside the workflow.

Parallel work and joins

Recover independent branches, such as observability and application fleet preparation, in parallel to reduce RTO. Join only after every mandatory branch provides its expected evidence. Define whether an optional branch can fail with accepted degradation.

Parallelism is bounded by API quotas, restore throughput, IP space, people, vendor capacity, and downstream dependencies. More concurrency can increase throttling and ambiguity.

Human callback

Step Functions callback patterns can pause a task until a holder returns a task token. Protect the token like a credential: send it only through approved channels, bind the decision to execution/correlation ID, authenticate and authorize the approver, set timeout/heartbeat, log decision/reason, and reject stale or reused callbacks.

Human approval is appropriate for disaster declaration, selected recovery point/data-loss acceptance, write-authority transfer, traffic move, material exception, destructive cleanup, and failback. Approval must show current evidence, not merely an “Approve” button.

Execution and evidence

Log inputs after redaction, state transitions, API request IDs, output hashes, alarms, approvals, and validation links. Encrypt logs and execution data, restrict access, set retention, and avoid secrets. Define what happens if Step Functions history or its Region is unavailable.

9. Use Systems Manager Automation as bounded workers

Systems Manager Automation runbooks define ordered mainSteps with actions, typed parameters, outputs, timeoutSeconds, maxAttempts, onFailure, onCancel, criticality, and branching. Pin approved document versions; do not run mutable “latest” during disaster unless that policy is deliberate.

Good Automation workers include:

  • verify and prepare a known subnet/security group;
  • validate EC2/agent/volume state;
  • execute a tested command against managed nodes;
  • deploy a known stack with controlled parameters;
  • perform a health query and return structured output; or
  • clean temporary recovery resources after approval.

The aws:approve action can pause for authenticated approval, but it does not support multi-account and multi-Region Automations in the documented model. Therefore do not assume one approval step governs a distributed recovery. Use a central approval/control pattern and per-account workers, or another reviewed mechanism.

Automation cancellation has bounded cleanup semantics, and not every action supports an onCancel jump. Design direct recovery and verify partial effects rather than assuming cancel equals rollback.

10. Put data authority ahead of DNS

Use an explicit writer-fencing protocol:

  1. identify current source of truth and last accepted sequence/transaction;
  2. stop or isolate source writers and schedulers;
  3. drain or capture in-flight work;
  4. reach the accepted replication/restore point;
  5. validate schema, counts, control totals, sequence values, files, queues, and business samples;
  6. record effective RPO and exceptions;
  7. grant target write authority exactly once; and
  8. enable target writers/consumers under observation.

Only then make the target eligible for client traffic. DNS and accelerator health can report endpoints healthy before the database is authoritative. Keep readiness separate from liveness and include data-authority state in the orchestration gate.

If the old Region returns during recovery, automation must not reactivate it automatically. Fencing must survive process restart, delayed messages, scheduler clocks, and human access.

11. Validate before and after traffic

Create a validation matrix for identity, key/secret access, network flows, DNS/TLS, database consistency, object/file availability, queue/stream position, cache behavior, application functions, external integrations, performance/capacity, observability/security, backup, and critical business transaction.

Each test records input, expected result, actual result, tolerance, timestamp, executor, verifier, and artifact. Prevent synthetic tests from charging cards, shipping goods, emailing customers, or polluting financial records.

Before traffic, run internal and synthetic tests. During canary traffic, compare target and source baselines, errors, latency, saturation, data correctness, retries, and business success. After full shift, retain a decision window and safe return strategy. Existing DNS caches and long-lived connections mean traffic distribution will not change instantly.

12. Five-tier orchestration workshop

Use fictional Meridian Reservations:

  • public API behind Route 53 and ALBs in two Regions;
  • normal identity uses external federation; emergency IAM roles exist in recovery;
  • Aurora-compatible primary and cross-Region secondary with asynchronous lag;
  • S3 documents, SQS work queues, ElastiCache, Step Functions bookings workflow, and ECS application/worker services;
  • KMS, Secrets Manager, ACM certificates, centralized logs/security account, Transit Gateway, and interface endpoints;
  • payment and email providers require Regional source-IP allowlists;
  • RTO 60 minutes, RPO 5 minutes, minimum service accepts but does not email reservations; and
  • suspected primary Region impairment, with source status initially unknown.

Produce:

  1. Recovery control-plane dependency and bootstrap-loop threat model.
  2. Top-level states, terminal states, transitions, and authorities.
  3. Step contract for every mutable or validating action.
  4. Identity, KMS, secrets, certificate, network, DNS, data, queue, cache, compute, observability, and provider dependency graph.
  5. Critical path plus bounded parallel branches and resource constraints.
  6. Source-fencing and one-writer protocol.
  7. Data/queue recovery point and reconciliation procedure.
  8. Step Functions state-machine pseudocode with Choice/Parallel/Wait/Retry/Catch/callback behavior.
  9. At least four bounded Systems Manager Automation worker designs.
  10. Idempotency ledger and error taxonomy.
  11. Compensation/fix-forward table for every mutation.
  12. Approval packet for disaster, data point, write authority, traffic, and failback.
  13. Validation matrix and canary/full-traffic thresholds.
  14. Evidence schema, correlation IDs, encryption, retention, and audit ownership.
  15. Failback state machine and cleanup/decommission gates.

Inject these failures: normal federation unavailable, KMS grant missing, subnet out of IPs, cross-Region DB lag exceeds RPO, promotion API times out after succeeding, redrive repeats a task, one SQS consumer starts early, payment allowlist is absent, approver callback expires, Route 53 changes while target is read-only, observability branch fails, and compensation cannot restore a route.

13. Read-only inspection

Only in an authorized account:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text

aws stepfunctions list-state-machines \
  --query 'stateMachines[].{Name:name,Type:type,Created:creationDate}'

aws ssm list-documents \
  --filters Key=DocumentType,Values=Automation Key=Owner,Values=Self \
  --query 'DocumentIdentifiers[].{Name:Name,Version:DocumentVersion,Platform:PlatformTypes}'

aws backup list-restore-jobs \
  --query 'RestoreJobs[].{Status:Status,Created:CreationDate,Completed:CompletionDate,Percent:PercentDone}'

aws route53 list-health-checks \
  --query 'HealthChecks[].{Id:Id,Type:HealthCheckConfig.Type,Disabled:HealthCheckConfig.Disabled}'

Inventory is not execution proof. For one assigned state machine, inspect definition/version/alias, execution role, logging/tracing, encryption, integrations, retry/catch, timeout, execution history, and alarms. For one Automation document, inspect exact version, parameters, assume role, steps, outputs, cancellation, and prior results. Do not start either.

14. Security, cost, quota, and lifecycle

Use separate least-privilege identities for orchestrator, domain workers, approvers, emergency operators, and evidence reviewers. Cross-account workers should trust only approved orchestrator sources and require resource tags/conditions where supported. Restrict StartExecution, redrive, stop execution, Automation start/cancel, pass role, DNS, KMS, secret, database promotion, restore, and traffic APIs independently.

Protect workflow definitions as code with review, signing/provenance, testing, immutable version references, and controlled deployment to recovery Regions/accounts. Detect configuration drift. Never let a compromised production pipeline overwrite the only recovery workflow.

Cost includes Step Functions transitions or execution duration by workflow type, Lambda/Automation/API calls, restored resources, parallel recovery capacity, data transfer, DNS/accelerator, logs/traces/metrics, secrets/KMS, SNS/callback infrastructure, test traffic, support, drills, and idle recovery control plane. Logging full payloads can add cost and expose data.

Check quotas for workflow executions/history, state transitions, Lambda concurrency, Automation concurrency, API rates, backup/restore, DRS, EC2/EBS/IPs, databases, Route 53, and external providers. Orchestration should throttle itself and preserve emergency reserve.

Version and retire workflows safely. Keep old versions while active executions or rollback plans depend on them. Revoke stale callbacks/roles, remove test resources, reconcile temporary routes/DNS/secrets, archive evidence, and verify the service returns to a known normal state.

Diagnose orchestration failure

ObservationLikely design errorCorrection
Workflow cannot start during outageIt depends on failed identity, Region, network, key, or pipelinePreposition and test an independent recovery control plane
Retry promotes twiceMutation lacks idempotency/state inspectionRecord intent/result and verify before repeating
DNS sends users to read-only targetTraffic is not gated by data authorityRequire one-writer and business validation before eligibility
Redrive corrupts stateCompleted side effects replayed with reset retry countsInspect execution/current state and make workers idempotent
Parallel recovery is slowerAPI, bandwidth, quota, or team is sharedBound concurrency from measured bottlenecks
Cancelled Automation leaves changesCancel was mistaken for rollbackDefine onCancel where supported plus direct compensation
Human approval arrives too lateNotification/access/timeout path was not testedUse out-of-band notice, backup approver, expiry, escalation
Workflow completes but service failsOutputs measured APIs, not business outcomeAdd end-to-end data and user acceptance gates

Knowledge check

  1. How does orchestration differ from automation?

Automation performs a bounded task; orchestration coordinates many tasks, decisions, dependencies, and evidence.

  1. What is an invariant?

A condition that must remain true through transitions, such as exactly one authoritative writer.

  1. When is a retry safe?

When the error is classified, attempts are bounded, and the step is idempotent or checks durable prior state.

  1. Why can Step Functions redrive be dangerous?

Rerun states can retry again while prior external side effects may already exist.

  1. What does compensation mean?

A new action that semantically counteracts a prior mutation when exact rollback is impossible.

  1. Why is aws:approve not a universal distributed approval mechanism?

Its documented action does not support multi-account and multi-Region Automations.

  1. When should DNS/traffic move?

After data authority, application readiness, critical integrations, and business validation pass.

  1. What proves orchestration success?

Accepted business service within RPO/RTO, complete evidence, controlled side effects, and recoverable/failback state.

Lesson acceptance

You may continue when your submission contains:

  • an independent, secured, testable recovery control plane;
  • explicit states, transitions, invariants, authorities, and terminal failures;
  • contracts for every action with idempotency, timeout, retry, evidence, and compensation;
  • complete dependency order and bounded parallelism;
  • source fencing and one-writer data protocol;
  • safe Step Functions retry/catch/redrive/callback design;
  • versioned bounded Systems Manager Automation workers;
  • manual approvals for irreversible business-risk gates;
  • pretraffic, canary, full-traffic, data, security, and business evidence;
  • failback and failed-compensation handling;
  • least privilege, quota, cost, versioning, and cleanup ownership; and
  • tabletop results for every supplied failure injection.

Official sources

Advertisement