Lesson 245 · AWS Learning Path

AWS 245: Convert the recovery steps into a tested runbook

· Published · 12 min read

Labelled process diagram for AWS 245: Incident diagnosis to Versioned runbook and typed input to Bounded automated repair to Verification, rollback, and execution evidence, with decision, proof and rejection evidence.

Why this lesson matters

AWS244 ended with a recovered application and evidence explaining three independent faults. A successful command history is not yet automation. Commands copied into a script can repeat a mistake faster, against more resources, under a more powerful identity.

In this lesson you convert one proven recovery - restoring an Application Load Balancer target group's health-check path from /healthz to /ready - into a bounded AWS Systems Manager Automation runbook. The runbook reads before it writes, proves ownership, checks the expected old state, pauses for approval, changes one field, waits, verifies target health, and explicitly compensates if verification fails.

The supplied runbook is design-only. It contains placeholders and must not be executed unchanged. The complete learning path requires no AWS mutation.

Outcomes

By the end, you can:

  • distinguish an incident playbook, human runbook, script, and Systems Manager Automation runbook;
  • explain Automation's control plane, execution identity, document version, steps, inputs, outputs, branching, retries, timeout, cancellation, and status;
  • design guard clauses and idempotent behavior instead of assuming the starting state;
  • explain why rollback is explicit compensation rather than a database transaction;
  • test healthy, failed, unauthorized, stale-state, canceled, and partial-failure cases;
  • bound fleet blast radius with targets, concurrency, and error thresholds;
  • collect execution evidence without leaking account identifiers or secrets;
  • estimate step, script-duration, attachment, logging, notification, and changed-resource costs.

From manual recovery to controlled automation

ArtifactPrimary questionTypical content
PlaybookHow does the team coordinate this class of incident?roles, severity, communication, escalation, decision points
RunbookHow is one repeatable operation performed safely?prerequisites, commands/actions, verification, rollback
ScriptHow are instructions executed?code with inputs and exit behavior
SSM Automation runbookHow does AWS orchestrate controlled steps and retain execution state?typed parameters, IAM context, actions, outputs, branches, approvals

Do not automate a recovery merely because it once appeared to work. First prove the causal chain, owner source, safe scope, acceptable rollback, and success signal. Automation should reduce toil and variation; it must not hide uncertainty.

A short history

Operations began with handwritten checklists and shell scripts. Configuration-management systems made desired state repeatable, while orchestration engines added workflow state and branching. AWS Systems Manager Documents provide machine-readable operational documents. Command documents commonly use schema 2.2 or later; Automation runbooks use schema 0.3. Automation adds AWS API actions, scripts, waits, assertions, approvals, branching, outputs, execution history, and fleet rate controls.

This evolution does not remove human responsibility. An approval after an unbounded destructive step is theater. Put scope checks and evidence before the approval, and the mutation after it.

The execution model

operator or event
      |
      | ssm:StartAutomationExecution + permitted document version
      v
Automation execution ---- assumes approved role ----> AWS service APIs
      |                              |
      |                              +--> exact resource permissions
      +--> step state and outputs
      +--> approval / wait / branch
      +--> verification or compensation
      +--> CloudTrail + optional EventBridge/CloudWatch evidence

An execution is not a shell session. Systems Manager advances named steps and records each step's status. A runbook document contains:

  • description and schemaVersion: '0.3';
  • optional top-level assumeRole;
  • typed parameters, defaults, descriptions, and validation patterns;
  • mainSteps, each with a unique name, an action, action-specific inputs, and shared controls;
  • step outputs selected from an action result;
  • optional top-level outputs exposed by the execution.

Reference an earlier output as {{ stepName.outputName }}. Types are static: do not pass a MapList where a String is expected.

Common actions

ActionUseImportant boundary
aws:executeAwsApicall one AWS APIinspect API names, pagination, throttling, and return shape
aws:executeScriptrun bounded Python or PowerShellAPI calls from a customer runbook require an Automation assume role
aws:runCommandrun an SSM Command document on managed nodesnode readiness and command permissions remain dependencies
aws:waitForAwsResourcePropertypoll until a property reaches an allowed valuealways choose a finite timeout
aws:assertAwsResourcePropertyfail if a resource property differsuseful for guard clauses, not a lock against later change
aws:branchselect the next stepinclude a safe default path
aws:approvepause for designated principalstimeout defaults to 7 days, maximum 30 days; unsupported in multi-account/Region automation
aws:invokeLambdaFunctioninvoke existing Lambda logicadds function IAM, retries, logs, and cost dependencies
aws:createStackcreate a CloudFormation stackstack rollback semantics are separate from Automation semantics

Every action supports shared controls such as timeoutSeconds and maxAttempts. onFailure can abort, continue, or jump to step:name. onCancel can abort or jump to a supported step, but the cancellation workflow has a two-minute maximum and cannot jump to certain long-running or approval actions. isCritical affects the final execution status. nextStep changes normal flow; isEnd terminates it. Defaults are behavior, so review them rather than relying on memory.

Identity and least privilege

Two identities matter:

  1. The starter needs permission to start the chosen document/version and pass the exact Automation role when one is supplied.
  2. The Automation role needs a trust relationship for Systems Manager and only the API/resource permissions used by the steps.

For many non-script runbooks, omitting assumeRole causes Automation to use the starter's IAM context. That makes behavior vary by operator and is unsuitable for a controlled production process. A customer runbook whose aws:executeScript calls AWS APIs requires an IAM service role. Passing a role requires iam:PassRole; restrict both the role resource and iam:PassedToService where applicable. Scope the trust policy with source-account/source-ARN conditions to reduce confused-deputy risk.

Do not solve access errors by attaching AdministratorAccess, AmazonSSMFullAccess, or a broad Resource: '*' mutation policy. Derive permissions from the exact read, write, approval, logging, and rollback API calls. Remember that rollback needs permission too.

Safe automation invariants

Write invariants before steps. This lab uses these:

  1. The account and Region are the approved sandbox.
  2. The target group ARN is explicit; discovery cannot select a look-alike resource.
  3. Its CourseOwner tag equals the approved value.
  4. The current path is either the expected broken value or the already-correct value.
  5. Any third value means the world changed; stop without mutation.
  6. The desired path is /ready, derived from application-owner evidence.
  7. Only the health-check path may change.
  8. Mutation requires a named change ID and human approval.
  9. Success means configuration convergence and healthy targets over an agreed observation window.
  10. A failed post-change verification attempts bounded compensation and reports failure even if compensation succeeds.

These are guard clauses, not decorative comments. A precheck followed by a delayed write still has a time-of-check/time-of-use race. Reduce that window, serialize execution for the same target, re-read immediately before mutation when risk requires it, and stop on unexpected state.

Idempotency, retries, and eventual consistency

An idempotent operation can be retried without creating additional unintended change. This runbook treats /ready as already converged and exits successfully without approval or mutation. It permits /healthz only when ownership and change inputs match. Any other current path is rejected.

Retries belong around transient operations, not business ambiguity. Retrying AccessDenied, an ownership mismatch, or a failed invariant does not make it safe. AWS APIs may be eventually consistent and may throttle. Use finite retries, waits, and total execution deadlines; distinguish “API accepted the change” from “the application is healthy.”

Approval is a control, not proof

aws:approve pauses until the minimum approval count is met, a denial arrives, or the step times out. Approvers can be IAM users or role principals; SNS notification is optional, and if used the topic name must begin with Automation. The approval message should contain resource, old state, desired state, change ID, evidence link, rollback, and expiry - never a secret.

Approval does not prove the approver read the evidence, prevent the resource changing afterward, or provide separation of duties automatically. Enforce who may start, pass the role, approve, and alter the document. For governed change workflows, evaluate Systems Manager Change Manager rather than treating one aws:approve step as a complete change-management system.

Rollback is compensation, not a transaction

Automation does not atomically undo earlier AWS API calls when a later step fails. onFailure merely controls workflow direction. Design a compensation step for every mutation and record what original value is safe to restore.

Compensation can also fail because permissions changed, an API is unavailable, or another actor changed the resource. Therefore:

  • capture and validate the exact pre-change state;
  • make the rollback step bounded and observable;
  • do not report the overall recovery as successful merely because rollback ran;
  • escalate with current state and execution ID when compensation fails;
  • prefer roll-forward when rollback would lose data or violate compatibility.

The supplied design restores /healthz only because the guard proved that exact starting state. It never guesses a rollback value.

Fleet and multi-account blast radius

Automation can target resources by IDs, tags, or resource groups. At fleet scale use:

  • MaxConcurrency to limit simultaneous resource executions;
  • MaxErrors to stop dispatching to new targets after the threshold;
  • location-level concurrency/error controls for account–Region pairs;
  • a canary target, then progressively larger waves;
  • explicit account, OU, Region, and exclusion lists.

An error threshold does not cancel executions already running. If the maximum acceptable failures must never exceed one, concurrency must be one and the operation must be safely serial. The aws:approve action is not supported for multi-account/Region executions, so place governance at an appropriate outer workflow boundary.

Practical lab: supplied-evidence track

Download the AWS245 recovery-runbook pack. It contains:

  • RECOVERY_RUNBOOK_WORKBOOK.md - design and evidence worksheet;
  • orders-target-health-recovery.DESIGN-ONLY.yaml - annotated schema 0.3 runbook;
  • SUPPLIED_RUNBOOK_TEST_CASES.md - expected results for safe and unsafe states;
  • validate_runbook.sh - local structural checks that make no AWS calls.

Run the local checks:

tar -xzf nw-p12-recovery-runbook.tar.gz
cd p12-recovery-runbook
sh validate_runbook.sh orders-target-health-recovery.DESIGN-ONLY.yaml

Then complete the workbook:

  1. Map each AWS244 manual recovery statement to a read, decision, approval, mutation, wait, verification, compensation, or evidence step.
  2. Identify every AWS API action and derive a least-privilege role outline.
  3. Trace all branches. Prove that Reject and AlreadyDesired cannot reach mutation.
  4. Execute the supplied test cases on paper. Include the exact last step and final status.
  5. Explain the time-of-check/time-of-use race and propose serialization.
  6. Calculate the bill for one successful change, one already-converged execution, and 100 targeted resources using current regional pricing.
  7. Design EventBridge notification for failed, timed-out, canceled, and rejected executions without creating a notification loop.

The validator checks syntax and required safety markers; it does not prove AWS API semantics, IAM adequacy, or production safety.

Optional live sandbox

Use only an owned disposable target group with account-owner approval. Do not use the AWS244 fictional identifiers.

Before execution:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws ssm get-document \
  --name nw-orders-target-health-recovery \
  --document-version '1' \
  --document-format YAML

Pin a reviewed numeric document version. $LATEST is useful during development but can change; $DEFAULT can be reassigned. Record document hash, version, caller, role ARN, Region, resource ARN, change ID, and UTC before starting.

Start only after replacing every placeholder and reviewing permissions:

aws ssm start-automation-execution \
  --document-name nw-orders-target-health-recovery \
  --document-version '1' \
  --parameters file://approved-parameters.json

Inspect without exposing sensitive output:

aws ssm get-automation-execution \
  --automation-execution-id "$EXECUTION_ID" \
  --query 'AutomationExecution.{Status:AutomationExecutionStatus,Document:DocumentName,Version:DocumentVersion,Target:Parameters.TargetGroupArn,Failure:FailureMessage}'

aws ssm describe-automation-step-executions \
  --automation-execution-id "$EXECUTION_ID" \
  --query 'StepExecutions[].{Step:StepName,Action:Action,Status:StepStatus,Failure:FailureMessage}' \
  --output table

Never auto-approve your own production change merely to make the lab finish. Test denial, timeout, stale state, and missing permission in a disposable environment.

Test matrix

CaseExpected behavior
owned target, /healthz, approval, targets become healthyone mutation, bounded wait, verified success
owned target already at /readyno approval and no mutation; AlreadyDesired success
wrong/missing owner tagfail closed before approval
current path is /custom-healthreject stale assumption; do not overwrite another change
approval denied or timed outno mutation
role lacks ModifyTargetGroupmutation fails; report exact denied API/resource
path changes but targets remain unhealthycompensate to proven old path, then report recovery failure
cancellation after mutationattempt supported bounded compensation; operator verifies final state
two concurrent executionsonly one may own the change; demonstrate lock/serialization design

Evidence and observability

Retain the document name/version/hash, execution ID/status, step status/failure, redacted inputs, approval decision/comment, CloudTrail management events, before/after resource state, health evidence, rollback result, and cleanup inventory. Automation and step status changes can be routed through EventBridge. aws:executeScript output can be sent to CloudWatch Logs. CloudTrail proves control-plane API activity; it does not prove customer transactions succeeded.

Do not emit secrets, parameter values, credentials, presigned URLs, full account IDs, or sensitive application responses into step outputs or logs. Execution history is an audit surface.

Troubleshooting by layer

SymptomFirst evidenceLikely boundary
AccessDeniedException starting executioncaller policy and document ARN/versionstarter lacks ssm:StartAutomationExecution
error mentioning iam:PassRolecaller policy and exact role ARNstarter cannot pass Automation role
role cannot be assumedrole existence and trust policyinvalid ARN/trust/source conditions
API action deniedfailed step and CloudTrailAutomation role lacks exact service action/resource
selector returns empty/type errorraw API response and JSONPathwrong output selector or static type
wait times outactual resource state and timeoutwrong expected value, delayed convergence, or failed service behavior
approval stays waitingapprover principal, SNS policy/topic name, timeoutnotification or authorization problem
overall success despite optional step failureonFailure and isCriticalstatus semantics were designed incorrectly
rollback step failedcurrent state, role permissions, concurrent changescompensation was not guaranteed

Cost and cleanup

As of this review, Systems Manager Automation charges per initiated step per resource, and aws:executeScript also charges by duration. The old general Automation free tier ended December 31, 2025; eligible new-account credits may differ. Attachments incur storage and possibly cross-account/Region transfer charges. Add costs for CloudWatch Logs, SNS, Lambda, and any resource changed by the runbook.

The supplied-evidence track creates nothing. For a live sandbox, export redacted evidence, delete only the course-owned test document and role/policies after checking dependencies, restore or delete the disposable target group through its owner stack, remove temporary topics/log groups according to retention policy, and query by course tags. A stopped execution is not cleanup.

Architecture and certification traps

  • A successful API response is not end-to-end recovery.
  • onFailure does not magically roll back earlier changes.
  • MaxErrors=0 can still leave already-running targets in flight.
  • “Latest” is not a stable production document version.
  • Approval after mutation is not preventive control.
  • A broad role makes execution easier but destroys least privilege.
  • A tag-based target set can change between review and execution.
  • Retrying permanent authorization or invariant failures increases noise, not reliability.
  • Automation history may contain sensitive inputs and outputs.

Knowledge check

  1. Why does /ready require both a configuration check and target-health verification?
  2. What permissions belong to the starter, and what permissions belong to the Automation role?
  3. Why must /custom-health fail closed instead of being replaced?
  4. What is the difference between onFailure, isCritical, and explicit compensation?
  5. Why can MaxErrors=0 still allow more than one failure at concurrency greater than one?
  6. When is AlreadyDesired success, and when could it hide an incomplete recovery?
  7. Why should production execute a numeric document version?
  8. Which evidence proves control-plane change, and which proves customer recovery?

Lesson acceptance

You pass when your workbook and review demonstrate all of the following:

  • correct playbook/runbook/script/Automation distinctions;
  • complete identity and least-privilege model, including iam:PassRole;
  • explicit invariants, idempotency, finite waits/retries, and stale-state rejection;
  • branch proof that unsafe inputs cannot reach mutation;
  • test results for every supplied case, including cancellation and concurrency;
  • explicit compensation and escalation for failed compensation;
  • document version/hash and redacted execution-evidence plan;
  • current cost calculation and no-create or verified cleanup evidence.

Official sources

Advertisement