AWS 245: Convert the recovery steps into a tested runbook
Why this lesson matters
AWS244 ended with a recovered application and evidence explaining three independent faults. A successful command history is not yet automation. Commands copied into a script can repeat a mistake faster, against more resources, under a more powerful identity.
In this lesson you convert one proven recovery - restoring an Application Load Balancer target group's health-check path from /healthz to /ready - into a bounded AWS Systems Manager Automation runbook. The runbook reads before it writes, proves ownership, checks the expected old state, pauses for approval, changes one field, waits, verifies target health, and explicitly compensates if verification fails.
The supplied runbook is design-only. It contains placeholders and must not be executed unchanged. The complete learning path requires no AWS mutation.
Outcomes
By the end, you can:
- distinguish an incident playbook, human runbook, script, and Systems Manager Automation runbook;
- explain Automation's control plane, execution identity, document version, steps, inputs, outputs, branching, retries, timeout, cancellation, and status;
- design guard clauses and idempotent behavior instead of assuming the starting state;
- explain why rollback is explicit compensation rather than a database transaction;
- test healthy, failed, unauthorized, stale-state, canceled, and partial-failure cases;
- bound fleet blast radius with targets, concurrency, and error thresholds;
- collect execution evidence without leaking account identifiers or secrets;
- estimate step, script-duration, attachment, logging, notification, and changed-resource costs.
From manual recovery to controlled automation
| Artifact | Primary question | Typical content |
|---|---|---|
| Playbook | How does the team coordinate this class of incident? | roles, severity, communication, escalation, decision points |
| Runbook | How is one repeatable operation performed safely? | prerequisites, commands/actions, verification, rollback |
| Script | How are instructions executed? | code with inputs and exit behavior |
| SSM Automation runbook | How does AWS orchestrate controlled steps and retain execution state? | typed parameters, IAM context, actions, outputs, branches, approvals |
Do not automate a recovery merely because it once appeared to work. First prove the causal chain, owner source, safe scope, acceptable rollback, and success signal. Automation should reduce toil and variation; it must not hide uncertainty.
A short history
Operations began with handwritten checklists and shell scripts. Configuration-management systems made desired state repeatable, while orchestration engines added workflow state and branching. AWS Systems Manager Documents provide machine-readable operational documents. Command documents commonly use schema 2.2 or later; Automation runbooks use schema 0.3. Automation adds AWS API actions, scripts, waits, assertions, approvals, branching, outputs, execution history, and fleet rate controls.
This evolution does not remove human responsibility. An approval after an unbounded destructive step is theater. Put scope checks and evidence before the approval, and the mutation after it.
The execution model
operator or event
|
| ssm:StartAutomationExecution + permitted document version
v
Automation execution ---- assumes approved role ----> AWS service APIs
| |
| +--> exact resource permissions
+--> step state and outputs
+--> approval / wait / branch
+--> verification or compensation
+--> CloudTrail + optional EventBridge/CloudWatch evidence
An execution is not a shell session. Systems Manager advances named steps and records each step's status. A runbook document contains:
descriptionandschemaVersion: '0.3';- optional top-level
assumeRole; - typed
parameters, defaults, descriptions, and validation patterns; mainSteps, each with a uniquename, anaction, action-specificinputs, and shared controls;- step
outputsselected from an action result; - optional top-level
outputsexposed by the execution.
Reference an earlier output as {{ stepName.outputName }}. Types are static: do not pass a MapList where a String is expected.
Common actions
| Action | Use | Important boundary |
|---|---|---|
aws:executeAwsApi | call one AWS API | inspect API names, pagination, throttling, and return shape |
aws:executeScript | run bounded Python or PowerShell | API calls from a customer runbook require an Automation assume role |
aws:runCommand | run an SSM Command document on managed nodes | node readiness and command permissions remain dependencies |
aws:waitForAwsResourceProperty | poll until a property reaches an allowed value | always choose a finite timeout |
aws:assertAwsResourceProperty | fail if a resource property differs | useful for guard clauses, not a lock against later change |
aws:branch | select the next step | include a safe default path |
aws:approve | pause for designated principals | timeout defaults to 7 days, maximum 30 days; unsupported in multi-account/Region automation |
aws:invokeLambdaFunction | invoke existing Lambda logic | adds function IAM, retries, logs, and cost dependencies |
aws:createStack | create a CloudFormation stack | stack rollback semantics are separate from Automation semantics |
Every action supports shared controls such as timeoutSeconds and maxAttempts. onFailure can abort, continue, or jump to step:name. onCancel can abort or jump to a supported step, but the cancellation workflow has a two-minute maximum and cannot jump to certain long-running or approval actions. isCritical affects the final execution status. nextStep changes normal flow; isEnd terminates it. Defaults are behavior, so review them rather than relying on memory.
Identity and least privilege
Two identities matter:
- The starter needs permission to start the chosen document/version and pass the exact Automation role when one is supplied.
- The Automation role needs a trust relationship for Systems Manager and only the API/resource permissions used by the steps.
For many non-script runbooks, omitting assumeRole causes Automation to use the starter's IAM context. That makes behavior vary by operator and is unsuitable for a controlled production process. A customer runbook whose aws:executeScript calls AWS APIs requires an IAM service role. Passing a role requires iam:PassRole; restrict both the role resource and iam:PassedToService where applicable. Scope the trust policy with source-account/source-ARN conditions to reduce confused-deputy risk.
Do not solve access errors by attaching AdministratorAccess, AmazonSSMFullAccess, or a broad Resource: '*' mutation policy. Derive permissions from the exact read, write, approval, logging, and rollback API calls. Remember that rollback needs permission too.
Safe automation invariants
Write invariants before steps. This lab uses these:
- The account and Region are the approved sandbox.
- The target group ARN is explicit; discovery cannot select a look-alike resource.
- Its
CourseOwnertag equals the approved value. - The current path is either the expected broken value or the already-correct value.
- Any third value means the world changed; stop without mutation.
- The desired path is
/ready, derived from application-owner evidence. - Only the health-check path may change.
- Mutation requires a named change ID and human approval.
- Success means configuration convergence and healthy targets over an agreed observation window.
- A failed post-change verification attempts bounded compensation and reports failure even if compensation succeeds.
These are guard clauses, not decorative comments. A precheck followed by a delayed write still has a time-of-check/time-of-use race. Reduce that window, serialize execution for the same target, re-read immediately before mutation when risk requires it, and stop on unexpected state.
Idempotency, retries, and eventual consistency
An idempotent operation can be retried without creating additional unintended change. This runbook treats /ready as already converged and exits successfully without approval or mutation. It permits /healthz only when ownership and change inputs match. Any other current path is rejected.
Retries belong around transient operations, not business ambiguity. Retrying AccessDenied, an ownership mismatch, or a failed invariant does not make it safe. AWS APIs may be eventually consistent and may throttle. Use finite retries, waits, and total execution deadlines; distinguish “API accepted the change” from “the application is healthy.”
Approval is a control, not proof
aws:approve pauses until the minimum approval count is met, a denial arrives, or the step times out. Approvers can be IAM users or role principals; SNS notification is optional, and if used the topic name must begin with Automation. The approval message should contain resource, old state, desired state, change ID, evidence link, rollback, and expiry - never a secret.
Approval does not prove the approver read the evidence, prevent the resource changing afterward, or provide separation of duties automatically. Enforce who may start, pass the role, approve, and alter the document. For governed change workflows, evaluate Systems Manager Change Manager rather than treating one aws:approve step as a complete change-management system.
Rollback is compensation, not a transaction
Automation does not atomically undo earlier AWS API calls when a later step fails. onFailure merely controls workflow direction. Design a compensation step for every mutation and record what original value is safe to restore.
Compensation can also fail because permissions changed, an API is unavailable, or another actor changed the resource. Therefore:
- capture and validate the exact pre-change state;
- make the rollback step bounded and observable;
- do not report the overall recovery as successful merely because rollback ran;
- escalate with current state and execution ID when compensation fails;
- prefer roll-forward when rollback would lose data or violate compatibility.
The supplied design restores /healthz only because the guard proved that exact starting state. It never guesses a rollback value.
Fleet and multi-account blast radius
Automation can target resources by IDs, tags, or resource groups. At fleet scale use:
MaxConcurrencyto limit simultaneous resource executions;MaxErrorsto stop dispatching to new targets after the threshold;- location-level concurrency/error controls for account–Region pairs;
- a canary target, then progressively larger waves;
- explicit account, OU, Region, and exclusion lists.
An error threshold does not cancel executions already running. If the maximum acceptable failures must never exceed one, concurrency must be one and the operation must be safely serial. The aws:approve action is not supported for multi-account/Region executions, so place governance at an appropriate outer workflow boundary.
Practical lab: supplied-evidence track
Download the AWS245 recovery-runbook pack. It contains:
RECOVERY_RUNBOOK_WORKBOOK.md- design and evidence worksheet;orders-target-health-recovery.DESIGN-ONLY.yaml- annotated schema 0.3 runbook;SUPPLIED_RUNBOOK_TEST_CASES.md- expected results for safe and unsafe states;validate_runbook.sh- local structural checks that make no AWS calls.
Run the local checks:
tar -xzf nw-p12-recovery-runbook.tar.gz
cd p12-recovery-runbook
sh validate_runbook.sh orders-target-health-recovery.DESIGN-ONLY.yaml
Then complete the workbook:
- Map each AWS244 manual recovery statement to a read, decision, approval, mutation, wait, verification, compensation, or evidence step.
- Identify every AWS API action and derive a least-privilege role outline.
- Trace all branches. Prove that
RejectandAlreadyDesiredcannot reach mutation. - Execute the supplied test cases on paper. Include the exact last step and final status.
- Explain the time-of-check/time-of-use race and propose serialization.
- Calculate the bill for one successful change, one already-converged execution, and 100 targeted resources using current regional pricing.
- Design EventBridge notification for failed, timed-out, canceled, and rejected executions without creating a notification loop.
The validator checks syntax and required safety markers; it does not prove AWS API semantics, IAM adequacy, or production safety.
Optional live sandbox
Use only an owned disposable target group with account-owner approval. Do not use the AWS244 fictional identifiers.
Before execution:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws ssm get-document \
--name nw-orders-target-health-recovery \
--document-version '1' \
--document-format YAML
Pin a reviewed numeric document version. $LATEST is useful during development but can change; $DEFAULT can be reassigned. Record document hash, version, caller, role ARN, Region, resource ARN, change ID, and UTC before starting.
Start only after replacing every placeholder and reviewing permissions:
aws ssm start-automation-execution \
--document-name nw-orders-target-health-recovery \
--document-version '1' \
--parameters file://approved-parameters.json
Inspect without exposing sensitive output:
aws ssm get-automation-execution \
--automation-execution-id "$EXECUTION_ID" \
--query 'AutomationExecution.{Status:AutomationExecutionStatus,Document:DocumentName,Version:DocumentVersion,Target:Parameters.TargetGroupArn,Failure:FailureMessage}'
aws ssm describe-automation-step-executions \
--automation-execution-id "$EXECUTION_ID" \
--query 'StepExecutions[].{Step:StepName,Action:Action,Status:StepStatus,Failure:FailureMessage}' \
--output table
Never auto-approve your own production change merely to make the lab finish. Test denial, timeout, stale state, and missing permission in a disposable environment.
Test matrix
| Case | Expected behavior |
|---|---|
owned target, /healthz, approval, targets become healthy | one mutation, bounded wait, verified success |
owned target already at /ready | no approval and no mutation; AlreadyDesired success |
| wrong/missing owner tag | fail closed before approval |
current path is /custom-health | reject stale assumption; do not overwrite another change |
| approval denied or timed out | no mutation |
role lacks ModifyTargetGroup | mutation fails; report exact denied API/resource |
| path changes but targets remain unhealthy | compensate to proven old path, then report recovery failure |
| cancellation after mutation | attempt supported bounded compensation; operator verifies final state |
| two concurrent executions | only one may own the change; demonstrate lock/serialization design |
Evidence and observability
Retain the document name/version/hash, execution ID/status, step status/failure, redacted inputs, approval decision/comment, CloudTrail management events, before/after resource state, health evidence, rollback result, and cleanup inventory. Automation and step status changes can be routed through EventBridge. aws:executeScript output can be sent to CloudWatch Logs. CloudTrail proves control-plane API activity; it does not prove customer transactions succeeded.
Do not emit secrets, parameter values, credentials, presigned URLs, full account IDs, or sensitive application responses into step outputs or logs. Execution history is an audit surface.
Troubleshooting by layer
| Symptom | First evidence | Likely boundary |
|---|---|---|
AccessDeniedException starting execution | caller policy and document ARN/version | starter lacks ssm:StartAutomationExecution |
error mentioning iam:PassRole | caller policy and exact role ARN | starter cannot pass Automation role |
| role cannot be assumed | role existence and trust policy | invalid ARN/trust/source conditions |
| API action denied | failed step and CloudTrail | Automation role lacks exact service action/resource |
| selector returns empty/type error | raw API response and JSONPath | wrong output selector or static type |
| wait times out | actual resource state and timeout | wrong expected value, delayed convergence, or failed service behavior |
| approval stays waiting | approver principal, SNS policy/topic name, timeout | notification or authorization problem |
| overall success despite optional step failure | onFailure and isCritical | status semantics were designed incorrectly |
| rollback step failed | current state, role permissions, concurrent changes | compensation was not guaranteed |
Cost and cleanup
As of this review, Systems Manager Automation charges per initiated step per resource, and aws:executeScript also charges by duration. The old general Automation free tier ended December 31, 2025; eligible new-account credits may differ. Attachments incur storage and possibly cross-account/Region transfer charges. Add costs for CloudWatch Logs, SNS, Lambda, and any resource changed by the runbook.
The supplied-evidence track creates nothing. For a live sandbox, export redacted evidence, delete only the course-owned test document and role/policies after checking dependencies, restore or delete the disposable target group through its owner stack, remove temporary topics/log groups according to retention policy, and query by course tags. A stopped execution is not cleanup.
Architecture and certification traps
- A successful API response is not end-to-end recovery.
onFailuredoes not magically roll back earlier changes.MaxErrors=0can still leave already-running targets in flight.- “Latest” is not a stable production document version.
- Approval after mutation is not preventive control.
- A broad role makes execution easier but destroys least privilege.
- A tag-based target set can change between review and execution.
- Retrying permanent authorization or invariant failures increases noise, not reliability.
- Automation history may contain sensitive inputs and outputs.
Knowledge check
- Why does
/readyrequire both a configuration check and target-health verification? - What permissions belong to the starter, and what permissions belong to the Automation role?
- Why must
/custom-healthfail closed instead of being replaced? - What is the difference between
onFailure,isCritical, and explicit compensation? - Why can
MaxErrors=0still allow more than one failure at concurrency greater than one? - When is
AlreadyDesiredsuccess, and when could it hide an incomplete recovery? - Why should production execute a numeric document version?
- Which evidence proves control-plane change, and which proves customer recovery?
Lesson acceptance
You pass when your workbook and review demonstrate all of the following:
- correct playbook/runbook/script/Automation distinctions;
- complete identity and least-privilege model, including
iam:PassRole; - explicit invariants, idempotency, finite waits/retries, and stale-state rejection;
- branch proof that unsafe inputs cannot reach mutation;
- test results for every supplied case, including cancellation and concurrency;
- explicit compensation and escalation for failed compensation;
- document version/hash and redacted execution-evidence plan;
- current cost calculation and no-create or verified cleanup evidence.