AWS 227: Automation runbooks, approvals, inputs, outputs, and rollback
Why this lesson matters
Systems Manager Automation coordinates AWS API calls, waits, branches, scripts, commands and approvals as a versioned runbook. It is not a transaction engine: retries can repeat side effects, cancellation has a short cleanup path, and rollback occurs only when the author explicitly models and verifies it.
What you will be able to do
By the end, you can:
- read schema 0.3 parameters, variables, assume role, main steps and outputs;
- distinguish
aws:executeAwsApi,aws:runCommand,aws:executeScript, waits, assertions, branches and nested automation; - define typed, constrained non-secret inputs and JSONPath-selected typed outputs;
- separate caller, Automation service role, node role and approver permissions;
- design retries, timeout, criticality, failure, cancellation and terminal paths;
- place meaningful human approval before irreversible/high-risk work;
- design compensating rollback and changed-state verification;
- inspect execution/step evidence without starting an automation.
Runbook execution model
caller + document/version + typed parameters
|
AutomationAssumeRole
|
v
step -> typed output -> branch/wait/assert -> next step
| |
+-> retry/timeout/onFailure path +-> approval
+-> onCancel path +-> child automation/Run Command
|
v
verified success OR verified compensation
Automation schema 0.3 uses sequential mainSteps unless nextStep, onFailure, onCancel, branching or terminal behavior changes the path. Draw every possible path, including timeout, rejection and cancellation.
Identity boundaries
| Identity | Required authority |
|---|---|
| initiating caller | start/stop/read exact runbook and pass the approved role |
AutomationAssumeRole | AWS APIs used by every step and child operation |
| managed-node role | node-side AWS calls made through aws:runCommand |
| approver principal | approve/reject the exact waiting execution |
| SNS topic policy/subscriber | route approval notification, not grant approval authority |
If no Automation role is specified, Automation uses the initiating user's permissions. That couples behavior to the operator and weakens repeatability. Production runbooks should use a least-privilege role, trust ssm.amazonaws.com, constrain confused-deputy risk where supported, and give callers only scoped iam:PassRole.
An approval does not elevate permissions. After approval, steps still use the Automation execution role.
Inputs, variables and outputs
Use the narrowest type and constraints: String, StringList, Integer, Boolean, StringMap, MapList, AWS resource types where supported, and strict allowed patterns/values. Keep secret plaintext out of parameters, execution views, outputs, approval messages and logs. Resolve secrets only inside a step that needs them with a scoped role and never return them.
Parameters are fixed execution inputs. Schema 0.3 variables are mutable workflow state and must retain their declared type. Prefer immutable step outputs over mutable variables when possible.
For aws:executeAwsApi, define each output with:
Name;- JSONPath
Selectoragainst the real API response; - declared
Type.
Reference it as {{ stepName.outputName }}. A selector that returns a list cannot be declared String merely because one result is expected. Empty arrays, pagination, eventual consistency and API shape changes need explicit handling.
Action selection
| Need | Action/direction |
|---|---|
| one AWS API operation | aws:executeAwsApi |
| wait until property reaches value | aws:waitForAwsResourceProperty with bounded timeout |
| fail unless property matches now | aws:assertAwsResourceProperty |
| choose path from parameter/output | aws:branch with explicit default |
| bounded code/data transformation | aws:executeScript; pin runtime/dependencies and constrain output |
| node operating-system action | aws:runCommand; inherit AWS224 target/output rules |
| reusable workflow | aws:executeAutomation with pinned child version and outputs |
| human risk decision | aws:approve |
| delay only | aws:sleep; prefer state polling over guessed sleep |
An API action's impact depends on the chosen API. DescribeInstances is read-only; TerminateInstances is destructive even though both use aws:executeAwsApi.
Step controls
timeoutSeconds: bound one attempt/step according to action behavior.maxAttempts: retries the step; use only when the operation/reconciliation is idempotent.isCritical: controls whether a failed step makes final execution fail under its path semantics.onFailure: defaults to abort; can continue or jump tostep:name.onCancel: abort or jump to an allowed cleanup step; cancellation workflow has a maximum two-minute window and cannot target several long-running/action types.nextStep: explicit success edge.isEnd: end execution after that step.
Never use onFailure: Continue merely to make the run green. A noncritical diagnostic failure can continue only when acceptance and final status report it explicitly.
Approval design
aws:approve pauses a simple execution for authenticated IAM users/roles:
- up to ten approvers and a positive
MinRequiredApprovalsno greater than that list; - optional SNS notification topic whose name must begin with
Automation; - default timeout seven days, configurable up to thirty days;
- output containing approval status/decisions;
- no support inside multi-account/Region automation executions.
Place approval after read-only prechecks and a generated change summary, but before the first high-impact mutation. The message must identify resource, account/Region, expected change, risk, evidence link, rollback and expiry - never secrets. Separate requester and approver. A timeout/rejection follows a tested safe terminal path and creates no mutation.
Rollback is compensation
AWS APIs do not share one transaction. A robust runbook captures before state before mutation, makes one bounded change, waits for the service's actual state, validates application behavior, and branches to compensation when needed.
Compensation can fail too. It needs its own retries/timeouts, permissions, outputs, alarm and human escalation. Some operations are irreversible: terminated instance-store data, deleted unversioned data or externally delivered messages cannot be rolled back. For those, use prevention, backup, replacement/cutover or explicit human gate.
Do not call “start the instance” a rollback unless the runbook proves it was running before the run and restores all relevant prior state.
Design exercise: safe instance-type change
Design, do not execute, a runbook for an approved EBS-backed non-Auto-Scaling EC2 instance:
- validate exact instance, owner tags, account/Region, EBS root and allowed current state;
- capture original type/state, attached volumes, ENIs and health evidence;
- assert new type architecture/virtualization/network/EBS compatibility and quota/capacity;
- create approved backup and verify completion;
- present before/after/risk/downtime/rollback to
aws:approve; - stop only if originally running; wait for stopped;
- modify instance type; verify property;
- start only if originally running; wait for EC2 and status checks;
- run application synthetic/metrics/log validation;
- on failure, stop, restore original type, start to original state and revalidate;
- if compensation fails, preserve evidence and escalate - never report success.
Reject this workflow for an Auto Scaling group member; change the launch template and replace through the group's deployment controls.
Output contract
Expose non-secret outputs such as original type, requested type, backup ID, final instance state, validation result and compensation status. Every output has a type and source step. Output alone is not proof: link it to execution ID and API/resource evidence.
Read-only inspection
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws ssm list-documents \
--filters Key=DocumentType,Values=Automation Key=Owner,Values=Self \
--query 'DocumentIdentifiers[].{Name:Name,Owner:Owner,Version:DocumentVersion,Platform:PlatformTypes}' \
--output table
aws ssm describe-automation-executions --max-results 20 \
--query 'AutomationExecutionMetadataList[].{Id:AutomationExecutionId,Document:DocumentName,Version:DocumentVersion,Status:AutomationExecutionStatus,Start:ExecutionStartTime,End:ExecutionEndTime}' \
--output table
For one exact approved historical execution:
execution_id="exact-approved-automation-execution-id"
aws ssm get-automation-execution \
--automation-execution-id "$execution_id" --output json
aws ssm describe-automation-step-executions \
--automation-execution-id "$execution_id" --output json
Capture runbook/version, mode, role, parameters (redacted), target, all step statuses/attempts/times/failure messages, outputs, approval decisions, compensation path and final resource behavior. Do not restart/stop/approve the execution.
Console review
Open Systems Manager > Automation > Executions, choose one approved historical execution, then:
- verify account/Region, document owner/name/version and execution mode;
- inspect parameters and assumed role;
- follow diagram/list through actual, skipped, failed and compensation steps;
- inspect every attempt, output and failure message;
- compare waiting approval identity/comment/time with separation-of-duties policy;
- verify final resource state directly;
- do not choose Execute, Retry, Stop or Approve/Deny.
Diagnose failures
| Symptom | First evidence | Correction |
|---|---|---|
| fails before first step | caller, PassRole, role trust and parameters | repair exact identity/input |
| API denied in step | Automation role and resource/context | add least privilege, not administrator |
| selector output empty/type error | raw API shape, pagination, JSONPath/type | correct selector and empty-result branch |
| wait times out | actual property, target ID and service events | repair root cause; don't add blind sleep |
| branch takes wrong path | typed variable/output and choice order/default | add deterministic conditions/default failure |
| approval never arrives | SNS topic prefix/policy/subscription and approver list | repair notification; approval authority remains IAM |
| retry duplicates resource | idempotency token/reconciliation | detect existing side effect before retry |
| cancelled but change remains | onCancel support, two-minute path and actual state | invoke tested compensation/escalation |
| rollback says success, app unhealthy | direct resource plus synthetic/metrics/log evidence | fail execution and escalate |
Cost and cleanup
Automation can charge by steps and trigger charged APIs, Run Command, Lambda, logs, SNS, backups and resources. Retries, nested/multi-account runs and long waits multiply operational load. Approval delay can retain temporary resources.
AWS227 creates/runs nothing. Do not alter documents, default versions, executions, approvals or roles inspected. The runbook design, path diagram, typed output table and compensation matrix are local deliverables.
Knowledge check
- Does Automation roll back automatically?
No; the author must implement and verify compensation.
- What identity calls AWS APIs in a role-based runbook?
The configured AutomationAssumeRole.
- Why type JSONPath outputs?
Later steps need a predictable value shape; wrong/empty selection must fail safely.
- Can
aws:approvegate multi-account/Region executions?
It does not support those execution modes.
- Does cancellation guarantee cleanup?
No; it is best effort with a short, action-limited cancellation path.
Lesson acceptance
- All identities, inputs, variables, steps, outputs and execution paths are mapped.
- Parameters/selectors are typed and secrets are excluded.
- Approval has bounded approvers, threshold, timeout, evidence and rejection path.
- Every mutating step has idempotency, timeout/retry and postcondition.
- Compensation restores captured prior state or explicitly escalates irreversible/failing recovery.
- A historical execution is inspected read-only or a complete no-resource evidence design is supplied.