AWS 402: Step Functions human approval, timeout, escalation, and rollback
Why this lesson matters
Incident remediation often needs automated evidence plus accountable human approval. Step Functions can orchestrate tasks, parallel collection, callbacks, timeouts, escalation, retries, catches and compensation, but it does not make an unsafe action reversible or an approval trustworthy by itself.
Workflow and service choice
Use Standard Workflows when durable execution history, exactly-once workflow execution semantics, long waits/callbacks and auditable incident orchestration are required. Express Workflows fit high-rate short workloads with different duration/delivery/history characteristics and are usually not the approval default. Verify current quotas and integrations.
| State/control | Incident purpose | Safety requirement |
|---|---|---|
| Parallel/Map | Collect independent evidence | Bound concurrency and tolerate partial sources explicitly |
| Choice | Branch on validated facts | Default/fallback branch and schema checks |
| Task integration | Invoke AWS API/Lambda/Automation | Least-privilege role, timeout and idempotency |
| Callback token | Wait for approval/external work | Secret single-use token, authorized decision binding |
| Wait/timeout/heartbeat | Bound delay and detect lost worker | Timeout shorter than incident objective, escalation |
| Retry/Catch | Handle transient/terminal errors | Retry only safe errors; route compensation |
| Succeed/Fail | Explicit final outcome | Verify user state before success |
Secure approval callback
The workflow creates a task token for .waitForTaskToken and sends it through a trusted backend, not directly into public chat/email logs. Store only encrypted/short-lived approval records needed to bind execution, action summary/hash, target, evidence, requester, allowed approver group, decision, UTC time and reason. The approver authenticates to a backend, which authorizes separation-of-duties and calls SendTaskSuccess or SendTaskFailure with the token.
Treat the token like a secret capability. Prevent logs, URLs, tickets and analytics from capturing it. Enforce one decision, expiry and replay rejection. A timeout produces a new token if retried; stale approvals must not authorize a new attempt. Approval text must show exact resources, action, blast radius, rollback limits and evidence freshness.
Timeouts, retries and compensation
Set top-level and task timeouts. Callback/worker tasks use heartbeat shorter than timeout where supported so lost work is detected. Escalate before expiry, but absence of approval defaults to no mutation. Use EventBridge/SNS for reminders without creating duplicate executions.
Retry only transient, idempotent operations with bounded attempts, exponential backoff and jitter. Do not retry authorization, validation or irreversible side effects blindly. Catch errors by class and record partial work. Compensation is a forward business action, not database transaction rollback: restore traffic/config when safe, release locks, reopen quarantine, notify, and reconcile external effects. Compensation can fail and needs escalation.
Pass large evidence through S3 with encrypted immutable references rather than state payload. Redact execution input/output and CloudWatch logs. State-machine role can invoke only approved resources; task roles remain separate. Version/alias the state machine and bind incident execution to definition revision.
Approval governance and recovery
Define the approval policy before an incident: which action classes may be automated, which require one or two approvers, who may approve each environment, and which emergency role is allowed when the normal identity path fails. The callback backend must evaluate current identity and authorization rather than trusting membership copied into the original message. Record policy version, evidence age, action digest and decision so an auditor can prove that the approved action is the action executed.
Recovery starts from observed state, not from the workflow's last successful state. Before retrying or compensating, inspect the target and external systems to discover what already happened. Use a stable incident and action key across restarts, but create a unique attempt identifier for evidence. If execution history approaches retention or size limits, export the authorized audit record and immutable evidence references; never resume by manually editing history. Rehearse operator procedures for a stuck execution, revoked approver, unavailable callback service and failed compensation.
Workshop and inspection
Design: receive normalized incident; deduplicate; parallel collect CloudWatch/Config/deployment evidence; calculate confidence; automatically quarantine only a tiny preapproved canary or request approval; wait with heartbeat/timeout; execute SSM Automation; verify SLI/security state; compensate or escalate; update ticket and notify; close only after bake.
aws stepfunctions describe-state-machine --state-machine-arn ARN
aws stepfunctions list-executions --state-machine-arn ARN --status-filter RUNNING
aws stepfunctions describe-execution --execution-arn EXECUTION_ARN
aws stepfunctions get-execution-history --execution-arn EXECUTION_ARN --reverse-order
aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=SendTaskSuccess --region ap-south-1
Create ASL/pseudocode, IAM and callback API design. Analyze supplied histories for approved, rejected, timeout, worker heartbeat loss, remediation failure, verification failure and compensation failure.
Failure game day
Test 20 cases: duplicate event starts workflow, evidence stale, parallel branch hangs, payload exceeds limit, execution role broad, token logged, approval link forwarded, approver unauthorized, requester self-approves, decision replayed, stale token after retry, no heartbeat, no top timeout, reminder duplicates action, retry repeats side effect, catch hides error, compensation ordering wrong, compensation fails, workflow succeeds before SLI, and history retention/evidence unavailable.
Cost and acceptance
Price Standard state transitions or Express requests/duration, Lambda/SSM/API, S3/KMS, logs, notifications and long-running executions under current pricing. This lesson creates nothing.
Submit state machine, trust/data-flow, callback threat model, approval record, timeout/retry table, compensation graph, seven history diagnoses, 20 failures, cost and retention. Pass requires no exposed token, fail-closed timeout, separation of duties, idempotent retry and verified post-action outcome.