Lesson 402 · AWS Learning Path

AWS 402: Step Functions human approval, timeout, escalation, and rollback

· Published · 5 min read

Labelled process diagram for AWS 402: Versioned intent to Automated validation to Controlled AWS change to Observed result and retained evidence, with decision, proof and rejection evidence.

Why this lesson matters

Incident remediation often needs automated evidence plus accountable human approval. Step Functions can orchestrate tasks, parallel collection, callbacks, timeouts, escalation, retries, catches and compensation, but it does not make an unsafe action reversible or an approval trustworthy by itself.

Workflow and service choice

Use Standard Workflows when durable execution history, exactly-once workflow execution semantics, long waits/callbacks and auditable incident orchestration are required. Express Workflows fit high-rate short workloads with different duration/delivery/history characteristics and are usually not the approval default. Verify current quotas and integrations.

State/controlIncident purposeSafety requirement
Parallel/MapCollect independent evidenceBound concurrency and tolerate partial sources explicitly
ChoiceBranch on validated factsDefault/fallback branch and schema checks
Task integrationInvoke AWS API/Lambda/AutomationLeast-privilege role, timeout and idempotency
Callback tokenWait for approval/external workSecret single-use token, authorized decision binding
Wait/timeout/heartbeatBound delay and detect lost workerTimeout shorter than incident objective, escalation
Retry/CatchHandle transient/terminal errorsRetry only safe errors; route compensation
Succeed/FailExplicit final outcomeVerify user state before success

Secure approval callback

The workflow creates a task token for .waitForTaskToken and sends it through a trusted backend, not directly into public chat/email logs. Store only encrypted/short-lived approval records needed to bind execution, action summary/hash, target, evidence, requester, allowed approver group, decision, UTC time and reason. The approver authenticates to a backend, which authorizes separation-of-duties and calls SendTaskSuccess or SendTaskFailure with the token.

Treat the token like a secret capability. Prevent logs, URLs, tickets and analytics from capturing it. Enforce one decision, expiry and replay rejection. A timeout produces a new token if retried; stale approvals must not authorize a new attempt. Approval text must show exact resources, action, blast radius, rollback limits and evidence freshness.

Timeouts, retries and compensation

Set top-level and task timeouts. Callback/worker tasks use heartbeat shorter than timeout where supported so lost work is detected. Escalate before expiry, but absence of approval defaults to no mutation. Use EventBridge/SNS for reminders without creating duplicate executions.

Retry only transient, idempotent operations with bounded attempts, exponential backoff and jitter. Do not retry authorization, validation or irreversible side effects blindly. Catch errors by class and record partial work. Compensation is a forward business action, not database transaction rollback: restore traffic/config when safe, release locks, reopen quarantine, notify, and reconcile external effects. Compensation can fail and needs escalation.

Pass large evidence through S3 with encrypted immutable references rather than state payload. Redact execution input/output and CloudWatch logs. State-machine role can invoke only approved resources; task roles remain separate. Version/alias the state machine and bind incident execution to definition revision.

Approval governance and recovery

Define the approval policy before an incident: which action classes may be automated, which require one or two approvers, who may approve each environment, and which emergency role is allowed when the normal identity path fails. The callback backend must evaluate current identity and authorization rather than trusting membership copied into the original message. Record policy version, evidence age, action digest and decision so an auditor can prove that the approved action is the action executed.

Recovery starts from observed state, not from the workflow's last successful state. Before retrying or compensating, inspect the target and external systems to discover what already happened. Use a stable incident and action key across restarts, but create a unique attempt identifier for evidence. If execution history approaches retention or size limits, export the authorized audit record and immutable evidence references; never resume by manually editing history. Rehearse operator procedures for a stuck execution, revoked approver, unavailable callback service and failed compensation.

Workshop and inspection

Design: receive normalized incident; deduplicate; parallel collect CloudWatch/Config/deployment evidence; calculate confidence; automatically quarantine only a tiny preapproved canary or request approval; wait with heartbeat/timeout; execute SSM Automation; verify SLI/security state; compensate or escalate; update ticket and notify; close only after bake.

aws stepfunctions describe-state-machine --state-machine-arn ARN
aws stepfunctions list-executions --state-machine-arn ARN --status-filter RUNNING
aws stepfunctions describe-execution --execution-arn EXECUTION_ARN
aws stepfunctions get-execution-history --execution-arn EXECUTION_ARN --reverse-order
aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=SendTaskSuccess --region ap-south-1

Create ASL/pseudocode, IAM and callback API design. Analyze supplied histories for approved, rejected, timeout, worker heartbeat loss, remediation failure, verification failure and compensation failure.

Failure game day

Test 20 cases: duplicate event starts workflow, evidence stale, parallel branch hangs, payload exceeds limit, execution role broad, token logged, approval link forwarded, approver unauthorized, requester self-approves, decision replayed, stale token after retry, no heartbeat, no top timeout, reminder duplicates action, retry repeats side effect, catch hides error, compensation ordering wrong, compensation fails, workflow succeeds before SLI, and history retention/evidence unavailable.

Cost and acceptance

Price Standard state transitions or Express requests/duration, Lambda/SSM/API, S3/KMS, logs, notifications and long-running executions under current pricing. This lesson creates nothing.

Submit state machine, trust/data-flow, callback threat model, approval record, timeout/retry table, compensation graph, seven history diagnoses, 20 failures, cost and retention. Pass requires no exposed token, fail-closed timeout, separation of duties, idempotent retry and verified post-action outcome.

Official sources

Advertisement