Lesson 231 · AWS Learning Path

AWS 231: Troubleshoot stack failure, rollback, drift, agent, patch, and runbook failures

· Published · 11 min read

Labelled process diagram for AWS 231: Failed deployment or operation to Events and per-target evidence to First causal boundary to Controlled retry, verified state, and cleanup, with decision, proof and rejection...

Why this lesson matters

A final red error is often a consequence, not a cause. CloudFormation may cancel dependent resources after one denial; Run Command may report aggregate success while a node plugin failed; patch compliance may describe a different baseline; and an Automation execution can start successfully before a later step fails.

An architect diagnoses from ordered evidence at the smallest failing boundary. Guessing, granting administrator access, rerunning every target, or deleting a failed stack destroys evidence and can increase damage.

Outcomes

You will be able to:

  • freeze change, establish account/Region/time scope, and preserve evidence;
  • construct a UTC timeline and identify first cause versus cascade;
  • distinguish CloudFormation validation, change-set, handler, rollback, deletion,

drift, hook/policy, IAM and underlying-service failures;

  • diagnose SSM Agent registration, credentials, DNS, route, endpoint, TLS and

service failures before blaming command content;

  • trace Run Command aggregate → invocation → plugin status and output;
  • prove patch execution type, baseline, snapshot, repository and node evidence;
  • separate State Manager schedule/target/compliance from command execution;
  • separate Automation start, assume-role, step input, API, timeout and

compensation failures;

  • make one bounded repair, retry from the correct boundary, verify state, and

preserve a regression control.

Safety and evidence rules

  • The supplied casebook is complete T0 work. Live investigation is read-only

unless the resource owner separately approves a repair.

  • Never disable rollback by habit, skip rollback resources casually, terminate

a managed node, open SSH/RDP broadly, detach policies, rerun patches fleet-wide, or delete logs/stacks merely to clear red status.

  • Record exact caller, account, Region, UTC window, resource IDs, execution IDs,

document/template versions and first error. Redact account IDs, private data, secrets and full ARNs from submissions.

  • Preserve CloudTrail, stack events, command/plugin output, agent logs, patch

snapshots, Automation steps and before/after state before cleanup.

Download the real failure set

The identifiers and timestamps are fictional. Diagnose them; do not execute a repair command copied from a card.

One method for every failure

freeze mutation
  -> scope caller/account/Region/resource/time
  -> preserve raw evidence
  -> state intended owner/source and last known good
  -> order events oldest-to-newest
  -> locate first failed boundary
  -> separate cause from cancellation/rollback/retry symptoms
  -> form one falsifiable hypothesis
  -> run one read-only discriminating query
  -> approve smallest reversible repair
  -> retry from correct boundary with new execution ID
  -> prove actual state + control-plane status + telemetry
  -> restore injected fault / clean up / add regression guard

Do not start with “What command fixes this?” Start with “Which component first departed from intended state, under which identity, and what evidence disproves the alternatives?”

Build an authoritative evidence envelope

export AWS_DEFAULT_REGION="ap-south-1"
date -u +%FT%TZ
aws sts get-caller-identity --output json
aws configure list

Record resource owner tags, IaC repository/version, recent deployments/changes, CloudTrail event time, service quota state, and correlation IDs. CLI pagination matters: --max-items can produce a token, so retrieve remaining pages before claiming no earlier failure. Service clocks/timestamps are more reliable than the order screenshots happened to be captured.

CloudFormation decision tree

Boundary 1: before stack execution

validate-template checks template syntax/basic structure, not every account permission, runtime dependency, policy, hook, quota, dynamic reference, or replacement outcome. Run local lint/policy tests, AWS validation, then create and inspect a change set.

aws cloudformation validate-template --template-body file://template.yaml
aws cloudformation create-change-set --stack-name "$stack_name" \
  --change-set-name nw-diagnosis-001 --change-set-type UPDATE \
  --template-body file://template.yaml --capabilities CAPABILITY_NAMED_IAM
aws cloudformation describe-change-set --stack-name "$stack_name" \
  --change-set-name nw-diagnosis-001 --output json

Failure here means no resource handler ran. Fix template/schema/transform/size or change-set inputs; do not search instance logs.

Boundary 2: stack/resource execution

aws cloudformation describe-stacks --stack-name "$stack_name" --output json
aws cloudformation describe-stack-events --stack-name "$stack_name" \
  --max-items 100 --output json
aws cloudformation list-stack-resources --stack-name "$stack_name" --output json

Sort by timestamp ascending for causality. Find the earliest *_FAILED, its logical/physical ID and reason. “Resource update cancelled,” rollback deletion, and parent-stack failure are usually cascades. Determine which identity failed:

  • caller authorizes stack/change-set actions and service-role passing;
  • CloudFormation service role, if configured, calls underlying services;
  • resource service-linked/execution roles may perform later work;
  • SCP, permissions boundary, session policy, resource policy, KMS key policy,

VPC endpoint policy and explicit denies can override an apparent allow.

Use CloudTrail and IAM simulation carefully; a simulator does not reproduce all resource policies, SCP conditions, service behavior, or runtime context.

Boundary 3: rollback states

StateMeaning and response
ROLLBACK_COMPLETEcreate failed and cleanup completed; diagnose original event; update is not allowed for a failed create stack
UPDATE_ROLLBACK_COMPLETEprior working stack restored; diagnose forward failure before retry
UPDATE_ROLLBACK_FAILEDrollback itself failed; fix rollback cause, then continue
DELETE_FAILEDat least one resource could not delete; inspect exact event/dependency/policy

For UPDATE_ROLLBACK_FAILED, normally repair the dependency/permission first:

aws cloudformation continue-update-rollback --stack-name "$stack_name"

--resources-to-skip is last-resort, owner-approved recovery. Specify the minimum logical resources that failed during rollback, not forward-update dependents. CloudFormation marks skipped resources complete while actual state remains inconsistent with the template. Reconcile before another update or the stack can fail again/become unrecoverable.

For DELETE_FAILED, retaining a failed resource can let stack deletion finish, but creates an orphan with separate security/cost/backup ownership. Record exact physical ID, tags, data retention and future deletion plan - absence of a stack does not prove absence of its resources.

Boundary 4: drift

Drift detection is asynchronous and only covers supported resource/property state. Poll its ID to DETECTION_COMPLETE; then inspect resource/property differences. NOT_CHECKED and unmodelled application data remain unknown.

drift_id="$(aws cloudformation detect-stack-drift --stack-name "$stack_name" \
  --query StackDriftDetectionId --output text)"
aws cloudformation describe-stack-drift-detection-status \
  --stack-drift-detection-id "$drift_id" --output json
aws cloudformation describe-stack-resource-drifts --stack-name "$stack_name" \
  --stack-resource-drift-status-filters MODIFIED DELETED NOT_CHECKED --output json

An unchanged source diff can coexist with actual drift. Decide whether to adopt actual state or restore intended state, encode the decision in owner source, deploy, query actual state, and rerun drift detection.

Systems Manager managed-node decision tree

A selectable managed node needs: supported SSM Agent installed/running with system privilege, valid managed-node identity/permissions, and outbound HTTPS DNS/network access to required regional endpoints.

aws ssm describe-instance-information \
  --filters Key=InstanceIds,Values="$instance_id" --output json
aws ec2 describe-instances --instance-ids "$instance_id" \
  --query 'Reservations[].Instances[].{State:State.Name,Profile:IamInstanceProfile.Arn,Vpc:VpcId,Subnet:SubnetId,IP:PrivateIpAddress}' \
  --output json

On an approved Linux console/session:

sudo systemctl status amazon-ssm-agent --no-pager
sudo journalctl -u amazon-ssm-agent --since '30 minutes ago' --no-pager
sudo tail -n 200 /var/log/amazon/ssm/amazon-ssm-agent.log
sudo tail -n 200 /var/log/amazon/ssm/errors.log
sudo ssm-cli get-diagnostics --output table

Check in this order:

  1. node is running and clock/TLS trust is sane;
  2. agent process/version/config/proxy and system/root execution;
  3. EC2 instance profile or hybrid activation identity and credential precedence;
  4. VPC DNS support/hostnames and resolver;
  5. route through NAT/internet or interface endpoint private DNS;
  6. node egress TCP 443 and endpoint security-group ingress TCP 443;
  7. regional ssm/ssmmessages path and feature-specific S3/KMS/Logs access;
  8. agent logs followed by refreshed PingStatus.

Modern Regions primarily use ssmmessages; endpoint needs can vary by Region and feature. An active (running) process proves neither credentials nor network. Do not add public SSH just because Session Manager is unavailable.

Run Command and State Manager

Run Command has command, per-target invocation, and per-plugin statuses. Aggregate Success can coexist with failed targets depending on delivery/error threshold; always enumerate details.

aws ssm get-command-invocation --command-id "$command_id" \
  --instance-id "$instance_id" --output json
aws ssm list-command-invocations --command-id "$command_id" \
  --details --output json

Classify delivery (Pending, delayed, undeliverable, target not connected), execution timeout, plugin response code/stdout/stderr, cancellation and error threshold. Retrying all targets can repeat successful side effects; rerun a canary or only proven failed targets with an idempotent document.

For State Manager, inspect definition, document/version, targets, schedule, apply-only-at-cron behavior, offset, concurrency/errors, association execution, target execution and compliance capture time. A successful old execution does not prove current desired state.

aws ssm describe-association --association-id "$association_id" --output json
aws ssm describe-association-executions --association-id "$association_id" --output json
aws ssm describe-association-execution-targets --association-id "$association_id" \
  --execution-id "$execution_id" --output json

Patch Manager decision tree

Separate four questions: Was the node targeted? Which operation executed? Which baseline/snapshot made patches eligible? Did the OS package manager install them?

aws ssm list-command-invocations --command-id "$command_id" --details --output json
aws ssm describe-instance-patch-states --instance-ids "$instance_id" --output json
aws ssm describe-instance-patches --instance-id "$instance_id" --output json

Record Operation (Scan or Install), ExecutionType (Command or PatchPolicy), baseline ID, snapshot ID, OS/product/classification/severity, approval delay, rejected patches, install override list, reboot option, command and association IDs, per-plugin output and compliance timestamp.

Common boundaries:

  • patch-group tag not registered to expected classic baseline;
  • Quick Setup patch policy association/override differs from assumed baseline;
  • patch is not yet approved, rejected, superseded or wrong OS/product/class;
  • repository DNS/network/proxy/S3 access or package-manager lock fails;
  • /var lacks space;
  • overlapping commands/windows/associations race the patch payload;
  • NoReboot leaves pending reboot and incomplete application state;
  • stale compliance is mistaken for current state.

Run one Scan first; install only with change/reboot approval. A snapshot ID keeps an operation's node set evaluating the same approved snapshot; do not reuse an arbitrary old snapshot across unrelated operations.

Automation runbook decision tree

aws ssm get-automation-execution --automation-execution-id "$automation_id" --output json
aws ssm describe-document --name "$document_name" \
  --document-version "$document_version" --output json
aws ssm get-document --name "$document_name" \
  --document-version "$document_version" --document-format YAML --output text

Classify:

  1. start denied: caller lacks scoped ssm:StartAutomationExecution;
  2. pass-role denied: caller cannot pass the exact Automation role to SSM;
  3. role assumption fails: malformed/missing role or trust lacks ssm.amazonaws.com;
  4. validation fails: missing/wrong type/pattern/value or unresolved output;
  5. step API denied/fails: Automation role lacks action or resource precondition;
  6. timeout/retry/cancel: inspect timeoutSeconds, attempts and step status;
  7. branch/output mismatch: JSONPath or type does not match payload;
  8. partial side effects: completed steps remain unless compensation is modelled.

Use exact document version and step execution evidence. A new retry creates a new execution; preserve the original. Do not assume approval or rollback adds permissions or reverses arbitrary API calls.

Correlation and smallest-repair table

EvidenceProvesDoes not prove
stack *_COMPLETECloudFormation operation completedworkload health/data correctness
command aggregate successcommand met aggregate semanticsevery target/plugin succeeded
PingStatus=Onlinerecent agent service communicationdocument can execute successfully
patch compliantevaluated items comply with recorded baseline/timeevery possible package is current
Automation Successdefined runbook path completedexternal business outcome unless verified
no alarmalarm expression/state has not breachedmetric exists or system is healthy

One change per hypothesis keeps cause and effect observable. Record before/after, approval, blast radius and rollback. Retry at the failed boundary - not by rebuilding unrelated resources - and give every new command/Automation/deployment its own correlation ID.

Cost and cleanup

Read-only APIs are usually low cost but logs, Insights queries, retained stack resources, Automation steps, Run Command fleet execution, patch downloads, snapshots, NAT traffic and idle failure resources can charge. Repeated retries multiply both cost and side effects. Preserve required evidence according to retention policy, remove only approved lab resources, and prove exact negative inventory; never delete incident evidence to reduce the bill without approval.

Practical casebook work and acceptance

Complete all eight cards, not a subset. For each, fill the worksheet with scope, timeline, first cause, cascade, identity/network/data boundary, one discriminating read-only query, smallest repair, retry boundary, success proof, regression guard, cleanup and rejected unsafe shortcut.

Acceptance requires technically correct classification for every card and must explicitly explain:

  • why cancelled/rollback events are not automatically root cause;
  • why rollback resource skipping creates inconsistency;
  • why empty source diff can coexist with drift;
  • why agent-running does not prove managed-node readiness;
  • why aggregate command success can hide a target failure;
  • why patch baseline/execution type/snapshot must be evidenced;
  • why Automation start success and step success are different;
  • why completed runbook side effects are not automatically rolled back.

Reject broad privilege, fleet-wide retries, unsupported “fixed” claims, missing UTC/correlation IDs, screenshots without raw fields, or cleanup without evidence.

Knowledge check

  1. First CloudFormation event to inspect? Earliest failed resource event in

time order; later cancellation/rollback may be effects.

  1. When use resources-to-skip? Only owner-approved last resort for minimum

resources failed during rollback, followed by explicit reconciliation.

  1. Agent process is active but node is offline - next boundary? Credentials,

DNS, route/endpoints, security groups, proxy, TLS/clock and agent logs.

  1. Patch Success but wrong baseline? The operation succeeded under recorded

inputs; it does not prove compliance with the intended policy.

  1. Automation started but API step denied - which identity? Usually the

Automation assume role at that step, after confirming exact execution data.

Official sources

Advertisement