Lesson 384 · AWS Learning Path

AWS 384: Troubleshoot template, dependency, rollback, drift, bootstrap, and cross-account failures

· Published · 4 min read

Labelled process diagram for AWS 384: Versioned intent to Automated validation to Controlled AWS change to Observed result and retained evidence, with decision, proof and rejection evidence.

Why this lesson matters

Infrastructure deployment failures cross source, transform, dependency graph, IAM, service quota/state, custom providers, rollback, drift, CDK bootstrap/assets/context, StackSets, accounts, Regions, and cleanup. Correct diagnosis starts with the first causal event and exact execution identity.

Failure localization ladder

LayerEvidenceTypical failure
AuthoringParser/linter/schemaindentation, type/property, unresolved reference
Transform/synthProcessed template/assemblymacro, context, logical ID, generated IAM
PlanDiff/change setunexpected replace/delete/capability
AuthorizationCloudTrail and policy chaintrust, pass-role, SCP, boundary, key/bucket policy
DependencyStack events/resource graphcycle, missing export, wrong order
Provider/serviceResource status/reasonquota, name, state, capacity, Region support
RecoveryRollback events/physical stateold config unavailable or external change
Multi-accountStackSet/CDK role sessionstarget, bootstrap, trust, artifact access

Freeze automatic retries and record account/Region, caller, source commit, template/assembly hash, stack/change set/execution/operation IDs, parameters without secret values, CDK version/context, bootstrap qualifier/version, target accounts, first failure time, and customer effect. Events are reverse chronological in some interfaces; reconstruct UTC order and dependency branches.

Template, dependency, and rollback diagnosis

Parsing and validation do not prove service constraints or runtime success. Inspect processed templates after transforms. For dependency cycles, map explicit DependsOn plus implicit references, attributes, exports/imports, security-group relationships, and nested outputs. Fix ownership/contracts rather than adding arbitrary ordering.

For authorization, evaluate actual principal and API against identity/resource policies, role trust, session policy, permissions boundary, SCP/RCP where relevant, endpoint policy, KMS key/grant, and conditions. Do not attach administrator access to test a theory.

Read the earliest CREATE_FAILED or UPDATE_FAILED; later Resource creation cancelled is usually consequence. UPDATE_ROLLBACK_FAILED needs repair of the rollback blocker and controlled continuation. Skipping a resource can leave template and physical state inconsistent. Preserve successful/failed partial resources only with an owner, security review, cost timer, and recovery plan.

Drift, CDK, and cross-account failures

Traditional change sets may miss actual drift; use supported drift-aware evidence when appropriate. Decide whether desired, previous, or actual state is authoritative before remediation. Unsupported/write-only/immutable properties require alternate evidence.

CDK failures may arise before CloudFormation: dependency install, synth, lookup role/context, Docker bundling, asset hash/publish, bootstrap version/qualifier, file/image role, deployment role, or CloudFormation execution role. Trace assembly manifest asset to S3/ECR and target role. Never refresh context blindly during incident response.

For StackSets and cross-account delivery, verify administration/home Region, permission model, delegated call mode, target OU/account filters, Region order, operation preferences, administrator/execution role trust, organization status, artifact/KMS access, and per-instance result. One failed account is not proof every target has the same cause.

Separate control-plane failure from resource-provider failure and application failure. CloudFormation may complete while initialization or user behavior is broken, and an application alarm may roll back technically valid infrastructure. Correlate stack events with service events, logs, metrics, configuration history, release identity, and synthetic transactions. Preserve the failed physical resource when forensics require it, but isolate exposure and assign expiry/cost ownership.

Read-only commands and lab

aws cloudformation describe-stack-events --stack-name STACK --region ap-south-1
aws cloudformation describe-stack-resource --stack-name STACK --logical-resource-id LOGICAL_ID --region ap-south-1
aws cloudformation describe-change-set --stack-name STACK --change-set-name CHANGE --region ap-south-1
aws cloudformation describe-stack-drift-detection-status --stack-drift-detection-id ID --region ap-south-1
aws cloudformation list-stack-set-operation-results --stack-set-name SET --operation-id OP --region ap-south-1
aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=AssumeRole --region ap-south-1

Diagnose eight supplied incidents: hidden transform error; circular cross-stack contract; KMS denial despite S3 allow; quota causes rollback; custom-resource callback missing; emergency drift would be overwritten; CDK asset role/qualifier mismatch; and delegated StackSet targets wrong OU. Produce timeline, causal event, policy/dependency proof, alternatives ruled out, smallest safe repair, retry/rollback decision, and verification.

Expand to 22 cases including name collision, unsupported Region, replacement, export in use, nested failure, macro permission, service role changed, rollback trigger, retained orphan, lookup stale, Docker unavailable, cross-account key, failure tolerance, suspended account, and cleanup dependency.

Cost and acceptance

Failures incur partial resources, retained data, build minutes, logs, NAT/data transfer, duplicate assets, snapshots, cross-Region copies, and engineer/customer time. This lesson creates nothing; a live sandbox must set cleanup timer and preserve audit evidence.

Submit eight full incident reports, 22-case matrix, dependency graph, one complete authorization evaluation, drift decision, bootstrap/asset chain, StackSet target proof, safe recovery commands, cost, and prevention backlog. Pass requires first-cause evidence, no blind retry, no privilege broadening, and verified final desired plus runtime state.

Official sources

Advertisement