# P11 troubleshooting fault casebook

Use with AWS231. All identifiers are fictional. Never run a repair from a fault
card; diagnose from the supplied evidence, state what additional read-only query
you need, and write a controlled repair/verification plan.

For every case submit: scope, UTC timeline, first causal event, cascading
symptoms, owner boundary, smallest repair, retry boundary, success evidence,
regression control, cleanup, and one tempting but unsafe response.

## Case 1: template rejected before a stack operation

```text
2026-09-15T09:00:01Z ValidateTemplate
E3002 Invalid Property Resources/ApplicationLog/Properties/RetentionDay
```

Template excerpt:

```yaml
ApplicationLog:
  Type: AWS::Logs::LogGroup
  Properties:
    RetentionDay: 7
```

Expected diagnosis: schema/property spelling failure. No stack operation and no
rollback occurred. Correct to `RetentionInDays`, rerun local/static policy tests
and `validate-template`, then create and inspect a change set. IAM changes cannot
repair a schema error.

## Case 2: CloudFormation authorization failure and cascading rollback

```text
09:12:04Z AppParameter CREATE_IN_PROGRESS
09:12:06Z AppParameter CREATE_FAILED
  Resource handler returned message: User ...:role/nw-cfn-service-role is not
  authorized to perform ssm:PutParameter on resource ...:parameter/nw/p11/app
09:12:07Z nw-p11-stack ROLLBACK_IN_PROGRESS
09:12:09Z AppParameter DELETE_COMPLETE
09:12:10Z nw-p11-stack ROLLBACK_COMPLETE
```

Expected diagnosis: the stack service role—not necessarily the console caller—
lacks a scoped underlying-service action. `ROLLBACK_IN_PROGRESS` and deletion
are consequences. Review caller permission to pass/use the service role, service
role policy, permissions boundary, session policy, SCP, resource policy and
CloudTrail denial context. Add only the required action/resource/conditions;
create a new change set. Do not grant AdministratorAccess.

## Case 3: update rollback is stuck

```text
09:30:02Z Database UPDATE_FAILED: requested previous DB instance no longer exists
09:30:05Z Api UPDATE_FAILED: Resource update cancelled
09:30:30Z nw-p11-stack UPDATE_ROLLBACK_FAILED
```

Expected diagnosis: `Database` is the primary rollback failure; `Api` is a
cancelled dependent. Restore the missing prerequisite or otherwise correct the
rollback cause, then `continue-update-rollback`. Use `--resources-to-skip` only
as a last-resort owner-approved recovery and only for logical resources that
failed during rollback. A skipped resource is marked complete but remains
inconsistent with the template; reconcile it before another update.

## Case 4: drift is real but ordinary template diff is empty

```text
CDK/CloudFormation desired RetentionInDays: 7
CloudFormation drift detection: DETECTION_COMPLETE / DRIFTED
ApplicationLog MODIFIED
  /Properties/RetentionInDays expected 7 actual 14
cdk diff: There were no differences
```

Expected diagnosis: source/template did not change, actual state did. Ordinary
template diff is not complete drift detection. Decide whether 14 should be
adopted or 7 restored, encode the reviewed value in the owner source, deploy,
rerun drift detection to completion, and query the actual property.

## Case 5: managed node is not online

```text
describe-instance-information:
  InstanceId=i-0example PingStatus=ConnectionLost AgentVersion=3.x
Run Command: TargetNotConnected
Linux service: amazon-ssm-agent active (running)
agent log: dial tcp: lookup ssmmessages.ap-south-1.amazonaws.com: no such host
```

Expected diagnosis: command document content never ran. The evidence points to
DNS/endpoint reachability, not a shell-script bug. Check VPC DNS attributes,
resolver path, route/NAT or interface endpoints, endpoint private DNS, endpoint
security-group inbound TCP 443, node egress TCP 443, proxy, TLS/clock, and
credentials. Run local `ssm-cli get-diagnostics` where supported. Do not open SSH
to the internet as an SSM repair.

## Case 6: patch command races and baseline assumptions

```text
Command A and Command B target the same node at 10:00Z
plugin output: No such file or directory patch-baseline-operations-*.tar.gz
df -h /var: 97%
latest compliance BaselineId=pb-0old ExecutionType=Command
operator expected PatchPolicy baseline pb-0new
```

Expected diagnosis: concurrent patch operations and nearly full `/var` can
explain package extraction failure. The compliance record also proves the old
baseline/Command path, not the assumed patch policy. Stop overlapping windows,
associations or commands; safely free space; verify registration/OS applicability,
baseline ID, approval rules, repository access and one operation's snapshot ID.
Run `Scan` first, inspect per-node/plugin output, then owner-approved `Install`.

## Case 7: aggregate command success hides a failed target

```text
Command.Status=Success TargetCount=10 CompletedCount=10 ErrorCount=1
node-07 Invocation.Status=Failed ResponseCode=1
node-07 plugin stderr: /opt/nw/config: Permission denied
```

Expected diagnosis: aggregate `Success` can describe delivery/error-threshold
semantics and is not proof every invocation/plugin succeeded. Query all
invocations with details; repair ownership/permissions only on the proven node;
rerun a canary and then the failed target, not all ten nodes.

## Case 8: Automation started, then failed in one step

```text
AutomationExecutionStatus=Failed
StepName=ChangeInstanceType Action=aws:executeAwsApi StepStatus=Failed
FailureMessage=InvalidAutomationExecutionParametersException: InstanceType value
AutomationAssumeRole=arn:...:role/nw-p11-automation
```

Expected diagnosis: start permission succeeded; the failure is step input or
runbook contract, not `ssm:StartAutomationExecution`. Inspect exact document
version, parameters/types/allowed patterns, resolved outputs, Automation role,
API request and preconditions. Correct source/input, create a new execution, and
verify resulting EC2 state. Completed earlier steps are not automatically rolled
back; run explicit compensation if the runbook defines and permits it.
