# AWS244 Supplied Broken Application Incident

All identifiers, names, and addresses are fictional/redacted. Times are UTC. Do not execute changes.

## User report

At 10:03, checkout begins returning mostly HTTP 503; after an apparent recovery, about 10% of requests still fail. The SLO is 99.9% successful checkouts over 30 days.

Expected path:

```text
checkout.example.invalid -> Route 53 weighted alias
  -> primary ALB -> orders target group -> EC2 service :8080
  -> SSM SecureString metadata -> KMS -> database -> 200
```

## Timeline and evidence

### Change events

```text
09:55 deployment nw-orders-2026.09.15 completed
09:56 PutKeyPolicy by pipeline role, ticket CHG-771
09:57 ModifyTargetGroup health path /ready -> /healthz, ticket CHG-771
09:58 Route53 ChangeResourceRecordSets by old migration automation, no ticket
10:01 target health transitions unhealthy
10:03 first alarm and user report
```

### EC2 and SSM

```text
InstanceState=running
SystemStatus=ok InstanceStatus=ok AttachedEbsStatus=ok
SSM PingStatus=Online
CPUUtilization=4% memory=31% df blocks=42% df inodes=18%
console output: normal boot; no kernel/OOM/I/O errors
```

Interpret what this proves and what it does not.

### Application

```text
systemctl: nw-orders active (running), restart count 23
ss: LISTEN 0 511 0.0.0.0:8080
journal:
09:59:12 dependency initialization failed:
AccessDeniedException kms:Decrypt; encryption context PARAMETER_ARN=/prod/orders/db
09:59:13 entering degraded mode; /ready returns 503
```

IAM allows `ssm:GetParameter` and `kms:Decrypt` on exact resources. KMS key policy allows the runtime role but its condition was changed to:

```text
kms:EncryptionContext:PARAMETER_ARN = /prod/payments/*
```

The owner source before deployment used `/prod/orders/*`.

### ALB

```text
listener/rule/target registration/SG/NACL/route: expected
health check: HTTP :8080 /healthz matcher 200
target reason: Target.ResponseCodeMismatch, observed 404
application endpoints:
  /ready -> dependency-aware 200 or 503
  /live  -> process-only 200
  /healthz -> no route, 404
ALB access metric: HTTPCode_ELB_503_Count increased
```

Explain the difference between the dependency fault and health-check configuration fault. Decide the intended readiness contract from owner evidence rather than making every response healthy.

### DNS after first two repairs are modeled

```text
Route 53 zone has weighted aliases:
  primary.example-alb.invalid Weight=90 EvaluateTargetHealth=true
  legacy.example-alb.invalid  Weight=10 EvaluateTargetHealth=false
legacy ALB was retained for rollback but has no healthy targets.
Resolver sampling after TTL expiry returns both answers according to weighting.
Migration repository says legacy weight should be 0 after acceptance.
```

This explains the residual approximate 10% failure. It was hidden while the primary path failed.

### Telemetry limitations

- Flow Logs show `ACCEPT` in both directions for ALB-to-target flows.
- Reachability Analyzer currently reports the path reachable.
- CloudTrail Event History contains the listed management changes.
- Config supports the target group but does not record the Linux process or parameter plaintext.
- Application logs do not contain secret values.

## Required work

1. Declare impact, roles, timeline, and last known good.
2. Write at least three falsifiable hypotheses before reading all evidence.
3. Classify the KMS context, target health path, and weighted DNS record as three faults; identify primary and residual impact.
4. Reject broad KMS permission, matcher 200–599, DNS deletion, and reboot as fixes.
5. Sequence controlled recovery with owner source, rollback, and propagation/TTL checks.
6. Prove new checkout, denied unauthorized decrypt, restart/replacement, target health, weighted DNS convergence, SLO, alarm/logs, and cleanup.
7. Write RCA and prevention actions for policy tests, target-group contract tests, DNS deployment ownership, and synthetic weighted-path testing.

## Answer-direction lock

Read only after completing your worksheet:

- First primary cause: key-policy encryption-context path changed to another application, so dependency initialization failed.
- Independent masking fault: target group checked a nonexistent path instead of the documented readiness endpoint.
- Residual fault: stale weighted alias still sent traffic to an unhealthy legacy ALB.
- Green EC2/SSM/network evidence correctly rejects infrastructure reachability as the primary cause.
