Lesson 244 · AWS Learning Path

AWS 244: Diagnose and recover a deliberately broken application

· Published · 9 min read

Labelled process diagram for AWS 244: User-visible failure to Layered control and behavior evidence to First causal fault and bounded repair to Recovered transaction and cleaned lab, with decision, proof and...

Why this lesson matters

Real incidents cross service boundaries. “ALB 503” might be caused by a dead process, bad health path, full disk, denied KMS decrypt, failed dependency, or deployment change. Restarting can erase evidence, temporarily hide the fault, or increase damage.

This capstone combines AWS231–AWS243. You receive symptoms and evidence, not a repair recipe. Your job is to restore the user transaction safely, prove why it failed, reject unrelated changes, and prevent recurrence.

Outcomes

You can:

  • establish impact, severity, scope, ownership, timeline, and last known good state;
  • separate detection, symptom, causal fault, contributing factor, and recovery action;
  • triage compute, Linux, storage, network, DNS, load balancer, identity, encryption, telemetry, and dependencies;
  • choose read-only evidence before mutation;
  • make one reversible, approved change with rollback;
  • prove recovery with positive, negative, restart/replacement, and observability tests;
  • conduct a blameless root-cause review and turn learning into controls;
  • prove all lab resources and costs are cleaned up.

Safety and authorization

  • The supplied-evidence track is complete and creates nothing.
  • A live version is permitted only in an instructor-owned sandbox with a recorded fault manifest, rollback snapshot/template, budget alarm, owner approval, hard stop time, and cleanup checker.
  • Never inject faults into staging shared by others or production.
  • Do not grant broad IAM, open 0.0.0.0/0, disable encryption, delete logs, reboot, terminate, detach volumes, or replace infrastructure before evidence.
  • Preserve logs, request IDs, timestamps, hashes, and original configuration. Redact account IDs, IPs, domains, secrets, tokens, payloads, and customer data.
  • Stop and escalate if impact exceeds the declared sandbox boundary.

How incident practice evolved

Traditional operations often repaired a long-lived server in place. Cloud operations increasingly replace unhealthy components from versioned images/templates and analyze failed components out of band. DevOps and SRE practices add measurable service objectives, automation, immutable deployment, and blameless learning.

AWS Well-Architected guidance distinguishes:

  • runbook: documented procedure for a known operational task;
  • playbook: investigation pattern for an event whose cause is not yet known;
  • incident management: coordinated detection, response, communication, recovery, and follow-up;
  • problem management/RCA: removal of systemic causes after service restoration.

AWS244 is the playbook exercise. AWS245 converts the proven recovery into a tested runbook. Diagnosis must come first; automating an unproven repair scales mistakes.

Incident vocabulary

TermMeaning
SignalRaw metric, log, event, trace, or user report
AlertEvaluated signal requiring attention
IncidentUnplanned service degradation requiring coordination
ImpactUser/business consequence, not component color
SymptomObservable consequence such as HTTP 503
Primary causeEarliest actionable condition sufficient to explain impact
Contributing factorIncreased likelihood, duration, or blast radius
RecoveryRestore acceptable service
ResolutionComplete approved corrective work and close incident
PreventionReduce recurrence/detection/recovery time

Do not confuse “first event in a query” with cause; coverage may start late. Do not call every correlated change causal.

The incident lifecycle

1. Detect and declare

Record:

  • incident ID, UTC start/detection/acknowledgment times;
  • affected user journey, Regions/accounts/tenants, and error rate/latency;
  • severity based on business impact;
  • incident commander, operations lead, communications lead, and service owner;
  • current containment boundary and next update time.

Freeze unrelated changes. Preserve evidence. A dashboard screenshot without time zone, query, and source is not durable evidence.

2. Build an evidence-based hypothesis

Write the expected transaction:

client -> DNS -> ALB listener/rule -> target SG/port/process
       -> local filesystem/config -> IAM/KMS/secret
       -> database/dependency -> response/log/metric

For each boundary state:

  • expected input/output;
  • owner;
  • current safe evidence;
  • hypothesis;
  • prediction that would be observed if true;
  • cheapest read-only test that can falsify it.

Prefer falsification. “Maybe networking” is not testable; “the target SG rejects ALB health checks on TCP 8080, so Flow Logs should show REJECT on the target ENI” is.

3. Contain

Containment limits harm before final repair: stop a broken rollout, remove a target from traffic, fail over, throttle, revoke a compromised credential, or disable a harmful automation. It requires approval and rollback.

Containment is not root cause. If traffic shifts to a healthy group, preserve failed-node evidence and continue analysis out of band.

4. Recover

Choose the smallest change that addresses proven cause. Define:

  • exact resource/configuration;
  • desired state and owner source;
  • blast radius and dependencies;
  • backup/snapshot/version;
  • command/change set;
  • stop condition;
  • rollback trigger/action;
  • observer and communication.

Change one variable where practical. Record request/change ID and UTC time. Do not silently make console drift; update the version-controlled owner source.

5. Verify and close

Recovery requires:

  1. Positive test: a new user transaction succeeds.
  2. Negative test: unauthorized/invalid access still fails.
  3. Component test: process, dependency, target, and alarms are healthy.
  4. Replacement test: restart/redeployment/new instance still works; otherwise repair may be ephemeral.
  5. Load/time test: service remains within SLO for an observation window.
  6. Telemetry test: expected logs/metrics/traces and alarm recovery are present.
  7. Cleanup test: temporary access, overrides, resources, and data are gone.

“Alarm OK” alone can mean missing data. “Target healthy” alone does not prove a real user transaction.

Layer-by-layer triage

User, DNS, and TLS

Compare synthetic/user evidence by location, resolver answer, certificate SAN/expiry/chain, TLS handshake, and HTTP status. Preserve request/trace IDs. A DNS answer proves only name resolution.

Load balancer

Check LB state/scheme/subnets, listener, rule priority/action, target group, registered target/AZ, health configuration, and reason code.

  • Target.ResponseCodeMismatch: target answered outside matcher.
  • Target.Timeout: connection or response did not complete in time.
  • Target.NotInUse: target group/listener/AZ relationship.
  • ELB-generated 5xx metrics/status differ from target-generated 5xx.

Use access and health-check logs. Do not widen matcher to 200–599 merely to show green.

Routing and filtering

Walk DNS-selected address, longest-prefix route, endpoint/NAT/peer/TGW, SG, both NACL directions, firewall, and return path. Flow Log ACCEPT is not application success; Reachability Analyzer models configuration, not runtime.

EC2 infrastructure

Separate:

  • system status: AWS host/platform reachability;
  • instance status: guest OS networking/boot/health;
  • attached EBS status: volume I/O impairment.

Inspect state reason, status checks, scheduled events, console output/system log, metrics, and Auto Scaling activity before rebooting. Console output is buffered and not a continuous application log. Nitro serial console can help approved recovery when normal network access is unavailable.

Systems Manager

PingStatus=Online proves SSM Agent recently communicated, not that the app is healthy. TargetNotConnected can mean wrong account/Region, agent/profile/endpoints/proxy/resources - not a shell command error.

Session Manager is interactive and harder to reproduce; Run Command supplies document/version, parameters, invocation/plugin status, output, and exit code. Aggregate command success may coexist with a failed target depending on thresholds. Scripts must return meaningful nonzero exits.

Linux process and operating system

Use read-only checks before mutation:

date -u
uptime
systemctl status nw-app --no-pager
journalctl -u nw-app --since '30 minutes ago' --no-pager
ss -lntp
df -hT
df -ih
free -m
ps -eo pid,ppid,stat,%cpu,%mem,cmd --sort=-%cpu
dmesg --level=err,warn
curl -fsS --connect-timeout 3 http://127.0.0.1:8080/health

Check process state, bind address/port, exit status, logs, CPU/memory, OOM, filesystem blocks and inodes, mounts/read-only state, file ownership/mode, SELinux, host firewall, time, DNS, certificates, and dependency connectivity.

Never run rm -rf, truncate logs, recursively chmod/chown, or kill unknown processes as a generic repair.

Storage and configuration

Differentiate root/EBS status, attachment/device mapping, filesystem mount, free blocks/inodes, I/O errors, configuration version, and secret/parameter metadata. A full /var can stop logs, package managers, PID files, sockets, or databases.

Preserve the failed configuration and hash. Compare with deployment artifact/Parameter Store/secret version without printing protected values.

Identity, KMS, and dependencies

Identify the runtime role/session. Evaluate SCP/RCP, boundary, session, IAM/resource/key/endpoint policies, conditions, key state, and service grants. Do not use the console operator as proof of workload permission.

For dependencies, independently test DNS, TCP/TLS, authentication, protocol, timeout, quota/throttle, dependency health, and application pool/circuit-breaker behavior.

Telemetry and change history

Align CloudWatch metrics/logs/alarms, ALB logs, Flow Logs, CloudTrail, Config timeline, deployment events, Auto Scaling, SSM, and application trace on UTC.

CloudTrail proves supported API calls within configured history/trails; it is not an application audit log. Config records supported resource configuration, not every packet or process change. Missing evidence may mean no coverage, delayed delivery, retention expiry, wrong Region/account, permissions, or query error.

Safe AWS inventory commands

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text

aws ec2 describe-instance-status --instance-ids i-REDACTED --include-all-instances
aws ec2 get-console-output --instance-id i-REDACTED --latest
aws ec2 describe-volumes --filters Name=attachment.instance-id,Values=i-REDACTED
aws ssm describe-instance-information --filters Key=InstanceIds,Values=i-REDACTED
aws elbv2 describe-target-health --target-group-arn replace-owned-arn
aws cloudwatch describe-alarms --state-value ALARM
aws cloudtrail lookup-events --start-time START --end-time END --max-results 50

Console output and events can contain secrets or user data; inspect privately and extract only redacted facts. Do not use unfiltered account-wide output as course evidence.

Supplied incident: three independent faults

Download:

The case contains three faults. Do not stop at the first recovery: prove the complete transaction and identify all independent and contributing failures. Do not read the answer direction until completing hypotheses.

Root-cause review

A useful RCA states:

  • impact and duration;
  • expected versus actual architecture;
  • detection and response timeline;
  • primary cause and mechanism;
  • contributing technical/organizational factors;
  • why existing controls did not prevent/detect/recover sooner;
  • recovery and verification;
  • corrective actions with owner/date/evidence;
  • lessons and what went well.

Avoid “human error.” Ask why the system allowed one action to create impact: missing test, unsafe default, excess privilege, absent deployment guard, weak SLO, or unclear ownership.

Track actions by prevention layer:

  • eliminate cause;
  • reduce blast radius;
  • detect sooner;
  • diagnose faster;
  • recover automatically/safely;
  • improve communication.

Cost and cleanup

For an optional live sandbox, cost can include EC2 runtime/public IPv4/EBS/snapshots, ALB capacity, NAT/endpoints, logs/metrics/queries, Route 53, KMS/secrets, and data transfer. Stopped EC2 still leaves EBS, snapshots, EIPs, logs, and other resources billable.

Cleanup requires tag-based inventory plus service-specific checks, not only “stack deleted.” Preserve approved evidence before deleting lab telemetry, then verify zero temporary IAM grants, SG rules, DNS records, targets, instances, volumes/snapshots, EIPs, LBs, NAT/endpoints, logs, alarms, secrets/parameters, and stacks.

Knowledge check

  1. Distinguish signal, symptom, primary cause, contributing factor, and recovery.
  2. Why should a hypothesis include a falsifying prediction?
  3. When is replacement better than in-place repair?
  4. What does each EC2 status-check category indicate?
  5. Why can Run Command show aggregate success with a failed target?
  6. Why check inodes as well as disk blocks?
  7. Why can ALB health recovery still fail acceptance?
  8. What does CloudTrail or Config not prove?
  9. Which tests prove security and restart durability?
  10. Why must temporary containment be removed?
  11. What makes an RCA blameless but accountable?
  12. What proves cleanup?

Lesson acceptance

A passing submission includes impact/severity/roles; UTC timeline; expected transaction; at least three falsifiable hypotheses; first causal fault and contributing factors; read-only evidence across every relevant layer; one-change recovery plan and rollback; positive, negative, replacement, SLO, telemetry, and cleanup tests; RCA/action owners; cost evidence; no protected data; and explicit production-untouched proof.

Official sources

Advertisement