Lesson 217 · AWS Learning Path

AWS 217: Troubleshoot missing metrics, delayed logs, alarm state, and notification delivery

· Published · 10 min read

Labelled process diagram for AWS 217: Missing or wrong observation to Publisher-to-storage evidence to Evaluation or delivery boundary to One repair and changed retest, with decision, proof and rejection evidence.

Why this lesson matters

A missing graph can mean no publisher, wrong account or Region, a different dimension set, an old time window, delayed ingestion, or genuinely no data. An incorrect alarm can instead be evaluating exactly the configuration it was given. This lab teaches you to find the first broken boundary and change one variable at a time.

What you will be able to do

By the end, you can:

  • classify a symptom as publication, identity/permission, scope, ingestion, evaluation or delivery failure;
  • prove exact metric identity: account, Region, namespace, name, complete dimensions, unit, statistic and time range;
  • compare log event time, ingestion time and query time without mistaking delay for loss;
  • explain alarm state reason, M-of-N evaluation and missing-data policy from alarm history;
  • separate CloudWatch-to-SNS action invocation, SNS subscription confirmation and endpoint receipt;
  • perform reversible agent and alarm faults, one at a time, with a changed retest;
  • restore P10 and prove cleanup from an inventory rather than memory.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This lesson has an optional live path. Check current prices, obtain the account owner's approval, set a hard timer, use course tags, and complete the stated cleanup. The evidence path is a complete alternative.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeDiagnose missing metrics, late logs, incorrect alarm state, and failed notification by following publisher, permission, scope, timestamp, dimensions, evaluation, and endpoint delivery.
Scope and boundaryFor CloudWatch telemetry troubleshooting, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup.
Evidence of successSuccess requires matching Console, command-line, behavior, monitoring, and owner evidence for CloudWatch telemetry troubleshooting. One green status is not enough.
Cost modelRequests, running capacity, stored telemetry, retained history, data transfer, and connected resources must be priced for the exact design.
Safe rejection ruleAvoid changing four settings at once, publishing storms, weakening permissions broadly, or calling delayed delivery data loss without timestamps.

How the request flows

+--------------------------------+
|  Missing or wrong observation  |
+--------------------------------+
                |
                v
+---------------------------------+
|  Publisher-to-storage evidence  |
+---------------------------------+
                |
                v
+-----------------------------------+
|  Evaluation or delivery boundary  |
+-----------------------------------+
                 |
                 v
+---------------------------------+
|  One repair and changed retest  |
+---------------------------------+

Troubleshoot left to right: producer process and local file, credentials and API result, account/Region/resource identity, CloudWatch ingestion, alarm evaluation, action invocation, SNS topic/subscription, then receiving endpoint. A downstream symptom does not prove its downstream component is at fault.

Evidence map

BoundaryStrong evidenceCommon false conclusion
Metric publishersuccessful API response or running agent plus recent datapointlist-metrics entry means data is current
Metric selectionexact namespace/name/all dimensions/unit/statistic/UTC windowsame metric name means same series
Log producerlocal one-line JSON and file offset/timeno query result means event was never written
Log ingestionlog stream's last ingestion time and matching eventevent timestamp equals arrival timestamp
Alarm evaluatorStateReasonData, history, threshold, M/N and missing policyALARM means the application is definitely broken
Alarm actionalarm history reports SNS action attemptendpoint received the message
SNS deliveryconfirmed subscription and endpoint-side receipt evidencetopic exists, therefore email works

Architecture decision table

SituationDirectionReason
Requirement matchesUse the sequence to find the first broken boundary instead of recreating telemetry blindly.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid changing four settings at once, publishing storms, weakening permissions broadly, or calling delayed delivery data loss without timestamps.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Confirm account and Region, then open CloudWatch > Metrics > All metrics. Select NitWings/P10 and compare the correct OrdersValidated / Environment=lab series with a deliberately wrong dimension.
  2. Open Logs > Log groups > /nw/p10/app. Use an absolute UTC range and compare @timestamp with @ingestionTime.
  3. Open Alarms, select each nw-p10- alarm, and read History plus the current state reason. Record threshold, statistic, period, evaluation periods, datapoints to alarm and missing-data policy.
  4. Open SNS > Topics > nw-p10-operations > Subscriptions. Record protocol and Confirmed status. Check the receiving system separately; the SNS page is not endpoint receipt evidence.
  5. Open CloudTrail > Event history for configuration calls such as PutMetricAlarm or Subscribe. CloudTrail management events prove control changes, not telemetry delivery.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Resolve exact P10 identities and capture the baseline before changing anything:

stack_name="nw-p10-observability"
instance_id="$(aws cloudformation describe-stacks --stack-name "$stack_name" --query 'Stacks[0].Outputs[?OutputKey==`ManagedNodeId`].OutputValue' --output text)"
topic_arn="$(aws sns list-topics --query 'Topics[?ends_with(TopicArn, `:nw-p10-operations`)].TopicArn | [0]' --output text)"

aws cloudwatch list-metrics --namespace NitWings/P10 --output json
aws logs describe-log-streams --log-group-name /nw/p10/app \
  --order-by LastEventTime --descending --limit 10 --output json
aws cloudwatch describe-alarms --alarm-name-prefix nw-p10- --output json
aws cloudwatch describe-alarm-history \
  --alarm-name nw-p10-actionable-warning --max-records 50 --output json
test "$topic_arn" = "None" || \
  aws sns list-subscriptions-by-topic --topic-arn "$topic_arn" --output json

Expected interpretation

Save timestamps and StateReasonData, not just screenshots. describe-log-streams --order-by LastEventTime cannot also use a log-stream-name prefix, and its last-event field is eventually consistent; confirm important events with an actual filtered query.

Practical work

Resolve the following four faults in order. Restore each before starting the next.

Fault 1: the metric is present but the graph is empty

Query the same metric twice over the last 30 minutes - first with the correct dimension and then with Environment=wrong. No data for the wrong series is the expected negative result, not a publication failure.

start="$(date -u -d '30 minutes ago' +%Y-%m-%dT%H:%M:%SZ)"
end="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
for environment in lab wrong; do
  aws cloudwatch get-metric-statistics \
    --namespace NitWings/P10 --metric-name OrdersValidated \
    --dimensions Name=Environment,Value="$environment" \
    --start-time "$start" --end-time "$end" --period 60 \
    --statistics Sum SampleCount --output json
done

If both are empty, inspect publication time and widen only the UTC window. Do not republish until account, Region and complete dimension set agree.

Fault 2: local logs exist but CloudWatch is delayed

Stop only the CloudWatch Agent, append a unique fake marker while stopped, prove it exists locally, then restart with the reviewed SSM configuration:

stop_id="$(aws ssm send-command --instance-ids "$instance_id" \
  --document-name AWS-RunShellScript \
  --parameters 'commands=["sudo /opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl -m ec2 -a stop","printf '\''{\"timestamp\":\"%s\",\"level\":\"WARN\",\"service\":\"p10-demo\",\"event\":\"agent_gap_test\",\"request_id\":\"p10-gap-001\"}\\n'\'' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" | sudo tee -a /var/log/nw-p10-app.log >/dev/null","sudo tail -n 1 /var/log/nw-p10-app.log"]' \
  --query Command.CommandId --output text)"
aws ssm wait command-executed --command-id "$stop_id" --instance-id "$instance_id"

restart_id="$(aws ssm send-command --instance-ids "$instance_id" \
  --document-name AmazonCloudWatch-ManageAgent \
  --parameters 'action=configure,mode=ec2,optionalConfigurationSource=ssm,optionalConfigurationLocation=/nw/p10/agent-config,optionalRestart=yes' \
  --query Command.CommandId --output text)"
aws ssm wait command-executed --command-id "$restart_id" --instance-id "$instance_id"
aws logs filter-log-events --log-group-name /nw/p10/app \
  --filter-pattern '"p10-gap-001"' --output json

Record event time, restart time and ingestion time. A later ingestion time with the original event time proves buffering/delay. If absent, inspect command output, agent status/configuration log, file permissions and outbound HTTPS before changing IAM.

Fault 3: silence produces the wrong alarm state

Capture the full warning alarm first. Change only missing-data treatment to breaching by repeating every required alarm property, observe after a full period, then restore notBreaching. put-metric-alarm replaces the configuration; omitting properties can unintentionally remove them.

aws cloudwatch describe-alarms --alarm-names nw-p10-validation-warning --output json

aws cloudwatch put-metric-alarm --alarm-name nw-p10-validation-warning \
  --alarm-description "P10 new WARN order-validation log event" \
  --namespace NitWings/P10 --metric-name ValidationWarnings --statistic Sum \
  --period 60 --evaluation-periods 1 --datapoints-to-alarm 1 --threshold 1 \
  --comparison-operator GreaterThanOrEqualToThreshold --treat-missing-data breaching

# After recording the incorrect silence=ALARM behavior, restore:
aws cloudwatch put-metric-alarm --alarm-name nw-p10-validation-warning \
  --alarm-description "P10 new WARN order-validation log event" \
  --namespace NitWings/P10 --metric-name ValidationWarnings --statistic Sum \
  --period 60 --evaluation-periods 1 --datapoints-to-alarm 1 --threshold 1 \
  --comparison-operator GreaterThanOrEqualToThreshold --treat-missing-data notBreaching

Fault 4: alarm action ran but notification was not received

Read composite alarm history, confirm the configured topic ARN, list subscriptions, and inspect the endpoint's own evidence. PendingConfirmation is not fixed by repeatedly subscribing. The address owner must use the confirmation link. If alarm history reports authorization failure, inspect the topic policy/KMS key policy and caller context; do not grant sns:*.

aws cloudwatch describe-alarm-history \
  --alarm-name nw-p10-actionable-warning --history-item-type Action \
  --max-records 20 --output json
aws sns get-topic-attributes --topic-arn "$topic_arn" --output json
aws sns list-subscriptions-by-topic --topic-arn "$topic_arn" --output json
aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=EventName,AttributeValue=PutCompositeAlarm \
  --max-results 20 --output json

Break-fix record

For every fault capture: expected/actual UTC time, producer, caller/role, Region, namespace or group, dimensions, unit, hypothesis, one change, changed test, result and restoration. Never change multiple settings to make a red dashboard green.

Diagnose this topic from its own evidence

SymptomFirst decisive checkDo not do
Metric name listed, no pointsget-metric-statistics on exact full identity/windowassume discovery means current data
Event local, absent remotelyagent state/config/log, offset and outbound 443broaden IAM before proving denial
Remote event appears lateevent versus ingestion timestamprewrite producer timestamp
Alarm disagrees with graphalarm statistic/period/dimensions/M-of-N/missing policyuse dashboard defaults as alarm truth
Composite seems invertedevery child state and full Boolean ruledisable all alarm actions
SNS action succeeded, no receiptsubscription confirmation and endpoint evidenceclaim delivery from topic existence

Cost and cleanup

Troubleshooting can increase cost through repeated custom metric points, Logs Insights scanned bytes, retained logs, alarms/SNS and longer EC2 runtime. Query a narrow absolute window and inject one changed event per hypothesis.

Restore the warning alarm to notBreaching, confirm the agent is running, and repeat AWS216 cleanup for dashboard, composite/child alarms, detector, metric filter and topic. Then delete P10 with its runbook. Custom metric history remains until CloudWatch retention expires and cannot be manually deleted.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Diagnose missing metrics, late logs, incorrect alarm state, and failed notification by following publisher, permission, scope, timestamp, dimensions, evaluation, and endpoint delivery.

  1. Which scope or ownership boundary must be proved first?

Expected direction: For CloudWatch telemetry troubleshooting, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup.

  1. What evidence is strong enough to accept the result?

Expected direction: Success requires matching Console, command-line, behavior, monitoring, and owner evidence for CloudWatch telemetry troubleshooting. One green status is not enough.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid changing four settings at once, publishing storms, weakening permissions broadly, or calling delayed delivery data loss without timestamps.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Requests, running capacity, stored telemetry, retained history, data transfer, and connected resources must be priced for the exact design.

Lesson acceptance

  • Each fault has a baseline, one hypothesis, one controlled change, changed retest and explicit restoration.
  • Correct and wrong metric dimensions produce explainable positive/negative results.
  • Agent-stop evidence proves local write, delayed ingestion and successful restart without opening inbound access.
  • Alarm history proves how missing-data policy changed state and that the original policy was restored.
  • Alarm action, SNS topic publication, subscription confirmation and endpoint receipt are reported as separate facts.
  • Manual alarm-chain resources and the P10 stack are deleted, or an approved owner/timer for immediate continuation is recorded.

Official sources

Advertisement