Lesson 206 · AWS Learning Path

AWS 206: CloudWatch alarms and EventBridge actions

· Published · 8 min read

Labelled process diagram for AWS 206: Metric datapoints to Alarm evaluation to EventBridge alarm-state event to Notification or controlled remediation, with decision, proof and rejection evidence.

Why this lesson matters

An alarm is an executable interpretation of telemetry. A poor alarm pages on noise, hides missing data, repeats without context or triggers unsafe automation. A useful alarm states the user harm, exact series/query, evaluation window, missing-data rule, owner and tested action path. EventBridge can route every state-change event, but routing is not proof that the destination received or completed the work.

What you will be able to do

By the end, you can:

  • distinguish metric, log and composite alarms and their permitted actions;
  • calculate period, evaluation periods and datapoints-to-alarm as an M-of-N rule;
  • choose breaching, notBreaching, ignore or missing from publisher semantics;
  • explain OK, ALARM and INSUFFICIENT_DATA transitions without treating state as resource health;
  • trace an alarm event through EventBridge pattern, target role, retry policy and dead-letter queue;
  • design idempotent, bounded automation with loop prevention and rollback;
  • test alarm and recovery paths while preserving timestamps and notification evidence.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeBuild alarms from an explicit failure condition, then route state changes to actions without creating loops or destructive automation.
Scope and boundaryFor CloudWatch alarms and EventBridge actions, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup.
Evidence of successSuccess requires matching Console, command-line, behavior, monitoring, and owner evidence for CloudWatch alarms and EventBridge actions. One green status is not enough.
Cost modelRequests, running capacity, stored telemetry, retained history, data transfer, and connected resources must be priced for the exact design.
Safe rejection ruleAvoid alarming on every noisy datapoint, treating missing as good by default, or sending remediation events back into an uncontrolled loop.

How the request flows

+----------------------+
|  Metric datapoints   |
+----------------------+
           |
           v
+----------------------+
|   Alarm evaluation   |
+----------------------+
           |
           v
+---------------------------------+
|  EventBridge alarm-state event  |
+---------------------------------+
                |
                v
+------------------------------------------+
|  Notification or controlled remediation  |
+------------------------------------------+

Alarm evaluation, deeply

A metric alarm evaluates one metric, metric-math expression, Metrics Insights query or supported anomaly model. A log alarm evaluates Logs Insights results directly; a metric filter plus metric alarm is a separate two-stage design. A composite alarm evaluates Boolean combinations of other alarm states and reduces notification noise. Composite alarms can notify and create operational work such as OpsItems/incidents, but they cannot perform EC2 or Auto Scaling actions.

For a 2-of-3 alarm with a five-minute period, CloudWatch evaluates whether at least two of the latest three evaluation periods breach. That describes up to a 15-minute evaluation window, but detection latency also includes sample publication, ingestion and evaluation timing. “Two consecutive points” is different from “two of three.” Document the exact M and N.

Missing data must follow publisher behavior:

Publisher behaviorCandidate treatmentRisk to discuss
continuous heartbeat/availability metricbreaching may expose silenceingestion outage can page even if workload is healthy
sparse metric emitted only on errornotBreachingzero and absent remain semantically different
temporary gaps should preserve stateignorestale ALARM or OK can persist
absence is unknownmissingall missing can produce INSUFFICIENT_DATA

The default is missing. CloudWatch uses real datapoints when enough are available, so the configured missing-data behavior is not necessarily applied on every evaluation. For destructive EC2 alarm actions, AWS recommends caution and triggering only from ALARM, not from a transient insufficient-data condition.

Alarm actions and EventBridge routing overlap but are not identical. CloudWatch sends alarm state-change events to EventBridge. The event contains previous/current state and configuration context. Match the narrow alarm ARN/name and target state; route to a target with a least-privilege role, finite retries, DLQ and correlation ID. An automated repair must check current state, mark its own attempt, refuse unknown resources, limit concurrency and verify postconditions so its corrective API event cannot trigger an infinite loop.

Architecture decision table

SituationDirectionReason
Requirement matchesUse alarms for actionable conditions with a named owner and tested response.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid alarming on every noisy datapoint, treating missing as good by default, or sending remediation events back into an uncontrolled loop.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open CloudWatch > Alarms > All alarms and select an approved alarm. Record state, state-update UTC time, metric/query, dimensions, statistic, period, M/N, threshold and missing-data setting.
  2. Use View in metrics with the same UTC interval. Explain every breaching, non-breaching and missing point used by the last transition; do not infer from graph color alone.
  3. Inspect alarm actions separately for ALARM, OK and INSUFFICIENT_DATA. Confirm whether actions are enabled and whether the target is SNS, EC2, Auto Scaling, Systems Manager or another supported destination.
  4. Open EventBridge > Rules and inspect only the rule associated with the alarm. Read the event pattern, bus, state, target, target role, retry policy and DLQ.
  5. Inspect the target's delivery/invocation evidence and DLQ depth. A successful rule match does not prove target completion.
  6. On this read-only track, use a supplied state-change event for testing; do not use set-alarm-state, create a rule or invoke remediation.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws cloudwatch describe-alarms --query 'MetricAlarms[].{Name:AlarmName,State:StateValue,Metric:MetricName,Period:Period,Eval:EvaluationPeriods,Missing:TreatMissingData}' --output table
aws events list-rules --query 'Rules[].{Name:Name,State:State,Pattern:EventPattern}' --output table
aws cloudwatch describe-alarm-history --alarm-name replace-with-approved-alarm \
  --history-item-type StateUpdate --max-records 20 --output table
aws events list-targets-by-rule --rule replace-with-approved-rule --output json

Expected interpretation

describe-alarms shows current configuration/state, while alarm history proves transitions and reasons. EventBridge list-rules can truncate the embedded pattern in table output, so inspect the exact rule in JSON. list-targets-by-rule proves configured destinations - not successful delivery, target execution or remediation. Correlate alarm history, EventBridge metrics, target logs/events and DLQ evidence by UTC time.

Practical work

Design an alarm for sustained ALB 5xx errors. Set metric math or dimensions, threshold, period, datapoints, missing data, alarm and OK actions, EventBridge pattern, target retry, DLQ, owner, test, and rollback.

Diagnose this topic from its own evidence

SymptomLikely boundaryProof
graph breaches but alarm stays OKdimensions/statistic/period differ or M-of-N not metcompare alarm configuration to exact datapoints
alarm is INSUFFICIENT_DATAnew alarm, no publisher, wrong dimensions or all points missinglast metric timestamp and alarm history reason
alarm changed but no notificationactions disabled, SNS policy/subscription or destination failurealarm action config, SNS delivery and subscription state
EventBridge rule did not matchevent source/detail-type/name/state pattern mismatcharchive/supplied event against exact pattern
remediation repeatsevent loop or non-idempotent targetcorrelated CloudTrail and invocation IDs; stop condition

Positive test: supplied breaching points produce the predicted 2-of-3 transition and one routed event. Negative test: one spike among three remains below the alarm gate. Dependency-failure test: alarm transitions but destination rejects the target role; route the failed event to the DLQ and preserve it before repair.

Cost and cleanup

Price standard/high-resolution metric alarms, composite/log alarms, underlying custom metrics or Logs Insights evaluation, EventBridge events, target invocations, SNS delivery, DLQ storage and retained logs. Noise is an operational cost even when API cost is small. Delete only course-owned alarms/rules after evidence export; disabling actions is not deletion, and deleting an alarm can break a composite alarm or automation dependency.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Build alarms from an explicit failure condition, then route state changes to actions without creating loops or destructive automation.

  1. Which scope or ownership boundary must be proved first?

Expected direction: For CloudWatch alarms and EventBridge actions, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup.

  1. What evidence is strong enough to accept the result?

Expected direction: Success requires matching Console, command-line, behavior, monitoring, and owner evidence for CloudWatch alarms and EventBridge actions. One green status is not enough.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid alarming on every noisy datapoint, treating missing as good by default, or sending remediation events back into an uncontrolled loop.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Requests, running capacity, stored telemetry, retained history, data transfer, and connected resources must be priced for the exact design.

Lesson acceptance

  • The learner calculates and explains one M-of-N evaluation from timestamped datapoints.
  • Missing data is selected from publisher semantics and a counterexample is documented.
  • Alarm history, metric graph, EventBridge rule/target and destination evidence form one UTC timeline.
  • Positive, single-spike negative and failed-target dependency tests have predicted and observed outcomes.
  • Automation includes scope check, idempotency token/marker, finite retry, DLQ, stop condition, rollback and postcondition verification.

Official sources

Advertisement