AWS 206: CloudWatch alarms and EventBridge actions
Why this lesson matters
An alarm is an executable interpretation of telemetry. A poor alarm pages on noise, hides missing data, repeats without context or triggers unsafe automation. A useful alarm states the user harm, exact series/query, evaluation window, missing-data rule, owner and tested action path. EventBridge can route every state-change event, but routing is not proof that the destination received or completed the work.
What you will be able to do
By the end, you can:
- distinguish metric, log and composite alarms and their permitted actions;
- calculate period, evaluation periods and datapoints-to-alarm as an M-of-N rule;
- choose
breaching,notBreaching,ignoreormissingfrom publisher semantics; - explain
OK,ALARMandINSUFFICIENT_DATAtransitions without treating state as resource health; - trace an alarm event through EventBridge pattern, target role, retry policy and dead-letter queue;
- design idempotent, bounded automation with loop prevention and rollback;
- test alarm and recovery paths while preserving timestamps and notification evidence.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Build alarms from an explicit failure condition, then route state changes to actions without creating loops or destructive automation. |
| Scope and boundary | For CloudWatch alarms and EventBridge actions, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup. |
| Evidence of success | Success requires matching Console, command-line, behavior, monitoring, and owner evidence for CloudWatch alarms and EventBridge actions. One green status is not enough. |
| Cost model | Requests, running capacity, stored telemetry, retained history, data transfer, and connected resources must be priced for the exact design. |
| Safe rejection rule | Avoid alarming on every noisy datapoint, treating missing as good by default, or sending remediation events back into an uncontrolled loop. |
How the request flows
+----------------------+
| Metric datapoints |
+----------------------+
|
v
+----------------------+
| Alarm evaluation |
+----------------------+
|
v
+---------------------------------+
| EventBridge alarm-state event |
+---------------------------------+
|
v
+------------------------------------------+
| Notification or controlled remediation |
+------------------------------------------+
Alarm evaluation, deeply
A metric alarm evaluates one metric, metric-math expression, Metrics Insights query or supported anomaly model. A log alarm evaluates Logs Insights results directly; a metric filter plus metric alarm is a separate two-stage design. A composite alarm evaluates Boolean combinations of other alarm states and reduces notification noise. Composite alarms can notify and create operational work such as OpsItems/incidents, but they cannot perform EC2 or Auto Scaling actions.
For a 2-of-3 alarm with a five-minute period, CloudWatch evaluates whether at least two of the latest three evaluation periods breach. That describes up to a 15-minute evaluation window, but detection latency also includes sample publication, ingestion and evaluation timing. “Two consecutive points” is different from “two of three.” Document the exact M and N.
Missing data must follow publisher behavior:
| Publisher behavior | Candidate treatment | Risk to discuss |
|---|---|---|
| continuous heartbeat/availability metric | breaching may expose silence | ingestion outage can page even if workload is healthy |
| sparse metric emitted only on error | notBreaching | zero and absent remain semantically different |
| temporary gaps should preserve state | ignore | stale ALARM or OK can persist |
| absence is unknown | missing | all missing can produce INSUFFICIENT_DATA |
The default is missing. CloudWatch uses real datapoints when enough are available, so the configured missing-data behavior is not necessarily applied on every evaluation. For destructive EC2 alarm actions, AWS recommends caution and triggering only from ALARM, not from a transient insufficient-data condition.
Alarm actions and EventBridge routing overlap but are not identical. CloudWatch sends alarm state-change events to EventBridge. The event contains previous/current state and configuration context. Match the narrow alarm ARN/name and target state; route to a target with a least-privilege role, finite retries, DLQ and correlation ID. An automated repair must check current state, mark its own attempt, refuse unknown resources, limit concurrency and verify postconditions so its corrective API event cannot trigger an infinite loop.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use alarms for actionable conditions with a named owner and tested response. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid alarming on every noisy datapoint, treating missing as good by default, or sending remediation events back into an uncontrolled loop. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Open CloudWatch > Alarms > All alarms and select an approved alarm. Record state, state-update UTC time, metric/query, dimensions, statistic, period, M/N, threshold and missing-data setting.
- Use View in metrics with the same UTC interval. Explain every breaching, non-breaching and missing point used by the last transition; do not infer from graph color alone.
- Inspect alarm actions separately for
ALARM,OKandINSUFFICIENT_DATA. Confirm whether actions are enabled and whether the target is SNS, EC2, Auto Scaling, Systems Manager or another supported destination. - Open EventBridge > Rules and inspect only the rule associated with the alarm. Read the event pattern, bus, state, target, target role, retry policy and DLQ.
- Inspect the target's delivery/invocation evidence and DLQ depth. A successful rule match does not prove target completion.
- On this read-only track, use a supplied state-change event for testing; do not use
set-alarm-state, create a rule or invoke remediation.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws cloudwatch describe-alarms --query 'MetricAlarms[].{Name:AlarmName,State:StateValue,Metric:MetricName,Period:Period,Eval:EvaluationPeriods,Missing:TreatMissingData}' --output table
aws events list-rules --query 'Rules[].{Name:Name,State:State,Pattern:EventPattern}' --output table
aws cloudwatch describe-alarm-history --alarm-name replace-with-approved-alarm \
--history-item-type StateUpdate --max-records 20 --output table
aws events list-targets-by-rule --rule replace-with-approved-rule --output json
Expected interpretation
describe-alarms shows current configuration/state, while alarm history proves transitions and reasons. EventBridge list-rules can truncate the embedded pattern in table output, so inspect the exact rule in JSON. list-targets-by-rule proves configured destinations - not successful delivery, target execution or remediation. Correlate alarm history, EventBridge metrics, target logs/events and DLQ evidence by UTC time.
Practical work
Design an alarm for sustained ALB 5xx errors. Set metric math or dimensions, threshold, period, datapoints, missing data, alarm and OK actions, EventBridge pattern, target retry, DLQ, owner, test, and rollback.
Diagnose this topic from its own evidence
| Symptom | Likely boundary | Proof |
|---|---|---|
| graph breaches but alarm stays OK | dimensions/statistic/period differ or M-of-N not met | compare alarm configuration to exact datapoints |
alarm is INSUFFICIENT_DATA | new alarm, no publisher, wrong dimensions or all points missing | last metric timestamp and alarm history reason |
| alarm changed but no notification | actions disabled, SNS policy/subscription or destination failure | alarm action config, SNS delivery and subscription state |
| EventBridge rule did not match | event source/detail-type/name/state pattern mismatch | archive/supplied event against exact pattern |
| remediation repeats | event loop or non-idempotent target | correlated CloudTrail and invocation IDs; stop condition |
Positive test: supplied breaching points produce the predicted 2-of-3 transition and one routed event. Negative test: one spike among three remains below the alarm gate. Dependency-failure test: alarm transitions but destination rejects the target role; route the failed event to the DLQ and preserve it before repair.
Cost and cleanup
Price standard/high-resolution metric alarms, composite/log alarms, underlying custom metrics or Logs Insights evaluation, EventBridge events, target invocations, SNS delivery, DLQ storage and retained logs. Noise is an operational cost even when API cost is small. Delete only course-owned alarms/rules after evidence export; disabling actions is not deletion, and deleting an alarm can break a composite alarm or automation dependency.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Build alarms from an explicit failure condition, then route state changes to actions without creating loops or destructive automation.
- Which scope or ownership boundary must be proved first?
Expected direction: For CloudWatch alarms and EventBridge actions, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup.
- What evidence is strong enough to accept the result?
Expected direction: Success requires matching Console, command-line, behavior, monitoring, and owner evidence for CloudWatch alarms and EventBridge actions. One green status is not enough.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid alarming on every noisy datapoint, treating missing as good by default, or sending remediation events back into an uncontrolled loop.
- Which cost dimensions and retained resources need an owner?
Expected direction: Requests, running capacity, stored telemetry, retained history, data transfer, and connected resources must be priced for the exact design.
Lesson acceptance
- The learner calculates and explains one M-of-N evaluation from timestamped datapoints.
- Missing data is selected from publisher semantics and a counterexample is documented.
- Alarm history, metric graph, EventBridge rule/target and destination evidence form one UTC timeline.
- Positive, single-spike negative and failed-target dependency tests have predicted and observed outcomes.
- Automation includes scope check, idempotency token/marker, finite retry, DLQ, stop condition, rollback and postcondition verification.