AWS 198: CloudWatch, CloudTrail, Config, and Systems Manager evidence for an architect
Why this lesson matters
Use CloudWatch, standard CloudTrail Event History or trails, AWS Config, and Systems Manager to prove metrics, API actions, configuration history, and managed-node state.
What you will be able to do
By the end, you can:
- explain cloudwatch, cloudtrail, config, and systems manager evidence for an architect in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Use CloudWatch, standard CloudTrail Event History or trails, AWS Config, and Systems Manager to prove metrics, API actions, configuration history, and managed-node state. |
| Scope and boundary | The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Architectural operations evidence. |
| Evidence of success | Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Architectural operations evidence. |
| Cost model | Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design. |
| Safe rejection rule | Avoid teaching CloudTrail Lake to a new customer after access closure, confusing API audit with data access not logged, or claiming compliance from one Config rule. |
How the request flows
+----------------------+
| Workload symptom |
+----------------------+
|
v
+----------------------------------+
| Metric and API-change evidence |
+----------------------------------+
|
v
+-------------------------------------------+
| Configuration and managed-node evidence |
+-------------------------------------------+
|
v
+---------------------------------------------+
| Root cause, repair, and verified recovery |
+---------------------------------------------+
Four services answer different questions
| Question | Primary evidence | Boundary |
|---|---|---|
| What is the workload doing? | CloudWatch metrics, logs, traces, alarms and SLOs | Only emitted/collected telemetry with correct dimensions, retention and access exists. |
| Who called which AWS API? | CloudTrail Event history, trails and supported analytics | Event history covers 90 days of regional management events; data/network events require selectors. |
| What configuration existed/changed? | Config items, relationships, timelines and rules | Only supported recorded types/Regions; compliance is rule-specific. |
| What managed-node action ran? | SSM command, automation, association/session evidence plus CloudTrail | Node health, role, network and output destination define coverage. |
An alarm is not root cause. A CloudTrail event is not application health. Config COMPLIANT proves one rule evaluation. Run Command Success means its plugin exited successfully, not that the business repair worked.
Construct a UTC incident timeline
- Define steady state and first user-visible symptom.
- Query metrics/SLOs for onset, blast radius and recovery, retaining dimensions/missing-data behavior.
- Query logs/traces with request/correlation IDs.
- Find CloudTrail actor/session, source, API, parameters, response/error and request ID.
- Compare Config before/after items and relationships.
- Inspect SSM document version, parameters, targets, per-node output, exit code and approval.
- Test a causal hypothesis, record remediation/rollback and prove restored steady state.
Normalize UTC and distinguish event, ingestion and processing times. Redact account/credential data without removing causal fields. Protected S3 trail delivery with validation/immutability and separate log-account access strengthens evidence.
Current CloudTrail analytics boundary
CloudTrail Event history remains available for 90-day management-event lookup. Trails provide ongoing S3 delivery and optional CloudWatch Logs/EventBridge integration. Data events and network activity events are not simply included by Event history and can add cost.
As of this review, CloudTrail Lake is closed to new customers from May 31, 2026; existing customers retain supported access under AWS's documented transition. New learners must not be instructed to create Lake event data stores. Use current CloudWatch ingestion/analytics guidance or S3/Athena architecture as appropriate. Existing Lake customers should assess AWS's CloudWatch migration path and preserve historical-data requirements.
Coverage and safe SSM evidence
Organization trails, Config aggregation and CloudWatch cross-account observability centralize views, but aggregation is not collection: every source account/Region still needs recording and telemetry coverage.
Prefer SSM Run Command/Automation over SSH when it improves identity/audit/network posture, but restrict documents, parameters and tag-selected targets. Send outputs to encrypted S3/CloudWatch because console output can truncate or expire. Session Manager API events do not automatically record every shell command; configure session logging explicitly.
CloudWatch cardinality/ingestion/query/retention, CloudTrail event selectors, Config items/evaluations and SSM logs all need cost ownership.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use several evidence types to reconstruct architecture behavior and ownership. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid teaching CloudTrail Lake to a new customer after access closure, confusing API audit with data access not logged, or claiming compliance from one Config rule. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Use the Console service search and open CloudWatch Alarms, CloudTrail Event history and Trails, Config resources, and Systems Manager Fleet Manager; confirm the account and Region before reading the page.
- Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
- Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
- Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws cloudwatch describe-alarms --query 'MetricAlarms[].{Name:AlarmName,State:StateValue,Metric:MetricName,Updated:StateUpdatedTimestamp}' --output table
aws cloudtrail describe-trails --include-shadow-trails false --query 'trailList[].{Name:Name,MultiRegion:IsMultiRegionTrail,LogValidation:LogFileValidationEnabled,S3:S3BucketName}' --output table
aws cloudtrail lookup-events --max-results 20 --query 'Events[].{Time:EventTime,Name:EventName,User:Username,Source:EventSource}' --output table
aws configservice describe-configuration-recorders --output table
aws ssm describe-instance-information --query 'InstanceInformationList[].{Id:InstanceId,Ping:PingStatus,Platform:PlatformName,Version:AgentVersion}' --output table
Expected interpretation
Metrics show numerical behavior, CloudTrail shows control-plane events, Config shows resource configuration history and compliance, and Systems Manager shows managed-node operations. None alone proves the complete application.
Practical work
Investigate a supplied outage timeline. Correlate an alarm, change API event, Config relationship change, SSM node state, application log, and recovery action. State what each source proves and cannot prove.
Diagnose this topic from its own evidence
- Missing metric/log: inspect emitter/agent, namespace/dimensions, account/Region, IAM, endpoint and time range.
- Empty CloudTrail lookup: check Region, 90-day window, management versus data event and selector/trail coverage.
- Empty Config timeline: check recorder, supported type, inclusion, delivery and recording start.
- SSM success but repair fails: inspect command/output/exit semantics and independent application health.
- Timestamp mismatch: normalize UTC and separate event, ingestion, processing and dashboard-period times.
Cost and cleanup
Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Use CloudWatch, standard CloudTrail Event History or trails, AWS Config, and Systems Manager to prove metrics, API actions, configuration history, and managed-node state.
- Which scope or ownership boundary must be proved first?
Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Architectural operations evidence.
- What evidence is strong enough to accept the result?
Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Architectural operations evidence.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid teaching CloudTrail Lake to a new customer after access closure, confusing API audit with data access not logged, or claiming compliance from one Config rule.
- Which cost dimensions and retained resources need an owner?
Expected direction: Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Lesson acceptance
- Map claims to the correct CloudWatch, CloudTrail, Config or SSM evidence.
- Produce a redacted UTC timeline with event/request/correlation IDs and causal reasoning.
- State collection gaps, retention, integrity, access and cost for each source.
- Prove SSM action through document/targets/output plus independent validation.
- Explain CloudTrail Lake's new-customer closure and current alternatives without teaching unavailable onboarding.