AWS 393: Metrics, logs, traces, deployments, and configuration changes in one timeline
Why this lesson matters
Incidents cross telemetry and change systems. Metrics show scope and timing, logs expose events, traces connect distributed requests, deployments/configuration explain candidate causes, CloudTrail identifies API actors, and tickets record decisions. Correlation is not causation; a defensible UTC timeline preserves identity, source and uncertainty.
Evidence model
| Evidence | Best question | Identity/time caveat |
|---|---|---|
| Metric | When and how broadly did behavior change? | aggregation, period, dimensions, ingestion delay |
| Structured log | What did one component observe/decide? | host clock, sampling, missing fields, retention |
| Trace/span | Which request path/dependency was slow/faulted? | sampling and propagation gaps |
| Deployment | Which immutable release reached which target? | status may precede user stability |
| Config/IaC | What desired/live control changed? | propagation and drift |
| CloudTrail | Which principal called which API? | event history scope/delay and data-event setup |
| Human record | Why was an action taken? | manual timestamps and retrospective bias |
Standardize UTC ISO-8601 with timezone and record source event time, ingestion time, and query time. Synchronize hosts, but never assume zero skew. Keep account, Region, environment, service, operation, resource ARN/ID, release/artifact/config/schema versions, request/trace IDs, target/container/instance, and safe tenant category. Never log credentials, tokens, secrets, full payloads, or unnecessary personal data.
Correlation design
Emit a release marker at startup and in responses/logs/traces. Propagate W3C trace context or chosen standard through HTTP, queues and events while generating new span/message IDs correctly. Preserve original event ID and idempotency key across retries. Connect async producer trace, message attributes, consumer processing, DLQ and replay without pretending they are one synchronous request.
Metrics need dimensions with bounded cardinality and statistic semantics. Logs should be JSON with timestamp, severity, service, event name, request/trace/span, release, resource, outcome, duration and error classification. Traces should capture dependency, status and safe attributes with sampling understood. Deployment pipelines, AppConfig, CloudFormation/Config and CloudTrail events should publish change markers into the operational view.
Create a telemetry ownership matrix with producer, schema/version, collection path, destination, access, retention, expected freshness, sampling, cardinality limit and loss alarm. Validate it during releases. If a service emits no logs, metrics or spans, the observability system must distinguish no traffic from broken collection. Keep an independent synthetic or ingestion heartbeat for critical paths, while avoiding a heartbeat that overwhelms or masks real user traffic.
Timeline method
- Define incident window generously and record detection plus claimed impact.
- Freeze queries and export raw evidence with source, Region, timezone and retrieval time.
- Establish user SLI and first/last known good/bad before searching causes.
- Add deployments/config/API changes, then component/dependency metrics, traces and logs.
- Normalize timestamps but retain originals and known ingestion/clock uncertainty.
- Link by exact IDs; label inference where only temporal association exists.
- Test competing hypotheses against evidence that should be present or absent.
- Record containment/recovery actions and verify user outcome plus backlog/delayed effects.
aws cloudwatch get-metric-data --metric-data-queries file://queries.json --start-time START --end-time END --region ap-south-1
aws logs start-query --log-group-names LOG_GROUP --start-time EPOCH_START --end-time EPOCH_END --query-string 'fields @timestamp, @message | sort @timestamp asc' --region ap-south-1
aws xray get-trace-summaries --start-time START --end-time END --region ap-south-1
aws cloudtrail lookup-events --start-time START --end-time END --region ap-south-1
aws configservice get-resource-config-history --resource-type TYPE --resource-id ID --earlier-time END --later-time START --region ap-south-1
Queries incur scope/cost and return limits/pagination. Preserve query text, IDs and completion status. Redact exports and restrict investigation workspaces.
Workshop and failure analysis
Given evidence for a latency incident during a canary and secret rotation, produce a minute-by-minute timeline. Determine whether deployment, rotation, database saturation, retry amplification or telemetry loss caused impact. Calculate release-specific SLI and compare traces for old/new versions. State confidence and missing evidence.
Analyze 18 traps: local timezone, host clock skew, delayed CloudTrail, log ingestion delay, dashboard aggregation, wrong metric statistic, missing dimension, cardinality cap, log truncation, trace sampling, broken propagation, async retry loses ID, release marker stale, config marker absent, same-time change assumed causal, query window too narrow, redaction removes join key, and recovery declared before queue drains.
Cost and acceptance
Price metric cardinality/resolution, log ingestion/storage/query scan, traces/spans/Application Signals, Config and CloudTrail events/data events, cross-account observability, dashboards and retention/export. This lesson creates nothing.
Submit telemetry contract, identity/retention map, reusable queries, raw-to-normalized timeline, hypothesis table, impact/recovery calculation, eighteen traps, privacy/access controls, and cost. Pass requires UTC with uncertainty, exact release/request linkage, user outcome first, explicit inference, and reproducible query evidence.