Lesson 392 · AWS Learning Path

AWS 392: Observability requirements, service-level indicators, objectives, and error budgets

· Published · 4 min read

Labelled process diagram for AWS 392: Versioned intent to Automated validation to Controlled AWS change to Observed result and retained evidence, with decision, proof and rejection evidence.

Why this lesson matters

Observability starts with decisions and user outcomes, not collecting everything. SLIs quantify a critical journey; SLOs set an owned reliability target over a window; error budgets make reliability tradeoffs explicit; burn rates detect dangerous consumption. None should promise impossible perfection or hide missing traffic.

From user journey to SLI

JourneyGood eventTotal eventImportant segmentation
CheckoutValid response within 800 ms and correct outcomeEligible checkout attemptsRegion, release, payment dependency
Async orderCompleted exactly once within 5 minutesAccepted valid ordersqueue age, retry, tenant
LoginAuthorized result within thresholdEligible auth attemptsidentity dependency, client type
Data freshnessRead serves data no older than 10 minutesEligible reads/periodssource, Region, pipeline version

Define observation point close to the user, event eligibility, success/latency threshold, timeout/cancel/retry/duplicate treatment, dependency errors, maintenance, low/no traffic, bot/test traffic, late data, and telemetry loss. Server 5xx alone misses incorrect 200s and client-visible latency. Avoid high-cardinality dimensions containing user IDs.

Request-based SLI is good eligible requests / total eligible requests. Period-based SLI classifies fixed time periods as good/bad. They answer different questions: one weights volume; the other weights time. CloudWatch Application Signals supports request-based and period-based SLOs using discovered latency/availability or a CloudWatch metric/expression, subject to current support and ingestion delay.

SLO and error-budget math

For a 99.9 percent 30-day availability objective, allowed bad fraction is 0.001, about 43.2 minutes if using time equivalence. For 10,000,000 eligible requests, request budget is 10,000 bad requests. Show exact calculation and rounding. An SLO is stricter than user expectation only when the service can measure and operate it; historical performance recommendations are a starting point, not the business objective.

Burn rate is current budget-consumption rate divided by the sustainable rate. Burn rate 1 exhausts the budget exactly over the objective window if sustained. Pair a short window for fast detection with a longer window to reduce noise; use different thresholds for paging and ticketing. Test low-volume and missing-data behavior.

Error-budget policy states responses as remaining budget and burn change: pause risky releases, require canary/recovery tests, prioritize reliability, tighten approval, or continue normal delivery. It is not permission to cause avoidable outages and should not punish teams for telemetry defects. Dependencies and shared platforms need aligned but separately owned objectives.

Assign a service owner, business approver, telemetry owner and review cadence. Publish the exact query/version and source-retention requirement. Recalculate after material architecture or journey change, but never rewrite history silently. Compare SLO results with complaints, support tickets, revenue/transaction outcomes and incident reviews to find blind spots. A consistently overachieved objective may justify investment elsewhere; a permanently breached objective needs architecture or expectation change, not normalization of failure.

Exclusions, ownership, and anti-patterns

Exclusions require reason, approver, exact window/events, customer impact, expiry, and audit. Do not automatically exclude maintenance, dependency failure, or “known issue” when users were affected. CloudWatch SLO time-window exclusions must reflect the approved policy rather than retroactively improving attainment.

Common failures: measuring infrastructure uptime instead of journey; averages hiding tail latency; denominator dropping failed requests; retries counted as new successes; no-traffic shown as perfect; objective copied from an SLA; 100 percent target; too many page-generating SLOs; and changing query during an incident. Version the SLI query and backtest changes.

Workshop

Using supplied 30-day checkout events, define availability, latency, correctness, and freshness candidates. Choose two SLIs, calculate attainment/budget, simulate one 20-minute outage and a slow burn, design multi-window alerts, and write an error-budget policy. Reconcile raw counts to CloudWatch metric math or Application Signals evidence.

aws applicationsignals list-service-level-objectives --region ap-south-1
aws applicationsignals get-service-level-objective --id SLO_ID --region ap-south-1
aws applicationsignals batch-get-service-level-objective-budget-report --timestamp START --slo-ids SLO_ID --region ap-south-1
aws cloudwatch get-metric-data --metric-data-queries file://queries.json --start-time START --end-time END --region ap-south-1

Analyze 16 traps: wrong observation point, missing denominator, retries, duplicate events, clock skew, late logs, low traffic, missing data, percentile aggregation, cardinality explosion, region aggregate hides failure, dependency exclusion, maintenance abuse, budget rounding, alert flapping, and SLO query changed without versioning.

Cost and acceptance

Price custom/high-resolution metrics, Application Signals/transaction search, logs, traces, canaries/RUM, alarms/SNS, cross-account observability, dashboards, and query scanning. This lesson creates nothing.

Submit journey map, SLI specification/query, calculation workbook, two SLOs, burn alerts, exclusion rules, error-budget policy, sixteen traps, ownership/review cadence, and cost. Pass requires auditable good/total events, user-centered thresholds, tested no/missing-traffic semantics, and reliability decisions linked to budget.

Official sources

Advertisement