Lesson 316 · AWS Learning Path

AWS 316: Operational Excellence pillar

· Published · 8 min read

Labelled process diagram for AWS 316: Workload and operating model to Evidence-based pillar review to Prioritized improvement action to Measured operation and learning loop, with decision, proof and rejection evidence.

Why this lesson matters

Operational Excellence is the ability to support development and run workloads effectively, gain insight into operations, and continuously improve supporting processes and procedures to deliver business value. AWS groups the pillar into Organization, Prepare, Operate, and Evolve.

Operations is not a dashboard or an operations team's problem. Product priorities, ownership, deployment safety, observability, runbooks, event response, learning, and improvement must be designed with the workload. A highly available service that nobody can diagnose or safely change is not operationally excellent.

This lesson turns the pillar into evidence. You will not mark a best practice complete because a tool exists. You will show who owns the outcome, what signal drives action, how a procedure was tested, what happened during a failure, and how learning changed the system.

What you will be able to do

By the end, you can:

  • explain the four Operational Excellence best-practice areas;
  • align team ownership and operational metrics to business outcomes;
  • define workload, dependency, stakeholder, compliance, and support knowledge;
  • use operations as code for infrastructure, configuration, policy, observability, and procedures;
  • design small reversible changes with deployment and rollback evidence;
  • create dashboards, alarms, runbooks, playbooks, and escalation paths;
  • distinguish service, workload, customer, and business health;
  • manage events from detection through recovery and communication;
  • run game days and blameless learning reviews;
  • maintain a risk-ranked improvement backlog; and
  • conduct a defensible Well-Architected review.

Before you start

  • Select one fictional or explicitly approved workload. Do not modify production in this T0 review.
  • Evidence must be timestamped and scoped by account, Region, environment, version, and owner.
  • Redact account IDs, customer data, incident details, internal endpoints, pager contacts, credentials, and commercially sensitive metrics.
  • A Well-Architected review finds and prioritizes risk; it is not an audit certificate or guarantee.

1. Understand the four areas

AreaCore questionEvidence
OrganizationDo people understand outcomes, priorities, roles, obligations, and support?RACI, service catalog, policies, objectives, escalation
PrepareCan the workload and procedures change safely and be operated as designed?IaC, CI/CD, observability, runbooks, game days, readiness
OperateIs workload/business health observed and are events handled effectively?SLOs, dashboards, alarms, tickets, incident timeline
EvolveDoes learning create verified improvement?retrospectives, backlog, experiments, completed actions

The areas form a loop, not project phases. An incident can expose an ownership gap, trigger preparation changes, and alter business objectives.

2. Organize around outcomes

Write a service ownership record:

  • customer and business outcome;
  • product/service owner;
  • development and operations ownership;
  • security, data, finance, compliance, and vendor dependencies;
  • operating hours and support severity;
  • SLOs and critical user journeys;
  • upstream/downstream dependencies;
  • RTO/RPO and degraded modes;
  • escalation and decision authority; and
  • lifecycle and decommission owner.

Use a RACI only where it clarifies decisions. A table with many names and no accepted duty is not ownership. Every alarm, runbook, deployment, exception, vendor escalation, and improvement item needs one accountable owner.

Leadership priorities must resolve tradeoffs. Teams cannot simultaneously maximize feature speed, availability, security, cost reduction, and migration pace without explicit priority and constraints.

3. Prepare with operations as code

Represent infrastructure, configuration, policies, alarms, dashboards, deployment, and repeatable procedures in version control where practical. Changes use peer review, automated tests, approval suited to risk, immutable artifacts, and deployment evidence.

change request -> code/review/tests -> staged deployment
 -> health gates -> progressive exposure -> complete
                        |
                        +-> stop/rollback/fix-forward

Small reversible changes reduce blast radius and make causality clearer. Define rollback before deploy, including data/schema compatibility. A code rollback cannot undo an irreversible data migration or external notification.

Operational readiness checklist:

  1. ownership and support accepted;
  2. architecture/dependencies/quotas documented;
  3. security and data controls tested;
  4. SLOs and user journeys instrumented;
  5. capacity/load and failure tests passed;
  6. dashboards/alarms/tickets connected;
  7. runbooks/playbooks exercised;
  8. deployment/rollback proven;
  9. backup/restore and DR evidence current;
  10. cost/budget/licensing owner accepted.

4. Design observability for decisions

Metrics answer quantitative questions, logs provide event detail, and traces connect distributed paths. Events and configuration history explain change. None alone proves user outcome.

Build a hierarchy:

business outcome
  -> critical user journey
     -> service-level indicator/objective
        -> dependency and resource signals

Examples: checkout success and latency before CPU; accepted telemetry freshness before stream iterator age; completed desktop task before instance count.

An alarm must identify affected outcome, threshold/window, severity, owner, runbook, deduplication, escalation, and expected action. Alarm only on actionable conditions. Use dashboards for exploration and trends, alarms for timely response.

Protect observability: structured schemas, correlation IDs, UTC clocks, retention, encryption, access, redaction, cardinality limits, and cost. Never put secrets or personal content in metric dimensions.

5. Write runbooks and playbooks

A runbook is a repeatable operational procedure, often for known maintenance or recovery. A playbook guides investigation and response where diagnosis branches.

Every procedure needs purpose, scope, prerequisites, permissions, safety checks, exact steps, expected output, stop conditions, escalation, rollback, verification, evidence, owner, and last exercise date.

Automation should be idempotent, bounded, observable, and abortable. Keep human approval for high-impact or uncertain steps. Do not automate an undocumented unsafe process.

Test procedures with a user unfamiliar with the author's assumptions. A runbook that only its writer can execute is fragile.

6. Operate through measurable health

Define normal, degraded, unavailable, recovering, and unknown states. Include external dependencies and data correctness, not just infrastructure availability.

Track:

  • SLO attainment and error-budget consumption;
  • throughput, latency, errors, saturation;
  • queue/backlog/age and data freshness;
  • deployment version and change failure;
  • customer contacts and business transaction outcome;
  • security and compliance exceptions;
  • dependency/provider health;
  • cost and quota headroom; and
  • operational workload such as alerts, tickets, and toil.

Review trends on a cadence. If an SLO repeatedly fails, change roadmap priorities rather than accepting alarm noise.

7. Respond to events

Incident roles can include incident commander, operations/technical leads, communications, scribe, business/security/vendor contacts. One person may fill several roles in a small event, but decision rights must be clear.

Timeline:

detect -> acknowledge -> classify -> contain -> diagnose
 -> mitigate/recover -> validate user outcome -> communicate
 -> close -> learn -> remediate

Preserve timestamps, evidence, commands, changes, hypotheses, decisions, and customer impact. Do not make simultaneous undocumented changes. Prefer reversible containment while root cause is uncertain.

Communication states known facts, impact, action, next update, and owner. Avoid speculation. Security incidents follow the approved security response plan and evidence handling.

8. Evolve through learning

A learning review asks:

  • what happened and customer impact;
  • expected versus actual controls;
  • contributing technical/organizational conditions;
  • what helped or delayed recovery;
  • where detection, procedure, ownership, or architecture failed;
  • which actions prevent, detect, contain, or recover; and
  • how completion will be verified.

Avoid individual blame. Accountability remains: actions need owner, due date, priority, acceptance evidence, and escalation. Close only after effectiveness is tested.

Run game days before real incidents. State steady state, hypothesis, injected failure, safety limits, stop conditions, observers, evidence, recovery, and learning. Progress from tabletop to isolated technical tests and production-safe experiments.

9. Improvement backlog

Score items by customer/business impact, likelihood, detectability, blast radius, compliance, effort, and dependency. Distinguish accepted risk from forgotten work. Risk acceptance needs authority, reason, expiry, compensating control, and review.

Track leading signals such as runbook exercise age, unowned alarms, rollback success, observability coverage, toil, and overdue actions, plus outcomes such as change failure rate, recovery time, SLO, and repeat incidents. Metrics must not incentivize hiding incidents or avoiding valuable change.

10. Read-only evidence

aws cloudwatch describe-alarms --max-records 20 --output table
aws cloudwatch list-dashboards --output table
aws cloudtrail describe-trails --output table
aws ssm list-documents --filters Key=Owner,Values=Self --output table
aws synthetics describe-canaries --output table
aws service-quotas list-requested-service-quota-change-history --output table

Inventory proves existence, not quality, ownership, actionability, exercise, or outcome.

11. Diagnose from evidence

SymptomEvidenceResponse
Alarms fire but nobody actsowner, paging route, runbook, action historyRemove non-actionable alarm or establish accepted response.
Dashboard green, users failuser-journey/business SLI, dependency and data healthAdd outcome-first observability.
Deploy rollback worsens incidentschema/data/external effects, compatibility testsDesign backward compatibility and compensation.
Same incident repeatsreviews, action backlog, completion testsEscalate systemic action and verify effectiveness.
Runbook step failsversion, prerequisites, permissions, expected outputStop safely, correct/exercise procedure, avoid improvising broadly.
Team spends all time on ticketstoil categories, automation value, alert volume, ownershipRemove demand/noise, automate bounded work, fund reliability.
Review has all answers “yes” but no proofevidence dates/scope/testsReclassify as unknown/risk and gather direct evidence.

12. Workshop

Review a fictional order platform. Produce:

  1. outcome/ownership/service catalog;
  2. priorities and obligations;
  3. architecture/dependency map;
  4. operations-as-code inventory;
  5. readiness checklist;
  6. three user-journey SLOs;
  7. telemetry and alarm matrix;
  8. deployment/rollback plan;
  9. two runbooks and two playbooks;
  10. incident role/communication plan;
  11. quota/capacity/cost controls;
  12. backup/restore/DR evidence map;
  13. game-day design;
  14. learning-review template;
  15. risk-ranked improvement backlog; and
  16. Well-Architected answers with evidence/confidence.

Inject stale DNS, exhausted database connections, malformed queue messages, expired certificate, bad schema deploy, and missing on-call response.

Cost and cleanup

Operational capability costs include telemetry ingestion/retention/query, synthetic tests, on-call, support, tooling, duplicate environments, game days, automation, and improvement work. Optimize noisy logs and alarms without removing required evidence.

T0 creates nothing. Preserve redacted findings and references. Any later test must have safety boundaries, owner, stop time, rollback, cleanup, and readback.

Knowledge check

  1. Four OE areas? Organization, Prepare, Operate, Evolve.
  2. Why outcome-first metrics? Resource health can be green while users fail.
  3. Runbook versus playbook? Repeatable procedure versus branching investigation/response.
  4. Why small reversible change? Lower blast radius, clearer causality, faster recovery.
  5. What makes an alarm actionable? Owner, meaningful threshold, runbook, routing, and expected response.
  6. Why game days? Test hypotheses, systems, people, and procedures before crisis.
  7. Blameless versus no accountability? Study systems without blame while assigning verified actions.
  8. What closes an improvement? Tested evidence that the intended risk/outcome improved.

Lesson acceptance

Pass when each Well-Architected answer links to current scoped evidence and a named owner, and high risks enter a funded/testable backlog. Fail if tool existence substitutes for operation, dashboards omit user outcomes, alarms lack response, runbooks are untested, rollback ignores data, incidents produce no verified improvement, or all unknowns are marked complete.

Official sources

Advertisement