AWS 340: Architecture review findings, priorities, owners, and measurable acceptance criteria
Why this lesson matters
An architecture review creates value only when observations become owned, prioritized, measurable improvements or explicitly accepted risks. “Enable monitoring,” “improve security,” and “make it highly available” are not actionable. A ticket marked done does not prove the customer risk was reduced.
AWS Well-Architected guidance recommends prioritizing by business value and effort, assigning owners, using measurable goals, implementing iteratively, and saving milestones as architecture improves. This lesson turns 25 noisy observations into a controlled improvement system.
Learning outcomes
By the end, you can:
- distinguish observations, findings, risks, causes, actions, and evidence;
- normalize and deduplicate review input without losing traceability;
- prioritize by business impact, urgency, dependency, confidence, and effort;
- assign accountable owners and decision authorities;
- write measurable acceptance and negative/failure retests;
- govern exceptions, residual risk, expiry, and escalation;
- track implementation through Well-Architected milestones;
- close findings only when outcome evidence passes.
1. Finding anatomy
| Field | Required content |
|---|---|
| ID and title | Stable, concise, outcome-oriented identifier |
| Scope | Workload, journey, environment, accounts, Regions, components |
| Condition | Factual current state without blame |
| Evidence | Source, timestamp, query/test, confidence, access boundary |
| Requirement | Affected business/NFR/control and owner |
| Scenario | How condition becomes failure or loss |
| Impact | User, data, security, recovery, cost, compliance, operations |
| Cause | Verified or explicitly hypothesized mechanism |
| Risk | Likelihood, impact, exposure window, rationale |
| Treatment | Avoid, reduce, transfer, accept, or research |
| Action | Smallest effective change and dependencies |
| Owner/date | One accountable owner and due/review date |
| Acceptance | Measurable test, threshold, environment, duration, approver |
| Rollback | Trigger and safe reversal/forward repair |
| Residual risk | What remains after action |
| Status | New, triaged, planned, active, validating, accepted, closed |
Example condition: “Restore test has not been executed for the production database backup created by policy X.” Avoid “the database team is careless.”
2. Observation versus risk
An observation is evidence: one subnet contains all application tasks. A risk explains scenario and impact: loss of that Availability Zone stops checkout, violating monthly availability and revenue requirements. An action changes the condition: distribute capacity and dependencies across independent AZs. Acceptance proves outcome during an impairment test.
Do not jump from a scanner finding directly to a favored tool. Verify scope, exploit/failure path, compensating controls, false-positive possibility, and business requirement.
3. Normalize noisy input
Review evidence may come from WAFR, Security Hub, Config, incidents, audits, cost analysis, support, penetration tests, operators, and customers. Preserve original source IDs, then normalize into one register.
Deduplicate only when condition, scope, scenario, owner, and remediation are materially the same. A public bucket finding and broad KMS access may affect one data journey but require different controls. Link parent/child findings when one root cause creates several symptoms.
Separate:
- verified finding;
- suspected condition requiring research;
- improvement opportunity without current violation;
- accepted risk;
- false positive with evidence;
- duplicate linked to canonical item;
- out-of-scope item transferred to a named owner.
4. Risk scoring without false precision
Use organization policy. A simple model can rate likelihood and impact from 1–5, but the number requires narrative evidence. Include exploitability/failure frequency, exposure, detection, recovery, data class, affected users, financial/regulatory impact, and confidence.
Mandatory legal/safety constraints and launch blockers cannot be averaged below cosmetic work. Low-confidence high-impact conditions may deserve urgent investigation rather than low priority.
Distinguish inherent risk before controls from residual risk after current/proposed controls. Record who has authority to accept residual risk.
5. Prioritization
Prioritize:
- immediate customer/security/safety containment;
- mandatory constraints and launch blockers;
- high impact/high likelihood or weak recovery;
- prerequisites that unlock several improvements;
- high-value, low-complexity improvements;
- strategic debt with funded roadmap;
- monitoring/research that reduces uncertainty.
Also consider implementation risk, effort, dependencies, delivery windows, reversibility, and available team capacity. Avoid “priority by finding count” or by loudest stakeholder.
Create a dependency graph. Logging/observability may precede automated remediation; tested backup may precede database migration; account/identity foundation may precede workload rollout.
6. Select treatment and solution
For each risk compare options and select the simplest solution that meets requirements. Prefer reusable pattern-based and reversible approaches where possible. Document consequences, operational ownership, cost, security, migration, and failure.
Containment is not closure. Disabling a feature may reduce immediate exposure while a durable fix is designed. Track both with separate acceptance.
Do not combine unrelated fixes into one change if it destroys diagnosability. Conversely, use a program-level parent where a foundational solution safely addresses many linked findings.
7. Owners and governance
Assign one accountable owner, even when many teams implement. Record business risk owner, technical implementer, control owner, validator, and approver. “Platform and app teams” is not accountability.
Owner must have authority/capacity or escalate. Due dates need dependency and business context. A risk committee does not become the implementation owner merely because it approves exceptions.
Use cadence:
- daily/incident for active critical exposure;
- weekly for blockers/high risks;
- sprint/fortnightly for active improvement;
- monthly/quarterly for accepted risk and strategic debt.
8. Measurable acceptance
Acceptance criteria state environment, starting condition, action/test, workload/failure, threshold, duration, evidence, and approver.
Weak: “Enable backups.”
Strong: “In isolated recovery account, restore the latest protected production-like database copy, run schema/checksum and five business queries, and make the application read-only journey available within 55 minutes; measured data loss is under five minutes; data owner and service owner approve the report.”
Include positive, negative, and failure tests:
- intended access succeeds;
- unauthorized access is denied and logged;
- dependency/AZ/timeout failure produces expected degradation/recovery;
- alarm reaches owner and runbook action works;
- cost and performance guardrails remain within bounds.
“Merged,” “deployed,” or “ticket closed” is implementation state, not outcome acceptance.
9. Exceptions and risk acceptance
An exception records control/requirement, scope, business reason, risk, compensating controls, evidence, accountable risk authority, start, expiry, review cadence, monitoring, and exit plan. It must not silently become permanent.
Reject acceptance by an unauthorized technical team. Escalate when residual risk exceeds tolerance or mandatory regulation. At expiry, close, renew with fresh evidence, or remediate; never auto-renew.
10. Implementation and validation
Move statuses with evidence:
New -> Triaged -> Planned -> Active -> Validating
-> Closed | Accepted risk | Reopened
Before change, preserve baseline and rollback. During change, capture version and unexpected effects. After change, run the agreed acceptance in the right environment and representative period. Independent validation is useful for high-consequence controls.
Close only when evidence proves acceptance and residual risk is recorded. Reopen on failed retest, regression, expired evidence, changed scope, or incident.
11. Well-Architected improvement loop
Map findings to the six pillars and workload context. Prioritize a feasible set, add to delivery backlog, implement, validate, and save a Well-Architected milestone. Compare risk state and customer outcomes, not simply HRI/MRI counts. Repeat as the workload changes.
Read-only discovery:
aws wellarchitected list-workloads --output table
aws securityhub get-findings --max-results 20 --output json
aws configservice get-compliance-summary-by-resource-type --output table
aws cloudwatch describe-alarms --state-value ALARM --output table
Results can be incomplete due to service enrollment, Region, aggregation, or permission. Record coverage before conclusions.
12. Metrics
Track aged critical/high risk, time to triage, time to containment, time to accepted closure, overdue items, exception age/expiry, reopened rate, acceptance failure rate, recurrence after closure, unowned findings, and customer/SLO outcome change.
Avoid incentivizing closure volume. It can encourage duplicate merging, false-positive labeling, or weak acceptance.
13. Guided workshop
Normalize 25 supplied observations from AWS335–AWS339. Produce:
- source/evidence register;
- canonical findings with source links;
- duplicate/parent-child map;
- fact versus hypothesis classification;
- affected requirements/journeys;
- scenario and impact statements;
- inherent/current/residual risk;
- confidence and research items;
- dependency graph;
- prioritized first iteration;
- treatment/options decisions;
- accountable owner/RACI;
- positive/negative/failure acceptance criteria;
- rollback and guardrail criteria;
- exceptions with expiry/authority;
- backlog and review cadence;
- milestone/retest evidence plan;
- executive risk summary and stakeholder approval.
14. Diagnostic examples
| Bad pattern | Correction |
|---|---|
| “Encrypt database” | Identify unencrypted data/key/identity path and measurable proof |
| 17 tickets for one logging gap | Parent foundational finding plus scoped child validation |
| “Shared owner” | One accountable owner with contributing teams |
| Closed when deployed | Validate required user/control outcome |
| Lowest effort always first | Respect mandatory/high-impact risk and dependencies |
| Exception without expiry | Add authority, monitoring, date, and exit |
Cost and cleanup
This T0 workshop creates no resources. Improvement plans must include engineering, testing, downtime, tools, operations, and residual-risk cost. Sanitize local findings and retain only approved evidence.
Knowledge check
- What turns an observation into a risk? A credible scenario and impact on a requirement.
- Is deployment proof of closure? No; acceptance outcome must pass.
- Who owns a finding? One accountable owner, supported by contributors.
- Should every HRI be fixed simultaneously? No; prioritize a feasible impact-led iteration without ignoring blockers.
- What happens at exception expiry? Close, renew with fresh authority/evidence, or remediate.
Lesson acceptance
Submit all 18 artifacts. Every finding must trace to evidence and requirement, have one accountable owner, risk rationale, dependency-aware priority, measurable positive/negative/failure acceptance, residual risk, and milestone/retest closure evidence.