Lesson 340 · AWS Learning Path

AWS 340: Architecture review findings, priorities, owners, and measurable acceptance criteria

· Published · 7 min read

Labelled process diagram for AWS 340: Review observations and evidence to Normalized risk findings to Owned remediation or exception to Measured acceptance and retest evidence, with decision, proof and rejection...

Why this lesson matters

An architecture review creates value only when observations become owned, prioritized, measurable improvements or explicitly accepted risks. “Enable monitoring,” “improve security,” and “make it highly available” are not actionable. A ticket marked done does not prove the customer risk was reduced.

AWS Well-Architected guidance recommends prioritizing by business value and effort, assigning owners, using measurable goals, implementing iteratively, and saving milestones as architecture improves. This lesson turns 25 noisy observations into a controlled improvement system.

Learning outcomes

By the end, you can:

  • distinguish observations, findings, risks, causes, actions, and evidence;
  • normalize and deduplicate review input without losing traceability;
  • prioritize by business impact, urgency, dependency, confidence, and effort;
  • assign accountable owners and decision authorities;
  • write measurable acceptance and negative/failure retests;
  • govern exceptions, residual risk, expiry, and escalation;
  • track implementation through Well-Architected milestones;
  • close findings only when outcome evidence passes.

1. Finding anatomy

FieldRequired content
ID and titleStable, concise, outcome-oriented identifier
ScopeWorkload, journey, environment, accounts, Regions, components
ConditionFactual current state without blame
EvidenceSource, timestamp, query/test, confidence, access boundary
RequirementAffected business/NFR/control and owner
ScenarioHow condition becomes failure or loss
ImpactUser, data, security, recovery, cost, compliance, operations
CauseVerified or explicitly hypothesized mechanism
RiskLikelihood, impact, exposure window, rationale
TreatmentAvoid, reduce, transfer, accept, or research
ActionSmallest effective change and dependencies
Owner/dateOne accountable owner and due/review date
AcceptanceMeasurable test, threshold, environment, duration, approver
RollbackTrigger and safe reversal/forward repair
Residual riskWhat remains after action
StatusNew, triaged, planned, active, validating, accepted, closed

Example condition: “Restore test has not been executed for the production database backup created by policy X.” Avoid “the database team is careless.”

2. Observation versus risk

An observation is evidence: one subnet contains all application tasks. A risk explains scenario and impact: loss of that Availability Zone stops checkout, violating monthly availability and revenue requirements. An action changes the condition: distribute capacity and dependencies across independent AZs. Acceptance proves outcome during an impairment test.

Do not jump from a scanner finding directly to a favored tool. Verify scope, exploit/failure path, compensating controls, false-positive possibility, and business requirement.

3. Normalize noisy input

Review evidence may come from WAFR, Security Hub, Config, incidents, audits, cost analysis, support, penetration tests, operators, and customers. Preserve original source IDs, then normalize into one register.

Deduplicate only when condition, scope, scenario, owner, and remediation are materially the same. A public bucket finding and broad KMS access may affect one data journey but require different controls. Link parent/child findings when one root cause creates several symptoms.

Separate:

  • verified finding;
  • suspected condition requiring research;
  • improvement opportunity without current violation;
  • accepted risk;
  • false positive with evidence;
  • duplicate linked to canonical item;
  • out-of-scope item transferred to a named owner.

4. Risk scoring without false precision

Use organization policy. A simple model can rate likelihood and impact from 1–5, but the number requires narrative evidence. Include exploitability/failure frequency, exposure, detection, recovery, data class, affected users, financial/regulatory impact, and confidence.

Mandatory legal/safety constraints and launch blockers cannot be averaged below cosmetic work. Low-confidence high-impact conditions may deserve urgent investigation rather than low priority.

Distinguish inherent risk before controls from residual risk after current/proposed controls. Record who has authority to accept residual risk.

5. Prioritization

Prioritize:

  1. immediate customer/security/safety containment;
  2. mandatory constraints and launch blockers;
  3. high impact/high likelihood or weak recovery;
  4. prerequisites that unlock several improvements;
  5. high-value, low-complexity improvements;
  6. strategic debt with funded roadmap;
  7. monitoring/research that reduces uncertainty.

Also consider implementation risk, effort, dependencies, delivery windows, reversibility, and available team capacity. Avoid “priority by finding count” or by loudest stakeholder.

Create a dependency graph. Logging/observability may precede automated remediation; tested backup may precede database migration; account/identity foundation may precede workload rollout.

6. Select treatment and solution

For each risk compare options and select the simplest solution that meets requirements. Prefer reusable pattern-based and reversible approaches where possible. Document consequences, operational ownership, cost, security, migration, and failure.

Containment is not closure. Disabling a feature may reduce immediate exposure while a durable fix is designed. Track both with separate acceptance.

Do not combine unrelated fixes into one change if it destroys diagnosability. Conversely, use a program-level parent where a foundational solution safely addresses many linked findings.

7. Owners and governance

Assign one accountable owner, even when many teams implement. Record business risk owner, technical implementer, control owner, validator, and approver. “Platform and app teams” is not accountability.

Owner must have authority/capacity or escalate. Due dates need dependency and business context. A risk committee does not become the implementation owner merely because it approves exceptions.

Use cadence:

  • daily/incident for active critical exposure;
  • weekly for blockers/high risks;
  • sprint/fortnightly for active improvement;
  • monthly/quarterly for accepted risk and strategic debt.

8. Measurable acceptance

Acceptance criteria state environment, starting condition, action/test, workload/failure, threshold, duration, evidence, and approver.

Weak: “Enable backups.”

Strong: “In isolated recovery account, restore the latest protected production-like database copy, run schema/checksum and five business queries, and make the application read-only journey available within 55 minutes; measured data loss is under five minutes; data owner and service owner approve the report.”

Include positive, negative, and failure tests:

  • intended access succeeds;
  • unauthorized access is denied and logged;
  • dependency/AZ/timeout failure produces expected degradation/recovery;
  • alarm reaches owner and runbook action works;
  • cost and performance guardrails remain within bounds.

“Merged,” “deployed,” or “ticket closed” is implementation state, not outcome acceptance.

9. Exceptions and risk acceptance

An exception records control/requirement, scope, business reason, risk, compensating controls, evidence, accountable risk authority, start, expiry, review cadence, monitoring, and exit plan. It must not silently become permanent.

Reject acceptance by an unauthorized technical team. Escalate when residual risk exceeds tolerance or mandatory regulation. At expiry, close, renew with fresh evidence, or remediate; never auto-renew.

10. Implementation and validation

Move statuses with evidence:

New -> Triaged -> Planned -> Active -> Validating
-> Closed | Accepted risk | Reopened

Before change, preserve baseline and rollback. During change, capture version and unexpected effects. After change, run the agreed acceptance in the right environment and representative period. Independent validation is useful for high-consequence controls.

Close only when evidence proves acceptance and residual risk is recorded. Reopen on failed retest, regression, expired evidence, changed scope, or incident.

11. Well-Architected improvement loop

Map findings to the six pillars and workload context. Prioritize a feasible set, add to delivery backlog, implement, validate, and save a Well-Architected milestone. Compare risk state and customer outcomes, not simply HRI/MRI counts. Repeat as the workload changes.

Read-only discovery:

aws wellarchitected list-workloads --output table
aws securityhub get-findings --max-results 20 --output json
aws configservice get-compliance-summary-by-resource-type --output table
aws cloudwatch describe-alarms --state-value ALARM --output table

Results can be incomplete due to service enrollment, Region, aggregation, or permission. Record coverage before conclusions.

12. Metrics

Track aged critical/high risk, time to triage, time to containment, time to accepted closure, overdue items, exception age/expiry, reopened rate, acceptance failure rate, recurrence after closure, unowned findings, and customer/SLO outcome change.

Avoid incentivizing closure volume. It can encourage duplicate merging, false-positive labeling, or weak acceptance.

13. Guided workshop

Normalize 25 supplied observations from AWS335–AWS339. Produce:

  1. source/evidence register;
  2. canonical findings with source links;
  3. duplicate/parent-child map;
  4. fact versus hypothesis classification;
  5. affected requirements/journeys;
  6. scenario and impact statements;
  7. inherent/current/residual risk;
  8. confidence and research items;
  9. dependency graph;
  10. prioritized first iteration;
  11. treatment/options decisions;
  12. accountable owner/RACI;
  13. positive/negative/failure acceptance criteria;
  14. rollback and guardrail criteria;
  15. exceptions with expiry/authority;
  16. backlog and review cadence;
  17. milestone/retest evidence plan;
  18. executive risk summary and stakeholder approval.

14. Diagnostic examples

Bad patternCorrection
“Encrypt database”Identify unencrypted data/key/identity path and measurable proof
17 tickets for one logging gapParent foundational finding plus scoped child validation
“Shared owner”One accountable owner with contributing teams
Closed when deployedValidate required user/control outcome
Lowest effort always firstRespect mandatory/high-impact risk and dependencies
Exception without expiryAdd authority, monitoring, date, and exit

Cost and cleanup

This T0 workshop creates no resources. Improvement plans must include engineering, testing, downtime, tools, operations, and residual-risk cost. Sanitize local findings and retain only approved evidence.

Knowledge check

  1. What turns an observation into a risk? A credible scenario and impact on a requirement.
  2. Is deployment proof of closure? No; acceptance outcome must pass.
  3. Who owns a finding? One accountable owner, supported by contributors.
  4. Should every HRI be fixed simultaneously? No; prioritize a feasible impact-led iteration without ignoring blockers.
  5. What happens at exception expiry? Close, renew with fresh authority/evidence, or remediate.

Lesson acceptance

Submit all 18 artifacts. Every finding must trace to evidence and requirement, have one accountable owner, risk rationale, dependency-aware priority, measurable positive/negative/failure acceptance, residual risk, and milestone/retest closure evidence.

Official sources

Advertisement