Lesson 341 · AWS Learning Path

AWS 341: Architecture: complete the SAP enterprise multi-account capstone

· Published · 7 min read

Labelled process diagram for AWS 341: Capstone requirements and evidence to Reviewed multi-account architecture to Defended corrections and trade-offs to Accepted decision, residual risk, and operating handoff, with...

Why this lesson matters

An enterprise architecture is not complete because its diagram looks convincing. It is complete only when requirements trace to controls, controls have owners, failure behavior is understood, costs are explainable, and reviewers can reproduce the evidence. This lesson is the final examination of the enterprise multi-account design built in AWS335.

You will act as both architect and review board. First prove the design is complete. Then challenge it with policy, identity, network, security, acquisition, and operating failures. Every discovered weakness becomes a governed finding through the workflow in AWS340. The result is a defensible decision package, not another service summary.

What you will be able to do

By the end, you can:

  • defend an AWS multi-account operating model from business requirement to technical evidence;
  • explain Organizations, organizational units, accounts, policies, delegated administration, and Control Tower without confusing their boundaries;
  • test whether preventive, detective, and responsive controls work together;
  • expose hidden dependencies in identity, centralized networking, logging, backup, and account vending;
  • distinguish a design assumption from measured or documented evidence;
  • manage exceptions and residual risk with named owners and expiry dates;
  • correct critical findings and obtain conditional or final approval.

Safety, scope, and required inputs

This is a T0 review. Use documents, sanitized read-only evidence, and local diagrams. Do not create or change AWS resources. Never submit account IDs, credentials, sensitive findings, or complete ARNs.

Bring the 24 artifacts from AWS335 and the findings register template from AWS340. At minimum you need requirements, stakeholder map, account and OU model, policy hierarchy, identity paths, network and DNS design, logging and security delegation, backup and recovery model, account-vending workflow, cost model, quota register, exception process, failure tests, and approval record. Missing input is itself a finding, not permission to invent evidence.

The defense model

Business outcomes and constraints
              |
              v
Requirements -> architecture decisions -> AWS controls
              |                         |
              v                         v
        accountable owners       reproducible evidence
              \                         /
               +-> failure tests <-+
                         |
                         v
              findings and corrections
                         |
                         v
             accepted residual risk

The chain must work in both directions. A requirement must reach an implemented or planned control. Every control must return to a requirement, owner, evidence source, cost, and lifecycle decision. An orphan control adds complexity; an orphan requirement creates unmanaged risk.

Stage 1: completeness preflight

Build a submission index with artifact name, version, owner, reviewer, evidence date, sensitivity, and related requirement IDs. Reject stale screenshots without context. A screenshot shows a moment, not intent or continuing compliance.

Select at least 20 requirements across security, reliability, operations, performance, cost, sustainability, compliance, and organizational constraints. For each record:

FieldRequired proof
RequirementClear, testable statement and business source
DecisionChosen pattern and rejected alternative
ScopeOrganization, OU, account, Region, VPC, or resource
ControlPreventive, detective, responsive, or recovery control
OwnerOne accountable role plus operating team
EvidenceDocument, query, log, test, or approved plan
FailureExpected behavior when the control or dependency fails
CostDriver, payer, allocation key, and forecast assumption
LifecycleProvision, change, exception, review, and retirement

Mark each row proven, partially proven, contradicted, or absent. “Designed” and “implemented” are different states.

Stage 2: four review panels

Panel A: governance and organization

Defend why each OU exists and what changes when an account moves between OUs. Explain policy inheritance and remember that service control policies set permission guardrails; they do not grant permissions. Show where tag policies, backup policies, declarative policies, and resource control policies are applicable and where another control is required.

Demonstrate management-account restrictions, root-user protection, break-glass ownership, delegated administrator choices, account closure, suspended accounts, and acquisition onboarding. Test a proposed SCP against normal deployment, incident response, logging, backup, and emergency access. A guardrail that prevents recovery is not successful governance.

Panel B: identity, security, and evidence

Trace workforce access from the identity provider through IAM Identity Center, permission sets, roles, sessions, and target resources. Trace workload identity separately. State session duration, privileged elevation, separation of duties, emergency access, and what happens when the identity provider is unavailable.

Then trace CloudTrail, AWS Config, GuardDuty, Security Hub, and access findings from member account to delegated security or log-archive accounts. Identify who can alter, suppress, delete, or query evidence. Prove onboarding checks detect a new account missing required trails, configuration recording, threat detection, contacts, or backup assignment.

Panel C: network, DNS, platform, and recovery

Trace north-south and east-west traffic through DNS, ingress, inspection, Transit Gateway or Cloud WAN, routing, security groups, network ACLs, endpoints, and egress. State ownership and blast radius for each shared component. Validate Availability Zone independence instead of drawing two identical boxes.

Defend IP address management, overlapping-CIDR handling, hybrid routes, DNS forwarding, certificate ownership, private endpoint strategy, quotas, and flow-log retention. Connect backup policies to restore ownership and tested recovery objectives. “Backed up” is not equivalent to “restorable.”

Panel D: operations, finance, and lifecycle

Explain account vending from request and approval through baseline application, validation, handover, drift response, quarantine, and closure. Show service quotas, patching, vulnerability response, observability, escalation, and platform product ownership.

Reconcile shared-service, security, network, support, logging, backup, and data-transfer costs. State allocation rules and unit metrics such as cost per account or workload. Model steady state, growth, incident retention, and acquisition. Identify commitment risk and the owner of forecast variance.

Stage 3: red-team scenarios

Run all scenarios as tabletop tests. Record injected event, detection source, first decision, responsible role, containment, recovery, evidence, elapsed time, and design change.

  1. An SCP update blocks the logging service and the emergency remediation role.
  2. The identity provider and IAM Identity Center path are unavailable during a security incident.
  3. One Availability Zone supporting centralized inspection or Transit Gateway attachments fails.
  4. A newly vended account has no effective Config recorder, GuardDuty enrollment, or central log delivery.
  5. The delegated security administration account becomes inaccessible.
  6. An acquired company arrives with overlapping CIDRs, unknown administrators, and incomplete inventory.
  7. A management-account credential or root recovery channel is suspected compromised.

For each scenario, ask what still works, what fails closed, what fails open, which evidence remains trustworthy, and whether a central dependency expands blast radius. Add every unsupported answer to the findings register.

Stage 4: scoring and critical fail conditions

Score each dimension from 0 to 4: requirements traceability, governance, identity, security evidence, networking, resilience, operations, cost, lifecycle, and communication. A score of 0 means absent; 1 means asserted; 2 means designed; 3 means evidenced; 4 means failure-tested and owner-approved.

The numerical score cannot override a critical failure. The capstone fails until corrected if it contains any of these:

  • no accountable owner for a critical control or residual risk;
  • daily administration through root or unrestricted management-account access;
  • no durable organization-wide audit path;
  • an untested emergency-access path;
  • centralized networking with no stated blast radius or recovery path;
  • no account onboarding validation or drift response;
  • recovery objectives claimed without restore evidence;
  • unresolved critical finding, fabricated evidence, or approval with hidden exclusions.

Stage 5: correction and approval

Process each gap through AWS340: observation, evidence, affected requirement, consequence, likelihood, priority, owner, treatment, due date, acceptance criteria, retest, and residual risk. Correct the design and all affected artifacts, not only the slide where the issue appeared.

Approval states are rejected, conditionally approved, or approved. Conditional approval requires explicit conditions, owners, deadlines, and expiry. Security, operations, finance, and business stakeholders approve only their risk domain; one signature does not silently accept another team’s risk.

Required submission artifacts

Submit these 18 items:

  1. final submission index;
  2. executive outcome and constraint summary;
  3. traceability matrix for at least 20 requirements;
  4. organization, OU, and account model;
  5. policy inheritance and policy-test evidence;
  6. workforce, workload, and emergency identity paths;
  7. delegated security and evidence flow;
  8. account-vending and baseline-validation proof;
  9. network, DNS, inspection, and hybrid flow;
  10. backup, restore, and regional recovery proof;
  11. service quota and capacity register;
  12. ownership and RACI matrix;
  13. cost model, allocation method, and forecast;
  14. seven completed red-team records;
  15. scored review rubric;
  16. findings register with correction and retest evidence;
  17. residual-risk and exception register;
  18. stakeholder decision record and oral-defense notes.

Lesson acceptance

All 18 artifacts must agree. Twenty requirements must be traceable in both directions, all seven scenarios must be completed, every critical finding must be corrected and retested, and every accepted residual risk must have an authorized owner and review date. The learner must explain the design without relying on unexplained product names.

Official sources

Advertisement