Lesson 335 · AWS Learning Path

AWS 335: Enterprise multi-account capstone

· Published · 8 min read

Labelled process diagram for AWS 335: Enterprise constraints to Landing-zone and operating-model design to Guarded workload-account consumption to Governance, operations, recovery, and lifecycle evidence, with...

Why this capstone matters

Enterprise AWS architecture is an operating model, not an OU diagram. Accounts, identities, policies, networks, DNS, logs, keys, security tools, backups, deployment, costs, quotas, incidents, exceptions, and lifecycle must work together across teams. Centralizing a capability without staffing, service levels, failure isolation, or an exit path creates a hidden critical dependency.

This capstone integrates earlier organization, security, network, operations, reliability, cost, and architecture-decision lessons into a reviewable 60-account platform. It is design-only and uses fictional identifiers so learners can work safely without creating an AWS Organization.

Scenario

Northstar Group has three business units: Retail, Logistics, and Analytics. It operates in India and the United Kingdom, recently acquired a company with 12 AWS accounts, and expects 60 accounts within 18 months. Payment and employee data have jurisdictional controls. Teams need autonomy, but current accounts use inconsistent identity, logging, networking, backup, tags, and root contacts. The central platform team has eight engineers; security operates 24x7; product teams own applications and on-call response.

Required outcomes:

  • separate production, nonproduction, sandbox, security, and infrastructure blast radius;
  • federated workforce access with controlled privilege and emergency recovery;
  • organization-wide audit/security coverage with protected evidence;
  • governed account vending in two business days;
  • shared hybrid/network/DNS capability without one uncontrolled path;
  • cost ownership and quota readiness for every workload;
  • acquired-account integration without implicit trust;
  • exceptions with authority, expiry, and compensating controls;
  • tested incident and account closure procedures.

Learning outcomes

By the end, you can produce and defend a complete enterprise platform, trace every control to a requirement, diagnose inherited-policy and shared-service failure, and define evidence that proves platform consumption works.

1. Requirements and decision rights

Create a register that distinguishes facts, constraints, assumptions, preferences, risks, decisions, and unknowns. Identify management-account owner, cloud-platform owner, identity owner, network/DNS owner, security/tooling owner, log/backup data owner, FinOps owner, product-account owner, risk authority, incident commander, and closure authority.

DecisionAccountable ownerRequired evidence
OU/control policyCloud governance and securityPolicy evaluation and staged tests
Workforce permissionsIdentity and workload ownerGroup lifecycle, assignment, access review
Network route/inspectionNetwork ownerBidirectional path, failure, change, rollback
Security exceptionRisk authorityImpact, expiry, compensating control
Account closureBusiness/data/finance ownersDependency, retention, cost, recovery sign-off

No central team owns application correctness merely because it owns the landing zone.

2. Account and OU architecture

Use accounts as resource, policy, quota, billing, and blast-radius boundaries. Keep normal workloads out of the management account. Define dedicated log archive and audit/security accounts plus network and shared-service accounts where justified.

Design OUs around stable governance differences, not the changing organization chart. A viable example starts with Security, Infrastructure, Workloads/Production, Workloads/NonProduction, Sandbox, PolicyStaging, AcquisitionQuarantine, and Suspended. Add jurisdiction-specific children only when controls differ materially.

For each account category state purpose, owner, permitted Regions, data class, network model, identity assignments, security/backup integration, budget, support, lifecycle, and exit. Avoid a deep hierarchy that makes inheritance opaque.

3. Policy architecture

SCPs and RCPs define maximum permissions for supported principals/resources; they do not grant access. IAM and resource policies still authorize. Build controls in layers:

  • organization policies for hard boundaries;
  • Control Tower preventive/proactive/detective controls where applicable;
  • IAM Identity Center permission sets and workload roles;
  • resource policies and KMS key policies;
  • Config/security findings and responsive automation;
  • application authorization and data controls.

Test restrictive policies in PolicyStaging with positive and negative cases. Preserve logging, backup, security response, service-linked roles, and account recovery. Record policy size/attachment limits and exception mechanism. Never edit the production OU first to “see what breaks.”

4. Workforce and workload identity

Federate workforce access from the approved identity provider through IAM Identity Center. Map groups to permission sets, account sets, session duration, approval, and review. Define joiner/mover/leaver behavior, privileged elevation, vendor sponsorship/expiry, and quarterly access review.

Design separate protected break-glass access and test it without sharing credentials. Root credentials, contacts, and recovery are centrally governed but not used for normal work.

Workloads use short-lived roles with explicit trust conditions. Map cross-account caller, role, external/source conditions, session policy, resource policy, KMS policy, and CloudTrail evidence. Avoid organization-wide wildcard trust merely because accounts are internal.

5. Delegated administration

Delegate supported organization-integrated services to specialist accounts so routine operations avoid the management account. Create a register for service, trusted access, delegated account, operators, visible data, Regions, failure impact, revocation, and incident owner.

Candidate capabilities include security aggregation, Config, GuardDuty, Security Hub, IAM Access Analyzer, Firewall Manager, Backup, IPAM, and account governance. Verify current service behavior; “delegated administrator” is not identical across services.

6. Network and DNS platform

Define global IPAM pools and nonoverlapping account/VPC allocation. Compare distributed, centralized, and hybrid ingress/egress/inspection. For every traffic class prove DNS, forward route, security controls, endpoint/listener, application authorization, return route, logging, and failure behavior.

Use Transit Gateway, Cloud WAN, VPC sharing, peering, PrivateLink, or service connectivity only where scale and ownership fit. Central network accounts can increase blast radius. Use multiple AZs, controlled changes, route/attachment ownership, quotas, flow logs, and rollback.

Hybrid DNS needs inbound/outbound Resolver endpoints/rules, namespace authority, sharing, query logging, cache/negative behavior, and on-premises forwarding. Split-horizon names and acquired-company overlap require explicit transition.

7. Logging, security, and evidence

Design organization trails and service logs to protected destinations with encryption, retention, immutability where required, separation of write/read/delete, and searchable incident access. Include all enabled Regions and verify new-account enrollment.

Define coverage and owner for Config, Security Hub, GuardDuty, Access Analyzer, Inspector, Macie where justified, vulnerability/patch evidence, and incident automation. A green aggregate can hide disabled Regions or failed forwarding. Maintain account-by-Region coverage inventory and reconcile it.

Protect sensitive logs from excessive central readers. Security visibility does not authorize unrestricted access to application data.

8. Backup and recovery governance

Apply tag/resource-based backup policies with workload-approved RPO, retention, vault, cross-account/Region copy, key, legal hold, and restore-test schedule. Use logically isolated or protected recovery evidence where required. Backup centralization must not let one compromised operator delete source and recovery copies.

Product teams own recovery acceptance; the platform team owns backup capability. Measure restore through application/data validation, not job completion alone.

9. Account vending and lifecycle

Account request fields include owner, business purpose, environment, data class, jurisdiction, cost center, Regions, network pattern, identity groups, support, expiration, and quota needs. Automated baseline applies contacts, OU, tags, budgets, logs/security/backup, identity, network/DNS, inventory, and deployment role.

Acceptance tests verify identity, allowed/denied policy, logging arrival, security enrollment, backup, DNS/network, budget, deployment, and incident contacts. Account status moves requested, approved, provisioned, validated, active, restricted, suspended, and closed.

Closure requires dependency and DNS/certificate checks, data retention/export/legal hold, backup decision, identities, shared resources, Marketplace/licenses, commitments, support cases, cost, root/contact custody, and the AWS closure waiting period. Never close by tag alone.

10. Platform delivery and guardrails as code

Version landing-zone extensions, StackSets, account baselines, policies, permission sets, network configuration, Config rules, backup policies, and tests. Use peer review, static policy checks, staged OUs/accounts, canary rollout, stop conditions, drift detection, and rollback.

Separate platform release from workload release. Publish supported patterns with version, service level, cost model, quota, onboarding, telemetry, incident route, deprecation, and consumer exit.

11. FinOps, quotas, and capacity

Require account metadata, cost categories/tags, allocation ownership, budgets, anomaly monitoring, forecast, commitment policy, and shared-cost allocation. Keep showback transparent; use chargeback only with finance governance.

Track per-account and regional quotas before launch and failover. Account separation can add quota pools but does not remove architecture limits. Central services need consumer and failure capacity models.

12. Exceptions, incidents, and acquisitions

An exception records requirement/control, reason, scope, risk, owner, approver, compensating control, monitoring, expiry, and closure. Expired exceptions trigger review, not silent permanence.

Define incident command across security, platform, network, identity, product, legal, communications, and vendor teams. Test loss of Identity Center, shared network, DNS, log destination, delegated security account, and management-account access.

Treat acquired accounts as untrusted. Inventory root/contact, payer/organization, identity, public exposure, logs, keys, domains, networks/CIDRs, data, support, contracts, and dependencies. Place them in a restricted integration path before connectivity or policy inheritance; test controls before moving OUs.

13. Read-only evidence track

aws sts get-caller-identity --query Arn --output text
aws organizations describe-organization --output json
aws organizations list-accounts --output table
aws organizations list-policies --filter SERVICE_CONTROL_POLICY --output table
aws sso-admin list-instances --output table
aws cloudtrail describe-trails --include-shadow-trails false --output table
aws backup list-backup-vaults --output table

These require privileged read permissions and may return AccessDenied. Use fictional supplied evidence instead of requesting management-account access. Redact all identifiers.

14. Capstone deliverables

Submit 24 artifacts:

  1. requirements and decision-rights registers;
  2. 60-account catalog and lifecycle states;
  3. OU hierarchy with inheritance rationale;
  4. management/log/audit account protection model;
  5. SCP/RCP/control catalog and evaluation tests;
  6. federation, permission-set, elevation, vendor, and break-glass design;
  7. workload cross-account trust matrix;
  8. delegated-administrator register;
  9. IPAM/VPC/account allocation;
  10. ingress/egress/inspection/hybrid network diagrams;
  11. DNS authority and Resolver design;
  12. logging/evidence architecture;
  13. security-service coverage matrix;
  14. backup, vault, copy, and restore-test plan;
  15. account vending workflow and acceptance tests;
  16. account closure/divestiture checklist;
  17. platform CI/CD, policy staging, drift, and rollback;
  18. shared-platform service catalog/SLOs;
  19. FinOps allocation/commitment/shared-cost model;
  20. quota and central-capacity register;
  21. exception lifecycle;
  22. acquisition quarantine/integration stages;
  23. five platform failure game days;
  24. ADRs, risks, phased roadmap, and stakeholder approval.

Failure scenarios

Prove response to an SCP blocking emergency logging, a compromised workload role attempting cross-account access, one central inspection AZ failing, an acquired overlapping CIDR, and a new account missing security enrollment. Each test needs expected deny/degrade, telemetry, owner, recovery, and prevention update.

Cost and cleanup

This capstone creates no AWS resources. Model Control Tower/Config, security tools, logging/storage/query, backup/copies, network hubs/inspection/NAT/transfer, DNS, support, CI/CD, shared platforms, and staffing. Sanitize and retain local evidence; there is no cloud cleanup.

Knowledge check

  1. Does an SCP grant access? No; it sets a maximum permission boundary.
  2. Who owns application recovery acceptance? The workload/product owner, using platform capability.
  3. Why use PolicyStaging? To test inheritance and failures before production rollout.
  4. Is centralized networking automatically safer? No; it can create broad blast radius and needs resilient ownership.
  5. What proves account vending? End-to-end identity, policy, logging, security, backup, network, cost, and deployment tests.

Lesson acceptance

All 24 artifacts must trace to scenario requirements. Every centralized capability needs owner, SLO, failure/recovery, capacity, cost, and consumer contract; policies need positive/negative tests; and account lifecycle must be proven from request through safe closure.

Official sources

Advertisement