AWS 265: Architecture: model an organization, OUs, accounts, SCPs, identity, logging, and shared services
Why this capstone matters
An AWS architect is not finished after selecting Organizations, Control Tower, Identity Center, Transit Gateway, CloudTrail, or Backup. A production foundation is a connected operating system: account boundaries affect policies and billing; identity depends on lifecycle and emergency recovery; network/DNS/key/logging platforms create shared blast radius; controls require exceptions and evidence; account vending must deliver a usable product; migration and decommissioning must preserve data, audit, cost, and recovery.
This capstone integrates AWS254–AWS264. You will receive an imperfect enterprise case, derive requirements, make explicit architecture decisions, prove every domain contract, diagnose supplied failures, and defend tradeoffs. A copied AWS reference diagram is not a submission.
Outcomes
By the end, you can:
- convert business, regulatory, workload, migration, resilience, and operating needs into testable requirements;
- design an organization/OU/account model based on common controls and lifecycle;
- map preventive, detective, proactive, and responsive controls without overlap;
- integrate workforce/workload identity, emergency access, and service delegation;
- separate immutable evidence custody from security analysis/remediation;
- design network, DNS, IPAM, hybrid, ingress/egress, endpoints, and inspection;
- define shared-service contracts for RAM, keys, backup, deployment, quotas, and licenses;
- build account-vending, exception, incident, migration, and decommission workflows;
- model platform cost, quotas, SLOs, RACI, dependency failure, and recovery;
- produce implementation waves, acceptance tests, evidence, rollback, and an executive decision record.
The supplied enterprise
Northwind Health & Retail has 38 AWS accounts, expects 120 in three years, and operates in ap-south-1, eu-west-1, and us-east-1. It runs payment, health, e-commerce, analytics, and internal workloads. Two data centers connect to AWS. An acquisition contributes eight accounts with unknown root contacts, local IAM users, overlapping CIDRs, public resources, and incomplete logs.
The current organization has Production and NonProduction OUs, broad permanent administrators, one central network, mutable logs in the security account, inconsistent backups, console deployments, optional tags, and no account catalog. A single egress/DNS path serves all Regions. Teams want self-service accounts in two business days. Audit requires seven-year evidence for regulated systems; executives require 30-minute recovery for checkout and four-hour recovery for health analytics. Finance requires 95 percent allocatable cost. Platform staffing is six engineers with 24x7 on-call shared across security/network/operations.
Additional constraints and conflicting stakeholder requests are in the downloadable case pack. Treat ambiguities as assumptions requiring owner validation - not permission to ignore them.
Required architecture method
Use this sequence:
stakeholders and facts
-> measurable requirements and assumptions
-> trust/failure/data boundaries
-> candidate options and ADRs
-> domain service contracts
-> dependency and threat/failure analysis
-> controls, evidence, cost, and ownership
-> implementation/migration waves
-> positive, negative, failure, recovery, and rollback tests
-> operating acceptance and lifecycle
Maintain requirement IDs. Every major diagram, ADR, control, service contract, test, and cost line must reference them. Every requirement needs an accountable owner and acceptance evidence. If no evidence can prove it, the requirement is not operationalized.
1. Requirements and constraints dossier
Capture:
- business units, products, owners, criticality, users, transaction/load forecasts;
- data classes, residency, payment/health/legal/audit obligations and retention;
- environments, SDLC, deployment frequency, autonomy, M&A/divestiture;
- RTO/RPO/SLO, dependency and degraded-mode requirements;
- identity source, workforce/vendor/workload identity, MFA/device/elevation;
- hybrid/sites/partners/internet, DNS, IPv4/IPv6, inspection and latency;
- account/Region growth, quotas, licensed software, staffing/skills/support;
- chargeback/showback, budget, commitment, Marketplace and allocation goals;
- brownfield inventory, cutover windows, rollback and coexistence.
Classify each as hard constraint, preference, assumption, risk, or open decision. Resolve contradictions or show the escalation/decision owner.
2. Organization, OU, and account architecture
OUs group accounts with common policy and lifecycle; they do not mirror the reporting org chart. Accounts provide stronger identity/resource/quota/billing/blast-radius boundaries.
At minimum evaluate:
Management (organization control plane only)
Security: clean Control Tower Security OU where applicable
Security extensions: Log Archive custody, Security Tooling/Audit
Infrastructure: Network, Shared Services, Operations, Backup/Recovery
Deployments: build/artifact/release orchestration
Workloads: Production and NonProduction by policy need/workload portfolio
Sandbox: budgeted, restricted, expiring
PolicyStaging: representative governance canaries
Transitional: acquisitions/migrations under quarantine
Exceptions: approved deviations with expiry
Suspended: controlled closure/incident state
BusinessContinuity: if separate recovery ownership is justified
Keep Security OU compatible with chosen Control Tower requirements; additional security accounts may need an adjacent OU. Explain each account boundary, who administers it, what may not run there, policy set, Regions, network, evidence, cost, and lifecycle. Define durable root email/phone/MFA/contact custody and management-account constraints.
Produce an account catalog schema: account ID/name, owner/service, environment, data class, OU, root/contacts, Regions, network/CIDR/DNS, identity groups, baseline version, delegated services, backup tier, log coverage, quotas, cost center, exceptions, creation/expiry/decommission state.
3. Landing-zone implementation ADR
Compare AWS Control Tower, Control Tower plus extensions, Landing Zone Accelerator, and a custom platform. Score managed ownership, requirements fit, brownfield enrollment, regulated patterns, release/upgrades, drift, skill, cost, rollback, support, and decommissioning.
Name one authoritative controller per managed resource. Define boundaries between Control Tower, LZA/CfCT/AFT, StackSets, Terraform/CDK/CloudFormation, security services, and console emergency changes. No design may let several controllers overwrite the same trail, Config rule, SCP, VPC endpoint, or IAM role.
4. Governance and policy architecture
Build a control catalog linking requirement to:
- SCP maximum identity permissions;
- RCP/resource-side perimeter where supported;
- declarative/tag/backup policies and inheritance;
- Control Tower preventive/detective/proactive controls;
- IAM boundaries/session policies/permission sets;
- Config/Security Hub/custom detection;
- network/service/KMS/data-perimeter controls;
- approved automatic or human remediation;
- exception, expiry, rollback, evidence and cost.
Show effective-policy paths from root through OUs to representative accounts. Include Region controls, root/management exceptions, service-linked roles/delegates, policy size/quotas, canary rollout, CloudTrail denial diagnosis, and recovery from a bad SCP. “We attach FullAWSAccess plus denies” is not a complete governance design.
5. Identity, delegation, and emergency access
Trace HR/IdP → MFA/device → SAML → SCIM → Identity Center group/attribute → permission set/account assignment → generated role/STS session → all policy layers → CloudTrail/service event.
Define baseline roles for developer, read-only, security, network, backup, billing, auditor, support, deployment, vendor, and platform. Separate short production elevation, approvals, source context, expiry, and active-session containment. Give the management account dedicated rare access and tightly governed direct assignments.
Separate workforce and workload identity. Pipelines use OIDC/temporary roles; compute uses service roles; no human keys in automation. Build joiner/mover/leaver/vendor lifecycle and access reviews.
Inventory every trusted service, service-linked role, and delegated administrator with Regions, permitted/management-only operations, coverage, data/billing, auto-enrollment, and deregistration. Design emergency paths for IdP, Identity Center Region, network/DNS/device, bad policy, and management compromise. Root recovery is separate. Test custody, allowed/denied action, alarms, closure, rotation, and review.
6. Evidence, security operations, and data protection
Separate Log Archive as protected/immutable source custody from Security Tooling as delegated administration, findings, SIEM, investigation, and response. Define organization CloudTrail management/data/network selectors, Config, identity, network/DNS/edge/firewall, OS/container/database/application/CI/CD, backup, and security-service telemetry.
For every source specify accounts/Regions, destination, source-conditioned S3/KMS policy, schema/partition, sensitivity/redaction/residency, latency/SLO, integrity/digest, versioning/Object Lock/legal hold, replication, retention, query, alert, failure alarm, cost, and restore/export.
Model GuardDuty, Security Hub CSPM home/linked Regions and central policies, Inspector, Macie, Detective, Access Analyzer, Security Lake/CloudWatch/SIEM as needed. Findings never replace raw evidence. Define incident roles, case-scoped query/export, containment authority, evidence preservation, communications, and post-incident review.
Encryption architecture covers key ownership/account/Region, administrator/users, key policy/grants, cross-account access, rotation, replicas, deletion protection, quotas, cost, and loss recovery. Data perimeter controls need service-to-service exceptions and negative tests.
7. Network, DNS, IPAM, and shared resources
Start with required/forbidden flows and forward/return paths. Choose per-account/shared VPC portfolio, TGW versus Cloud WAN, route domains/segments, Direct Connect/VPN resilience, IPv4/IPv6 and delegated IPAM, hybrid/private DNS, endpoints/PrivateLink/Lattice, public ingress, egress/NAT, firewall/GWLB inspection, and logging.
For centralized components quantify blast radius, AZ/Region failure, stateful symmetry/appliance mode, source attribution, capacity, policy size, quotas, latency, and per-hop processing/transfer. Define Regional degraded operation; one global egress/DNS/firewall is not resilient by declaration.
Use RAM contracts for subnets, TGW/core, Resolver rules, IPAM pools, licenses, and other supported resources: owner, principal, managed permission/version, Region/invitation, consumer IAM, service responsibility, cost, association state, monitoring, revocation and deletion order.
8. Backup, resilience, quotas, and licenses
Map workload RTO/RPO to backup/restore, pilot light, warm standby, or multi-site. Design cross-account/Region vaults, keys, immutability, organization backup policies, role/vault prerequisites, protected-resource reconciliation, copies, isolated restore, application consistency, cutover/failback, and evidence.
Recover the platform too: IdP/emergency access, management contacts, organization/IaC configuration, pipelines/artifacts, DNS/IPAM/network, keys, account catalog, logs/query, security services, and runbooks.
Build quota forecasts for normal peak, deployment surge, one-AZ loss, Region failover, restore, acquisition, and incident telemetry. Distinguish administrative quota from actual capacity/address/application/license limits. Request early and prove applied values plus bounded scale/failover.
Translate license terms only with procurement/legal approval; reconcile entitlement, License Manager rule, AMI/resource discovery, running/stopped/host use, invoice/vendor record, hard/soft enforcement, DR rights, and offboarding.
9. Deployment and software supply chain
Separate untrusted build from production release. Require protected source, review, ephemeral runner, temporary OIDC role, dependency/secret/malware/IaC scans, SBOM, signed provenance/artifacts, immutable repositories, environment approval, narrow deployment role, CloudTrail source context, deployment verification, rollback, and break-glass change reconciliation.
The landing-zone repository is production code. Validate policy schemas/effective results, synthesized IaC, routes, permissions, managed-resource conflicts, account/Region matrix, cost and quotas. Deploy to PolicyStaging, canary accounts/OUs/Regions, then bounded batches with failure thresholds and automatic/manual rollback.
10. Account vending and service catalog
The request captures owner/sponsor, workload/environment, data/criticality, cost center, contacts, OU/policies, Regions, VPC/IP/DNS, identity groups, deployment, quotas/licenses, backup/RTO/RPO, logging/security, budget, exceptions, and expiry.
Workflow:
validate -> create/invite -> catalog/tag/contact/OU
-> baseline/control -> identity -> network/DNS -> log/security
-> backup/budget/quota/license -> positive/negative/recovery acceptance
-> owner handoff and SLO
Make it asynchronous, idempotent, retryable, observable, and recoverable from partial failure. Define dead-letter/manual repair, lead time, status, cleanup, and no-ready-until-all-gates behavior.
11. FinOps and platform capacity
Design payer/billing ownership, account/resource tags, tag policies and activation/backfill, account tags, Cost Categories, unallocated/shared charge rules, CUR/data exports, budgets/anomaly detection, commitments/discount sharing, Marketplace/support/credits, showback/chargeback, and unit economics.
Model platform cost by account × Region × resource/change/event/GB/query. Include Config/security services, CloudTrail/data events, logs/SIEM, NAT/TGW/Cloud WAN/endpoints/firewall/transfer, keys, backup/copies/restore tests, pipelines, support, licenses, and people. Reconcile allocation to total bill and assign Unallocated an owner/SLO.
Track central quotas and capacity: accounts, policies/attachments, Identity Center assignments, StackSets, IPAM, TGW/Cloud WAN/routes, Resolver, endpoints/NAT/firewall, log/KMS APIs, security members/findings, backup jobs/vaults, pipeline concurrency and support throughput.
12. Operating model, SLOs, and dependency analysis
Create RACI and service contracts for platform product, management/root, identity/IdP, Organizations/controls, security operations, evidence custody, keys, network/DNS/IPAM, shared services, backup/DR, deployment, account vending, FinOps, compliance/legal/procurement, workload teams, and vendor/AWS support.
Define SLOs and escalation for account delivery, access/elevation/revoke, policy change/exception, DNS/network, security alert, log delivery/query, backup/restore, quota, deployment, incident, and decommission. Six engineers cannot own every 24x7 central dependency without automation, scope decisions, and support escalation.
Build a dependency graph. Identify components whose failure affects three or more domains, common credentials/keys/DNS/routes/pipelines, and circular recovery dependencies. For each specify blast radius, detection, degraded mode, RTO/RPO, recovery authority, and game day.
13. Brownfield, acquisition, and migration plan
Discovery precedes enrollment: ownership/root/contacts/payer, accounts/OUs/policies, IAM/users/keys, Regions/resources/public exposure, network/CIDR/DNS, data/keys, trails/Config/security, backup, pipelines, quotas/licenses, cost/commitments, drift and unsupported resources.
Acquisitions enter Transitional quarantine with constrained identity/trust/network, immediate evidence onboarding, threat review, IP/DNS collision analysis, and a dated exit. Migrate in waves: platform test, foundational accounts, low-risk non-production, representative production canary, dependency groups, then legacy-control retirement. Every wave has entry/exit criteria, coexistence, data consistency, rollback trigger, owner and communications.
14. Exceptions, incidents, and decommissioning
Exceptions need control/requirement, reason, risk owner, compensating controls, exact scope, evidence, issue/expiry, review, automatic detection, and removal. An Exceptions OU is not an uncontrolled bypass.
Account decommission inventories business/data owner, records/legal hold, backups/recovery points, keys/secrets/certificates/domains, identity/users/sessions, shared resources/RAM/ENIs/routes/DNS, delegates, pipelines/artifacts, licenses/commitments/Marketplace/support, billing/credits, public endpoints, quotas, and root contacts. Freeze, archive/migrate, prove consumers removed, revoke, close through supported organization workflow, retain evidence and update catalog. Suspension is not deletion.
Capstone evidence gates
| Claim | Minimum authoritative evidence | What is not sufficient |
|---|---|---|
| account is ready | catalog plus baseline, identity, network/DNS, logs/security, backup, budget, quota and positive/negative acceptance | account-created status |
| policy is safe | effective policy for representative paths, canary behavior, denied/allowed events and independent rollback | attachment list or JSON lint |
| identity is controlled | IdP/SCIM, assignment/provisioning, session, target action, revoke and emergency test timeline | portal screenshot |
| logs are durable | source canary, destination delivery, KMS, digest/integrity, immutability/retention, query and failure alarm | bucket or organization trail exists |
| network path works | DNS, forward and return routes, security/inspection, active probe and target application evidence | attachment available or flow-log ACCEPT |
| resource share works | owner/principal/permission/version/association plus consumer IAM and service action | RAM share status alone |
| recovery meets objective | isolated restore, keys/identity/network/DNS/quota, data consistency, cutover/failback and measured RTO/RPO | successful backup/copy job |
| deployment is trustworthy | reviewed source, scans, immutable signed artifact/provenance, narrow release session, verification and rollback | pipeline succeeded |
| cost is allocated | authoritative total reconciled to direct/shared/unallocated categories with owners | percentage of only tagged resources |
| account is retired | consumers/data/evidence/keys/contracts/identity/network/billing cleared and closure retained | empty console or suspended OU |
Required deliverables
Download the AWS265 enterprise governance capstone pack. Submit:
- executive summary and requirement/assumption/risk register;
- six diagrams: organization/accounts, identity, network/DNS, evidence/security, deployment, and backup/recovery;
- landing-zone and major domain ADRs;
- account catalog, control/evidence catalog, delegate register, RAM contracts, and RACI/SLO;
- account-vending and exception/decommission workflows;
- cost/allocation/quota/license model;
- phased brownfield/M&A migration with rollback;
- ten supplied failure diagnoses and six game-day plans;
- requirement-to-design-to-test traceability matrix;
- oral defense answers and residual-risk signoff.
Read-only evidence audit
No command proves the platform. With approved roles, combine and paginate:
aws organizations describe-organization
aws organizations list-accounts --query 'Accounts[].{Id:Id,Name:Name,State:State}'
aws organizations list-aws-service-access-for-organization
aws organizations list-delegated-administrators
aws sso-admin list-instances --region ap-south-1
aws cloudtrail describe-trails --include-shadow-trails false --region ap-south-1
aws configservice describe-configuration-aggregators --region ap-south-1
aws ram get-resource-shares --resource-owner SELF --region ap-south-1
aws ec2 describe-transit-gateways --region ap-south-1
aws backup list-backup-vaults --region ap-south-1
aws service-quotas list-requested-service-quota-change-history --region ap-south-1
Redact organization/account/policy/identity/network/data/key/log/cost/license details. Correlate configuration with behavior, negative tests, telemetry, cost and recovery.
Supplied integration failures
Diagnose these without granting broad administrator access or disabling governance:
- Control Tower is green; workload launch fails on quota, DNS loop and KMS grant.
- Security analysts can delete original seven-year logs and suppress alerts.
- One egress/DNS/inspection path fails both production Regions.
- Control Tower, LZA, Terraform and console automation overwrite the same resources.
- Pull-request builds can assume production deployment and mutate artifacts.
- Backup copies succeed; recovery lacks keys, identity, DNS, network, quota and app config.
- Acquisition connects before root/contact/threat/CIDR/IAM/log discovery.
- Broad SCP blocks a service-linked role and the same operator needed for rollback.
- Tag-selected backups and cost allocation fail after an unauthorized tag change.
- Emergency access depends on the compromised IdP and creates no independent alert.
For each show failed requirement/domain contract, evidence order, root cause, containment, reversible correction, owner, cost, positive/negative/failure/recovery tests, and residual risk.
Architecture review and oral defense
Expect reviewers to change a constraint: double account growth, remove one Region, acquire overlapping networks, lose the IdP, require legal hold, reduce staffing, block a service, or halve cost. Explain which decisions change and which invariants remain.
Defend at least:
- why each OU/account exists and what it isolates;
- who can alter identity, policies, logs, keys, routes and pipelines;
- how a workload reaches dependencies and how forbidden paths fail;
- how evidence survives operator compromise;
- how production recovers when central services fail;
- how costs and quotas scale;
- how a brownfield account enters and exits Transitional;
- how a bad platform release rolls back;
- which residual risks executives accept.
Knowledge check
- Why should OUs follow common controls rather than the org chart?
- Which workloads must never run in the management account?
- How do preventive, detective and responsive controls combine?
- Why separate Log Archive and Security Tooling?
- Which identity/session paths survive an IdP or policy failure?
- How can a central network/key/DNS/pipeline become correlated failure?
- What makes a RAM/shared-service contract complete?
- Why do backup jobs and green platform dashboards not prove recovery?
- Which gates make an account ready for workload handoff?
- How do platform costs scale and reconcile to the bill?
- Why must M&A accounts begin in quarantine?
- What evidence makes decommissioning safe?
- How does traceability expose architecture gaps?
- Which residual risks require executive acceptance?
Capstone acceptance
Pass requires every supplied requirement classified and traced; explicit assumptions/open decisions; complete diagrams/ADRs/contracts/catalogs; effective policy and identity evidence; protected logs/security operations; route/DNS/IPAM and failure paths; backup/restore; pipeline/account product; cost/quota/license; RACI/SLO; all ten diagnoses; migration/decommission; game days; rollback; and oral defense. Any missing owner, evidence, denied-path test, recovery path, cost, or lifecycle for a critical requirement is a revision - not a pass.