AWS 324: Designing for organizational complexity
Why this lesson matters
Enterprise architecture must fit real ownership. Business units, products, environments, countries, acquisitions, vendors, budgets, regulations, and support teams rarely share identical needs. A technically elegant platform fails if nobody can approve access, pay the bill, operate an incident, grant an exception, or retire an account.
AWS accounts are resource, isolation, quota, and billing boundaries. Organizational units group accounts for governance; they are not application folders. AWS Organizations, Control Tower, IAM Identity Center, delegated administration, and organization-integrated services create an operating system for the cloud estate, but only when responsibilities and lifecycle processes are explicit.
Learning outcomes
By the end, you can:
- derive account and OU boundaries from security and operational needs;
- distinguish management, security, shared-service, network, and workload ownership;
- design workforce and workload access across accounts;
- compare preventive, proactive, detective, and responsive controls;
- define delegated administration without uncontrolled central privilege;
- design account vending, funding, exceptions, incidents, and closure;
- integrate acquired accounts and regulated geographies safely;
- produce an evidence-backed organizational architecture.
1. Organizational requirements map
Before drawing OUs, map:
- legal entities, jurisdictions, contracts, and data residency;
- business units, product teams, platform teams, and vendors;
- production, pre-production, sandbox, security, and shared services;
- risk classifications and regulatory regimes;
- funding, cost allocation, commitments, and chargeback/showback;
- identity sources, joiner/mover/leaver ownership, and emergency access;
- support hours, incident authority, and recovery ownership;
- mergers, divestitures, temporary programs, and account retirement.
An organization chart changes frequently and is usually a poor OU blueprint. Organize OUs around stable policy and operational differences. Apply controls to OUs where practical and avoid a hierarchy so deep that policy inheritance becomes difficult to reason about.
2. Accounts as boundaries
Use separate accounts to isolate workloads and teams, reduce blast radius, separate billing, and obtain independent quota pools. Common account categories include:
| Account category | Typical purpose | Important boundary |
|---|---|---|
| Management | Organizations and consolidated billing | Run no normal workloads; tightly restrict access |
| Log archive | Immutable centralized log destinations | Security-owned write/read separation |
| Audit/security tooling | Security administration and investigation | Delegated controls and protected evidence |
| Network | Central connectivity, DNS, inspection, IPAM | Shared-path change and incident ownership |
| Shared services | Directory, CI/CD, artifacts, observability | Consumer isolation and dependency SLOs |
| Production workload | One or a small set of related workloads | Product ownership and blast radius |
| Non-production | Development, test, staging | Distinct controls and no production data by default |
| Sandbox | Time/cost-bounded experimentation | Strong egress, service, and budget guardrails |
| Quarantine/suspended | Acquired or noncompliant accounts | Restricted connectivity and investigation |
Do not host workloads in the management account. Its privileges and trust relationships create organization-wide risk. Keep root-user credentials protected, enable contact and recovery governance, and use controlled break-glass procedures.
3. OU design and policy inheritance
An OU is a policy attachment and administration boundary. Example top-level structure:
Root
|- Security
|- Infrastructure
|- Workloads
| |- Production
| |- NonProduction
|- Sandbox
|- PolicyStaging
|- Suspended
Nested OUs can represent materially different controls, such as regulated production, but start small. An account belongs to one OU at a time and inherits applicable organization policies from root and parents.
Service control policies (SCPs) set maximum permissions for principals in member accounts; they do not grant permissions. IAM policies and resource policies still need to allow an action. SCPs generally do not constrain the management account, another reason to protect it. Resource control policies (RCPs) establish maximum permissions for supported resources, but also do not grant access. Verify supported services, principals, exceptions, and evaluation behavior before deployment.
Test restrictive policies in a policy-staging OU with noncritical accounts. Preserve required service-linked roles and organization service access. A broad deny can block logging, patching, backup, incident response, or account recovery.
4. Landing zone and Control Tower
A landing zone is the governed multi-account foundation. AWS Control Tower orchestrates Organizations, IAM Identity Center, Config, CloudTrail, account provisioning, and controls to establish and govern one. It does not eliminate architecture decisions, inherited estates, custom networks, or operating ownership.
Plan home Region, governed Regions, log/audit accounts, OU enrollment, existing-account registration, identity, account factory, network model, baseline configuration, control rollout, and drift remediation. Understand which resources Control Tower manages before modifying or importing them.
Controls can be:
- preventive: block disallowed actions;
- proactive: evaluate infrastructure before provisioning where supported;
- detective: identify deployed noncompliance;
- responsive: trigger defined remediation or incident handling.
Controls need owners, evidence, severity, remediation SLO, exception path, and false-positive handling. A dashboard showing green is not proof that every workload risk is controlled.
5. Identity and delegated administration
Federate workforce access through a centralized identity provider and IAM Identity Center where appropriate. Map groups to permission sets, accounts, duration, and approval. Avoid long-lived IAM users for workforce access. Separate human roles from workload roles and use short-lived credentials.
Define:
- joiner, mover, and leaver synchronization;
- privileged and read-only permission sets;
- just-in-time or approved elevation;
- break-glass identities protected by strong authentication and tested recovery;
- vendor identities with sponsor, expiry, scope, and review;
- session logging and access-review evidence.
Delegated administrators let specialist accounts manage supported organization-integrated services without routine management-account use. Delegation is service-specific and powerful. Record service, delegated account, allowed operators, data visibility, failure impact, revocation, and incident owner.
6. Shared platforms without shared confusion
Centralization can improve consistency, but creates dependencies and bottlenecks. For network, DNS, security tooling, CI/CD, artifacts, observability, and data platforms define producer, consumers, onboarding contract, authorization, quotas, SLO, cost allocation, change windows, incident escalation, and exit path.
AWS Resource Access Manager can share supported resources across accounts. Sharing does not replace resource policies, security groups, routes, data authorization, or cost ownership. Prefer explicit contracts over informal cross-account access.
Decide where autonomy is safe:
- central platform owns foundation and mandatory controls;
- product teams own workload code, data behavior, on-call response, and cost within guardrails;
- security owns risk policy and investigation, not every application deployment;
- finance owns commercial policy while teams own explainable usage;
- exception authority is separate from requesters for material risks.
7. Account lifecycle
An account-vending process should collect owner, business purpose, environment, data class, cost center, support group, Region needs, network pattern, and expiration. Automation applies baseline contacts, tags, budgets, logging, Config, security integrations, identity assignments, backup policy, DNS/network attachment, and inventory registration.
Lifecycle stages are requested, approved, provisioned, validated, active, restricted, suspended, and closed. Offboarding must address legal hold, backup, export, DNS, certificates, identities, network, shared resources, commitments, marketplace products, support cases, and the account-closure waiting period. Never close an account merely because it appears unused.
8. Mergers, acquisitions, and regulated geography
Treat an acquired AWS estate as untrusted until assessed. Inventory organization membership, payer commitments, root/contact control, identity providers, public exposure, logging, encryption keys, domains, support, legal obligations, and critical dependencies. Place accounts in a restricted transition model before normal connectivity.
Account transfer between organizations changes billing and can affect organization policies, delegated services, shares, licenses, support, and commitments. Build coexistence, migration, rollback, and evidence plans. For divestiture, separate identities, data, encryption keys, domains, contracts, and shared services before transfer.
Geographic controls must identify exact data and legal entities, not simply create one OU per country. Validate Region enablement, service availability, cross-border logs, backups, support access, and central-security visibility.
9. Read-only discovery
These commands require organization or delegated permissions and may correctly return AccessDenied:
aws sts get-caller-identity --query Arn --output text
aws organizations describe-organization --output json
aws organizations list-roots --output table
aws organizations list-organizational-units-for-parent --parent-id r-example --output table
aws organizations list-accounts --output table
aws organizations list-policies --filter SERVICE_CONTROL_POLICY --output table
aws sso-admin list-instances --output table
aws ram get-resource-shares --resource-owner SELF --output table
Never paste real account IDs into course submissions. CLI inventory shows configuration, not whether ownership, incident response, access review, or account closure works.
10. Failure diagnosis
| Symptom | Possible cause | Evidence path |
|---|---|---|
| Allowed IAM action is denied | SCP/RCP, permissions boundary, session, resource policy, or KMS policy | Policy evaluation and CloudTrail |
| New account lacks controls | Provisioning failed, OU not governed, baseline drift | Account Factory/Control Tower events and Config |
| Central security has gaps | Region/service not enabled or delegation incomplete | Organization configuration and coverage inventory |
| Team cannot deploy | Guardrail is too broad or exception expired | Deny context, policy ancestry, change record |
| Shared network outage affects many accounts | Central blast radius and change failure | Routes, attachments, change timeline, telemetry |
| Bill is unowned | Account metadata/cost allocation missing | Payer data, tags, account registry |
Diagnose explicit deny before adding permissions. Do not move a production account between OUs as an ad hoc test because policy inheritance changes immediately.
11. Guided workshop: acquisition across two countries
Design for three business units operating in India and the UK after acquiring a company with 18 AWS accounts. Produce:
- stakeholder and decision-rights map;
- account inventory and confidence gaps;
- target account categories;
- target OU hierarchy with policy rationale;
- management/log/audit account ownership;
- workforce identity and permission-set model;
- workload identity and cross-account trust model;
- preventive/proactive/detective/responsive control catalog;
- delegated-administrator register;
- network, DNS, security, CI/CD, and observability ownership map;
- cost allocation, commitment, and funding model;
- account vending and baseline acceptance process;
- exception and expiry process;
- acquired-account quarantine and integration stages;
- incident-command and break-glass plan;
- account decommission/divestiture checklist;
- phased rollout with policy-staging tests;
- risk register, metrics, and review cadence.
Cost and cleanup
This design lesson creates no resources. Model Control Tower/Config logging, security services, support, central networking, NAT/inspection, inter-Region and inter-AZ transfer, shared platforms, and operational staffing. Central billing does not remove per-team accountability.
Knowledge check
- What does an SCP grant? Nothing; it limits maximum available permissions.
- Why avoid workloads in the management account? Its organization-wide privileges create exceptional blast radius.
- Should OUs mirror the company org chart? Usually no; use stable governance and operational differences.
- What does delegated administration solve? Routine service administration from a specialist account, reducing management-account use.
- Why use a policy-staging OU? To test inherited controls safely before broad enforcement.
Lesson acceptance
Submit all 18 artifacts. Every account and platform needs business, technical, security, cost, and incident ownership. OU and policy choices must follow stable control needs, access must be federated and reviewable, acquisitions must use staged trust, and lifecycle must cover provisioning through verified closure.