Lesson 255 · AWS Learning Path

AWS 255: Accounts, organizational units and workload separation

· Published · 12 min read

Labelled process diagram for AWS 255: Workload and ownership requirements to Account and OU boundary to Policy and shared-service relationship to Isolation, billing, and operations evidence, with decision, proof and...

Why this matters to an architect

Putting every system in one AWS account makes identity mistakes, quota exhaustion, billing confusion, and incidents share one blast radius. Creating thousands of accounts without automation creates a different failure: inconsistent baselines, abandoned owners, networking sprawl, and unmanageable cross-account dependencies.

The design task is to place each workload environment and foundational capability at the right boundary. An AWS account is an isolation boundary; an organizational unit is a policy and lifecycle grouping for accounts. Neither automatically creates network paths, copies resources, or moves a workload.

Outcomes

By the end, you can:

  • distinguish organization, OU, account, Region, VPC, workload, environment, and resource boundaries;
  • decide when production, non-production, regulated data, tenants, teams, or shared services need separate accounts;
  • design Security, Infrastructure, Workloads, Sandbox, Deployments, PolicyStaging, Transitional, Exceptions, and Suspended OUs;
  • place log archive, security tooling, network, identity/shared services, deployment, backup, and billing functions safely;
  • map cross-account identity, network, DNS, data, key, logging, backup, and deployment dependencies;
  • compare decentralized VPCs with VPC sharing without confusing ownership responsibilities;
  • design account vending, baseline deployment, drift detection, quota, metadata, and decommissioning workflows;
  • migrate a mixed account incrementally with evidence and rollback.

Begin with precise boundaries

BoundaryWhat it separatesWhat it does not automatically separate
AWS accountIAM namespace, many quotas, billing records, service configuration, API blast radius, resourcesorganization payer, approved sharing, physical facilities, every global service dependency
OUinherited organization controls and lifecycle groupingresources, users, VPCs, invoices, direct network traffic
Regionservice failure/data-placement scope for Regional servicesaccount/IAM/Organizations global control planes
VPCrouting and network policy domainIAM, billing, all service APIs, account ownership
subnetIP/routing/AZ placementa complete security or account boundary
workloadcomponents/data managed for a business outcomenecessarily one microservice or one account
environmentone running instance/stage of a workloadautomatically isolated unless deliberately placed

By default, accounts cannot access one another's resources. Cross-account trust, resource policies, AWS RAM shares, networking, encryption grants, or service integrations must create explicit relationships. This is a valuable fail-closed starting point, but only if teams inventory those relationships.

The default: workload-oriented accounts

A practical starting model is one account per significant workload environment:

Payments-Development account
Payments-Test account
Payments-Production account
CustomerData-Development account
CustomerData-Production account

This lets production have stricter access, change control, backups, detection, quotas, budgets, and network reachability. A developer's experimental resources cannot consume a production account quota or accidentally modify a similarly named production bucket.

This is a default, not a law. Combine small components when they share owner, data classification, lifecycle, access model, controls, recovery objective, cost center, and failure domain. Split when any of those differ materially. A microservice-per-account pattern can be appropriate for strong team/tenant/regulatory independence, but it magnifies account vending, observability, networking, DNS, deployments, quotas, and cost-allocation operations.

Never put production and non-production together merely to save an account. Avoid mixing unrelated workloads with different owners or compliance. Conversely, do not create one account per Lambda function or bucket without a boundary requirement.

The account-boundary decision record

For every placement evaluate:

  • identity and separation of duties: who administers, deploys, reads data, and responds;
  • data: classification, residency, tenant promises, encryption keys, retention, legal hold;
  • blast radius: credential compromise, destructive automation, configuration error, incident containment;
  • quotas and scaling: account/Region quota competition and increase ownership;
  • change and lifecycle: release cadence, maintenance windows, acquisition/divestiture, retirement;
  • resilience: backup account, alternate Region, recovery credentials, shared-dependency failures;
  • network: ingress/egress ownership, segmentation, overlapping CIDRs, inspection, DNS;
  • policy: distinct SCP/RCP/declarative controls, approved services/Regions, exceptions;
  • cost: owner, budget, tags, commitment sharing, showback/chargeback;
  • operations: logs, alerts, patching, inventory, support, automation maturity.

Write which requirement is satisfied by the account boundary and which still needs a resource, VPC, Region, policy, or application control. Accounts reduce scope; they are not a substitute for secure application design.

Foundational OUs and accounts

AWS guidance groups common governance/infrastructure capabilities separately from business workloads.

Security

  • Log Archive account: immutable destinations for organization CloudTrail, Config snapshots, security logs, and retention. Workload administrators should not alter destinations or delete evidence. Restrict human access and protect KMS keys/buckets from accidental teardown.
  • Security Tooling/Audit account: delegated administration and investigation for GuardDuty, Security Hub, Inspector, Macie, Detective, or SIEM integrations as chosen. Read/security-remediation access is distinct from log ownership.

Keep a Control Tower Security OU clean according to its requirements; place additional security workloads in an appropriate separate OU when guidance requires it.

Infrastructure

  • Network account: Transit Gateway/Cloud WAN, egress/inspection, IPAM, Route 53 Resolver, Direct Connect/VPN, or shared VPC ownership where centralized.
  • Shared Services account: directory/DNS/package repositories/operations tools that genuinely serve multiple workloads. Split high-risk or differently owned capabilities instead of creating one “everything shared” account.

Centralization improves consistency but creates a shared failure and privilege domain. Define availability, tenant isolation, capacity, emergency change, service ownership, and a way for workloads to operate safely during failure.

Workloads and deployments

Separate production and non-production accounts and, when their controls differ, OUs. A Deployments/Tooling account may host CI/CD orchestration and artifact promotion, but target roles must be narrow, builds untrusted by default, artifacts signed/scanned, and production promotion independently approved. One compromised deployment account must not become unrestricted organization administrator.

Procedural and experimental OUs

  • Sandbox: budgeted experimentation with restricted data/connectivity and automated expiry.
  • PolicyStaging: representative accounts for testing organization policies/baselines - not application IAM development.
  • Transitional: acquired/migrating accounts under temporary controls with an exit date.
  • Exceptions: rare accounts with approved deviations, compensating controls, owner, expiry, and enhanced monitoring.
  • Suspended: temporary lifecycle holding area with access/cost controls; moving here does not change AWS account State or delete resources.
  • Business Continuity: recovery capabilities/accounts when policy and operational separation justify them.

An OU should exist because accounts below it need a shared control or lifecycle process. An empty hierarchy that mirrors departments but carries no distinct policy adds confusion.

Tenant and regulatory isolation

Separate accounts can strengthen SaaS tenant isolation, sovereign/regulated workloads, or contractual divestiture. Decide whether isolation is pooled, siloed, or hybrid based on tenant risk, scale, cost, noisy-neighbor risk, customization, and operational automation. Account-per-tenant is not automatically compliant; identity, encryption, network, logging, support access, and deletion proof still matter.

AWS partitions (aws, aws-us-gov, aws-cn) have separate accounts/credentials/endpoints and do not support ordinary cross-partition IAM delegation. Region restrictions and data residency need SCP/application/service configuration, not just an OU name.

Networking does not follow the OU tree

Moving an account between OUs does not create/remove routes, VPC peering, Transit Gateway attachments, PrivateLink endpoints, DNS rules, firewall policies, or public access. Maintain a separate network topology and reconcile it with account lifecycle.

Common models:

  • workload-owned VPCs connected through governed transit, preserving account autonomy;
  • centralized inspection/egress with workload VPCs;
  • VPC subnet sharing, where a network owner shares subnets through AWS RAM and participant accounts create supported resources there;
  • PrivateLink or service-specific endpoints for narrow provider/consumer relationships.

In VPC sharing, the owner controls VPC/subnet/route tables/network ACLs/endpoints and pays owner resources; participants own and pay their created resources and control their security groups where supported. Participants cannot view/control everything the owner can. Both sides need quotas, IP capacity, incident, deletion, and change contracts. Shared VPC failure can affect many accounts, while per-account VPCs increase address/routing operations.

Do not use broad peering merely because accounts share an OU. Prefer explicit connectivity requirements, least routes/ports, DNS ownership, flow logs, and egress attribution.

Cross-account data, keys, backups, and logs

For each dependency record provider, consumer, identity, resource policy, KMS key/grant, network path, Region, data classification, cost, availability, and revocation.

  • Central log buckets require service/source conditions, bucket/KMS policies, Object Lock/retention where required, and monitored delivery failures.
  • Cross-account backups need source copy permissions, destination vault/key policy, retention/legal hold, restore testing, and protection against source-account compromise.
  • Shared artifacts/configuration need immutable versions, provenance, malware/signature controls, and target read-only access.
  • Cross-account databases/data lakes require service-specific authorization; a route and IAM role alone may not grant data-plane access.
  • KMS key ownership is an availability decision. A centralized key administrator can protect separation but can also stop many workloads.

Minimize synchronous cross-account dependencies in a workload's critical path. If shared DNS, directory, egress, artifact, or key service fails, the recovery design must say what production does.

Quotas and scale

Many quotas are per account per Region, so account separation reduces competition. Some quotas are global, shared through a networking hub, organization-level, service-managed, or not adjustable. A new account often starts with default quotas and no request history. Account vending must request needed quotas before deployment and track approval lead time.

Do not multiply accounts to evade intended service controls or abuse limits. Central resources such as Transit Gateway, NAT, IPAM pools, shared subnets, resolver endpoints, build fleets, and security ingestion have their own aggregate capacity and cost bottlenecks.

Account vending and landing-zone automation

An account is not ready when CreateAccount returns success. A vending workflow should collect:

business/workload/environment/owner
unique durable root email and contacts
requested OU and data classification
Regions/CIDRs/connectivity/DNS
identity groups and break glass
logging/security/config/backup baselines
budgets/tags/cost center/commitment policy
service quotas and lifecycle/expiry

Then asynchronously create/invite, place the account, deploy baselines through infrastructure as code, enroll security services, configure identity/network, run acceptance tests, and publish metadata to a catalog. Make retries idempotent and detect partial failure.

AWS Control Tower can establish a landing zone, governed OUs, controls, Account Factory, and account baselines. Account Factory for Terraform (AFT) can add a Git/Terraform workflow. These are not substitutes for ownership and design. Direct Organizations moves can create Control Tower governance drift unless current auto-enrollment/baseline behavior is configured and prerequisites match. Inspect landing-zone/OU/account baseline status after every move.

Store authoritative account metadata outside names alone. Names and emails change; account IDs are stable identifiers but sensitive in shared evidence. Track owner, purpose, environment, data class, OU, contacts, baseline version, network, budget, state, exception expiry, and decommission status.

Migration from a mixed account

You cannot move an arbitrary running resource by moving its account to another OU. If one account contains production, test, and shared services needing different controls, create target accounts and migrate/recreate resources service by service.

  1. Discover resources, identities, data flows, DNS, certificates, keys, quotas, costs, and unsupported/move-limited services.
  2. Define target account, Region/VPC, ownership, baseline, and rollback for each workload.
  3. Rebuild stateless infrastructure from IaC; replicate data with service-supported methods.
  4. Establish least-privilege cross-account transition paths and dual logging.
  5. Test data consistency, performance, security, backup/restore, and forbidden paths.
  6. Cut over DNS/traffic with an explicit rollback window.
  7. remove temporary trust/connectivity, retain required evidence, and decommission old resources only after acceptance.

Certificates, encrypted snapshots, domains, Marketplace products, reserved commitments, IP addresses, service-linked resources, and resource policies have service-specific migration constraints. Inventory first; “copy everything” is not a plan.

Practical lab

Download the AWS255 account-topology pack. It contains:

  • ACCOUNT_PLACEMENT_WORKBOOK.md with 15 placement decisions;
  • SUPPLIED_BOUNDARY_CASES.md with six failures;
  • CROSS_ACCOUNT_DEPENDENCY_REGISTER.md;
  • ACCOUNT_VENDING_AND_MIGRATION_CHECKLIST.md;
  • ANSWER_DIRECTIONS.md, opened after independent analysis.

Submit an OU/account tree, boundary decision record, dependency graph, RACI, quota/network/cost plan, account-vending acceptance, and mixed-account migration wave. The no-create track is complete.

Read-only inspection

With authorized Organizations and account-catalog access:

aws organizations list-roots
aws organizations list-children --parent-id r-REDACTED --child-type ORGANIZATIONAL_UNIT
aws organizations list-children --parent-id ou-REDACTED --child-type ACCOUNT
aws organizations list-accounts \
  --query 'Accounts[].{Id:Id,Name:Name,State:State,Method:JoinedMethod,Joined:JoinedTimestamp}'
aws organizations list-tags-for-resource --resource-id 111122223333
aws service-quotas list-service-quotas --service-code ec2 --region ap-south-1

Recursively traverse every OU and paginate every list; one parent query is not an inventory. Redact IDs, emails, names, CIDRs, owner data, and policy identifiers. Organizations output does not prove network paths, deployed baselines, resource inventory, quota readiness, or cost ownership - join those evidence sources.

Troubleshooting matrix

SymptomEvidenceLikely issue
prod change damages test or reverseaccount/resource inventoryenvironments share boundary
account moved but connectivity unchangedroutes/TGW/RAM/DNS/security policiesOU does not control network
new account cannot deploy scaleService Quotas by account/Regionvending missed quota lead time
Control Tower account becomes driftedbaseline/enrollment/provisioned-product statusunmanaged OU move/baseline mismatch
central service outage affects all appsdependency graph/SLO/failover testshared blast radius unmanaged
logs disappear during incidentdestination/KMS/service policy + delivery statuscentral logging dependency failure
cross-account restore failsvault/key/copy/role policies and restore testbackup was never end-to-end tested
suspended account still runs/incurs costaccount State, resources, access, budgetOU move is not shutdown/closure
one workload needs conflicting OU controlsresource inventory inside accountaccount contains mismatched workloads
cost cannot be assignedcatalog, account/usage tags, CURownership metadata missing

Cost and lifecycle

AWS accounts and OUs have no direct hourly charge. More accounts increase operational tooling, logs, Config rules, security products, NAT/endpoints, support allocation, backup, network transfer, and human effort. Fewer accounts can increase incident and cost-allocation risk. Compare total operating model, not account count.

Before retirement: stop intake; preserve logs/legal holds; remove production traffic; inventory commitments/domains/data/keys/backups; revoke workforce/pipeline/vendor access; stop resources; verify invoices; move to Suspended when appropriate; wait through retention; then obtain owner/security/legal/finance approval for closure. Prove that shared consumers no longer depend on the account.

Knowledge check

  1. Why is an account a stronger boundary than an OU?
  2. What criteria justify splitting two workloads or environments?
  3. When is combining small components reasonable?
  4. Why must production and non-production normally use separate accounts?
  5. Which functions belong in Log Archive versus Security Tooling?
  6. How can a shared services account become an enterprise blast radius?
  7. Why does an OU move not alter network topology?
  8. How do VPC owner and participant responsibilities differ?
  9. Which dependencies must accompany a cross-account data path?
  10. Why do new accounts need quota readiness before deployment?
  11. What is Control Tower governance drift, and how can account moves trigger it?
  12. Why can a mixed account require resource migration rather than an OU move?
  13. What distinguishes Suspended OU placement from AWS account suspension/closure?
  14. Which evidence proves an account is ready for use or retirement?

Lesson acceptance

Pass requires defensible placement of all 15 scenarios; OU/account tree based on control differences; production/non-production and foundational separation; cross-account identity/network/data/KMS/log/backup map; VPC sharing ownership; shared-service failure treatment; quota and FinOps model; automated vending/baseline/drift acceptance; six supplied diagnoses; staged mixed-account migration and rollback; lifecycle proof; and no-create or verified cleanup evidence.

Official sources

Advertisement