AWS 258: Landing zone architecture
Why this matters to an architect
A landing zone is the operated foundation into which workloads arrive. It combines organization/account structure, identity, network, DNS, logging, security, controls, data protection, backup, deployment, cost management, account vending, incident response, and lifecycle automation.
Buying or deploying a landing-zone tool is not the same as having a production-ready landing zone. The architect must translate business, regulatory, migration, resilience, and operating requirements into boundaries and service contracts, then prove that a newly vended account can run safely and recover when central services fail.
Outcomes
By the end, you can:
- define a landing zone and distinguish platform capabilities from workload responsibilities;
- choose Control Tower, Control Tower plus extensions, Landing Zone Accelerator, or a custom implementation;
- design identity, OU/accounts, controls, network/DNS/IPAM, logging, security, backup, deployment, FinOps, and operations as connected domains;
- identify central-service blast radius and design degraded operation/recovery;
- build a RACI, service catalog, account-vending contract, exception process, and evidence model;
- plan greenfield, brownfield, merger/acquisition, multi-Region, and regulated adoption;
- gate workload migration on tested identity, connectivity, observability, backup/restore, quotas, cost, and rollback;
- version, test, deploy, monitor, upgrade, and retire the platform as a product.
What a landing zone is - and is not
AWS defines a landing zone as an orchestration framework for the foundational AWS environment. It establishes an initial multi-account baseline for identity/access, governance, data security, network, and logging.
It is not:
- one management account full of shared workloads;
- only an Organizations OU tree;
- only a Control Tower installation;
- a one-time CloudFormation deployment with no owner;
- proof that workloads are secure, resilient, or compliant;
- a replacement for application threat modeling, IAM least privilege, backups, or incident response.
Use the “paved road” model: the platform supplies safe defaults, reusable services, evidence, and self-service automation. Workload teams remain accountable for data, code, architecture, access requests, service configuration, recovery objectives, and costs within that road.
Begin with requirements, not a reference diagram
Capture measurable requirements:
| Domain | Questions that change architecture |
|---|---|
| business/ownership | business units, products, acquisitions/divestitures, autonomy, 24x7 owners |
| risk/compliance | data classes, sovereignty, PCI/health/financial scope, legal hold, audit evidence |
| scale | accounts now/three years, Regions, resources, deployment rate, log volume, tenants |
| identity | authoritative IdP, MFA, workforce/workload split, JIT elevation, contractors, break glass |
| network | on-premises/cloud connectivity, inspection, egress, CIDR/IPAM, DNS, IPv6, partner access |
| resilience | workload RTO/RPO, alternate Region, central-service failure, ransomware/credential compromise |
| operations | platform team, support tiers, change windows, incident ownership, automation/IaC maturity |
| financial | payer, budgets, showback, commitments, Marketplace, support, log/security/network cost |
| migration | existing accounts/resources, downtime, domains/certificates/keys, rollback, coexistence |
An architecture is accepted only when each requirement maps to an owner, implementation, test, evidence source, cost, and recovery path.
Select the implementation approach
AWS Control Tower
Use as the managed default for many new or existing Organizations environments. It orchestrates landing-zone versions, baselines, controls, governed Regions, enrollment, Account Factory, and integrations. It reduces undifferentiated implementation but imposes supported workflows and managed-resource boundaries.
Control Tower plus customer-owned extensions
Common enterprise choice. Keep the Control Tower lifecycle authoritative while separate IaC deploys networking, security services, budgets, backup, DNS, pipelines, and workload baselines. Extensions must not modify Control Tower-managed resources and need compatibility tests for every landing-zone update.
Landing Zone Accelerator on AWS (LZA)
LZA is an AWS Solution that uses configuration files, CodePipeline/CodeBuild, CDK, and CloudFormation to orchestrate a prescriptive multi-account environment. Its configuration repository manages accounts, organization, global settings, IAM, network, security, and optional customizations. It can use Control Tower Account Factory where integrated. It is valuable for regulated or complex repeatable architectures, but the customer owns configuration review, pipeline privileges, solution upgrades, failure recovery, and conflicts with other automation.
Custom landing zone
Justified when supported managed patterns cannot meet hard requirements. The customer owns Organizations/account lifecycle, every policy/integration, deployment ordering, drift, release compatibility, documentation, support, rollback, and decommissioning. “We already have scripts” is not sufficient. Avoid multiple controllers managing the same resource.
Write an ADR comparing fit, gaps, lock-in, managed/custom ownership, release cadence, skill, cost, migration, recovery, and exit strategy.
Organization and account architecture
Build from AWS254–255:
management (control plane only)
Security: log archive, security tooling/config integration
Infrastructure: network, shared services, operations, backup
Deployments: build/artifact/release orchestration
Workloads: production and non-production accounts
Sandbox: isolated/budgeted/expiring experimentation
PolicyStaging: representative governance tests
Transitional/Exceptions/Suspended: process-bound, owned, expiring
An OU groups accounts needing common controls/lifecycle; an account is the stronger resource/identity/quota/billing boundary. Give every account a catalog record, durable email/contacts, owner, data class, environment, regions, network, budget, baseline version, exceptions, and decommission date.
Keep log storage separate from security analytics/remediation and keep both outside the management account. Security responders may query logs without gaining deletion rights. Network and deployment accounts are highly privileged shared services, not ordinary workload accounts.
Identity architecture
The workforce path should be:
authoritative HR/IdP -> MFA/conditional access -> SAML + SCIM
-> IAM Identity Center groups -> permission sets/account assignments
-> temporary role session -> CloudTrail/source identity
Define standard job-function permissions, short production elevation, management-account exceptions, ABAC attribute governance, joiner/mover/leaver SLAs, session containment, access reviews, and independent break glass. Do not create per-account IAM users or pipeline long-term keys.
Workload identity uses service roles, instance/task roles, IRSA/Pod Identity where appropriate, OIDC federation for CI, Roles Anywhere for approved external workloads, and Secrets Manager/Parameter Store for actual secrets. Human and machine trust paths must be separate.
Central identity failure planning includes cached/active session behavior, emergency administrator credentials, IdP/DNS/network outage, certificate/SCIM rotation, and periodic recovery tests.
Governance and control architecture
Translate invariants into the right layer:
- SCP: maximum permissions for member-account principals;
- RCP/resource policy: maximum/resource-side access including external principals where supported;
- declarative policies: central supported service configuration;
- Control Tower preventive/detective/proactive controls;
- IAM boundaries/session policies and workload IAM;
- Config/Security Hub/custom detection and automated or approved remediation;
- network/service perimeter controls;
- pipeline policy-as-code checks before deployment.
Maintain a control catalog: requirement, control ID/policy, owner, target OUs/accounts/Regions/resources, behavior, evidence, remediation, exception, test, cost, and review date. A single objective may need prevent + detect + respond. Avoid duplicated Config rules and conflicting automation.
PolicyStaging mirrors representative OU paths. Canary changes, monitor denials/compliance, batch rollouts, and preserve management-account rollback. Exceptions require risk owner, compensating control, expiry, and continuous evidence.
Network, DNS, and IPAM architecture
Choose centralized, distributed, or hybrid connectivity deliberately:
- AWS Transit Gateway or Cloud WAN for scalable routing domains;
- Direct Connect and Site-to-Site VPN with redundant paths and BGP design;
- centralized inspection/egress only when shared blast radius and cost are acceptable;
- VPC sharing where central subnet ownership and participant responsibilities fit;
- PrivateLink/service endpoints for narrow service exposure;
- AWS Network Firewall/third-party appliances with symmetric-routing and fail-open/closed decisions;
- VPC endpoints to reduce public paths, with endpoint/resource policy and DNS design.
Use VPC IPAM for hierarchical pools, non-overlap, Region/environment allocation, IPv4 conservation, and IPv6 strategy. Reserve growth and M&A ranges. Document route propagation/limits, MTU, appliance mode, NAT/endpoint/TGW/data-transfer cost, and quota ownership.
DNS needs public/private zone ownership, Route 53 Resolver inbound/outbound endpoints/rules, split-horizon behavior, DNSSEC where needed, domain registrar/root credentials, forwarding loops, query logs, and failure recovery. Organization hierarchy does not create network/DNS connectivity.
Test east-west, north-south, on-premises, internet egress, service endpoints, partner links, denied paths, DNS failure, inspection failure, and alternate connectivity.
Logging, evidence, and observability
Design a log source matrix for CloudTrail organization management/data/network events, Config, VPC/TGW flow logs, Route 53 DNS, load balancers, WAF/Network Firewall, operating systems, containers, databases, applications, identity provider, CI/CD, security services, and platform operations.
For every source define account/Region coverage, destination, source-policy conditions, KMS ownership, format/partition, latency, integrity, retention/legal hold, Object Lock/versioning, replication, query/SIEM route, access, failure alarm, cost, and restore/export.
The Log Archive account should be a durable, tightly controlled sink. Security Tooling runs analytics/remediation. Workload teams may need local operational logs but cannot erase central evidence. Preventive controls protect destinations; delivery health tests prove data arrives. “Organization trail exists” is not enough - test a canary event from every account/Region through query and retention.
Central observability includes metrics, traces, logs, SLOs, cross-account observability, incident routing, and tenant/access separation. Do not send sensitive application payloads into a broadly accessible central log system without classification and redaction.
Security operations and data protection
Delegate organization security services to Security Tooling as supported: Security Hub CSPM, GuardDuty, Inspector, Macie, Detective, Security Lake, Firewall Manager, Access Analyzer, or third-party SIEM/SOAR. Record enrollment status, Regions, new-account auto-enable, standards, finding aggregation, suppression ownership, response SLA, and cost.
Encryption design distinguishes AWS-owned, AWS-managed, and customer-managed keys; key account/Region, administrators/users, grants, rotation, deletion protection, cross-account policy, replicas, quotas, cost, and recovery. A centralized key can enforce separation but becomes a shared availability dependency.
Data perimeter design combines trusted identities, resources, networks, and encryption context using SCPs/RCPs/resource/endpoint policies and supported global keys. Monitor missing context and service-to-service exceptions before enforcement.
Backup, resilience, and disaster recovery
Landing-zone backup includes more than workloads:
- cross-account/Region backup vaults and logically air-gapped or immutable controls where required;
- customer-managed key and grant recovery;
- configuration repositories, IaC state, pipelines, policies, account catalog, IPAM/DNS data, contacts, and runbooks;
- central log durability/replication;
- management account, IdP, network, DNS, artifact, security tooling, and deployment outage procedures.
Define workload RTO/RPO and choose backup/restore, pilot light, warm standby, or multi-site per workload. Landing-zone services need their own SLO/RTO/RPO. Test restore into an isolated recovery account, identity/network/DNS cutover, keys, quotas, data consistency, security controls, and failback. A successful backup job is not recovery proof.
Deployment and software supply chain
Separate untrusted build from production deployment authority. Use source review, protected branches, ephemeral runners, OIDC/temporary roles, dependency/secret/malware scans, SBOM, artifact signing/provenance, immutable artifact repositories, environment approvals, narrow target roles, CloudTrail source identity, and rollback.
The platform configuration repository - Control Tower settings, LZA YAML, Terraform/CDK/CloudFormation, SCPs, network, identity, and security automation - is production code. Require schema/lint/policy tests, synthesized change review, representative organization test, canary, concurrency controls, change window, drift detection, and known-good rollback.
No console-only change should become an undocumented permanent baseline.
Account vending as a product
The service catalog request collects owner, workload/environment, data class, cost center, contacts, OU, Regions, network/CIDR/DNS, identity groups, quotas, backup, monitoring, exceptions, and expiry. The workflow:
validate request -> create/invite account -> catalog/tag/place OU
-> deploy baseline -> identity/network/DNS/security/log/backup/budget/quota
-> positive and negative acceptance -> handoff with SLO/support
Use asynchronous polling, idempotency, retry, dead-letter/manual recovery, and partial-failure cleanup. Publish lead time and status. A new account is not ready merely because Organizations returns SUCCEEDED.
FinOps and capacity model
Define consolidated payer/billing transfer, cost allocation tags/categories, account ownership, budgets/anomaly detection, CUR/data export, showback/chargeback, RI/Savings Plans ownership/sharing, Marketplace/private offers, support, credits, and unit economics.
Model platform cost by account × Region × resource/change volume. Config evaluations, CloudTrail data events/Lake, Security Hub/GuardDuty/Inspector/Macie, log ingestion/storage/query, NAT/endpoints/TGW/Cloud WAN/data transfer, backup copies, KMS, SIEM, and support can dominate. A control with no cost owner will be disabled under pressure.
Track service quotas for workload and central hubs. Vending requests increases before migration. Monitor IP space, TGW routes/attachments, endpoints, DNS queries, log bucket objects/throughput, KMS/API rates, Config recorders/evaluations, security member limits, account creation, StackSet concurrency, and pipeline throughput.
Operating model and RACI
Name accountable owners for platform product, management/root, identity/IdP, Organizations/controls, network/DNS/IPAM, security services/incident response, log archive/evidence, keys, backup/DR, CI/CD, account vending/catalog, FinOps, compliance, workload teams, and vendor support.
Define SLOs and escalation for account delivery, identity assignment, network change, DNS, quota, control exception, incident containment, restore, and decommission. Use dual control for organization-wide destructive changes. Run game days for IdP outage, bad SCP, TGW/DNS/egress failure, log delivery break, KMS denial, compromised deployment role, Region loss, and account quarantine.
Greenfield, brownfield, and M&A adoption
Greenfield still needs staged acceptance before production. Brownfield starts with discovery: accounts, ownership, identities, root contacts, policies, Regions, network/CIDRs/DNS, trails/Config/security services, data/keys, pipelines, quotas, costs, drift, and unsupported resources.
Do not enroll/move everything at once. Use waves:
- management/foundational design and PolicyStaging;
- log/security/network/shared services with tested recovery;
- low-risk sandboxes/non-production;
- representative production canary;
- remaining workloads by dependency group;
- transitional cleanup and old-control retirement.
Acquired accounts enter Transitional with restricted trust/connectivity, inventory, root/contacts/payer confirmation, threat review, log onboarding, CIDR/DNS collision analysis, and an exit plan. Divestiture requires data/log/key/commitment/domain separation and removal of organization trust.
Workload migration readiness gates
A production workload moves only after:
- account baseline/control and owner/contact acceptance;
- workforce/workload identity, least privilege, MFA/elevation, break glass tested;
- required quotas/capacity and Regions approved;
- network/DNS/endpoint/egress and denied paths tested;
- logs/metrics/traces/security findings arrive and alert end to end;
- backup copy plus isolated restore meets RTO/RPO;
- keys/secrets/certificates and rotation/recovery work;
- deployment/provenance/rollback and incident containment work;
- budget/cost allocation/support/on-call are live;
- data migration consistency, cutover and rollback criteria are signed.
No “landing zone ready” dashboard substitutes for these workload-level gates.
Practical lab
Download the AWS258 landing-zone architecture pack. It contains:
LANDING_ZONE_ADR_WORKBOOK.mdfor a 50-account enterprise;SUPPLIED_INTEGRATION_CASES.mdwith eight architecture failures;DOMAIN_CONTRACT_AND_RACI.md;MIGRATION_READINESS_AND_GAMEDAY.md;ANSWER_DIRECTIONS.md, opened after independent design.
Submit architecture diagrams for organization, identity, network, logging/security, deployment, and recovery; a dependency graph; tool-selection ADR; RACI/SLO; cost model; account product; control/evidence catalog; migration waves; and game-day results.
Read-only evidence audit
No single command proves a landing zone. With approved read access, combine:
aws organizations describe-organization
aws organizations list-accounts --query 'Accounts[].{Id:Id,Name:Name,State:State}'
aws sso-admin list-instances --region ap-south-1
aws cloudtrail describe-trails --include-shadow-trails false --region ap-south-1
aws configservice describe-configuration-aggregators --region ap-south-1
aws ec2 describe-transit-gateways --region ap-south-1
aws route53resolver list-resolver-rules --region ap-south-1
aws backup list-backup-vaults --region ap-south-1
Then inspect every account/Region/service, paginate, and correlate behavior. Redact IDs, names, emails, CIDRs, endpoints, domains, key/bucket ARNs, and architecture-sensitive data.
Troubleshooting matrix
| Symptom | Evidence | Architecture gap |
|---|---|---|
| account vended but unusable | workflow stages, baseline, identity/network/quota | no acceptance contract |
| security has logs but workloads cannot troubleshoot | access/query tiers and data classification | archive/operations needs conflated |
| one DNS/egress/key outage stops organization | dependency graph/game day | central blast radius lacks recovery |
| policy says compliant, resource remains risky | behavior/evidence/remediation | detect confused with prevent/respond |
| duplicate rules/trails and rising bill | control/source-of-truth inventory | multiple controllers overlap |
| migration fails on certificate/key/domain | service-specific inventory | brownfield discovery incomplete |
| DR backup exists but recovery misses RTO | isolated restore timeline | job success confused with recovery |
| new account uses public/overlapping CIDR | IPAM/vending validation | network contract absent |
| production deploy role compromised from PR build | trust/provenance/pipeline stages | build and release authority mixed |
| acquired account exposes transitive network/trust | Transitional controls and graph | M&A quarantine missing |
Cost, change, and lifecycle
Landing-zone tools may have no direct fee, but the architecture has continuing service and people cost. Price all underlying services, environments, Regions, evidence retention, network paths, security coverage, backup, and pipeline operations. Measure cost per account and per workload onboarding.
Treat every platform change as versioned production release. Record architecture decisions and deprecations. Decommission only after all consumers, data/evidence retention, keys/domains, commitments, identity, connectivity, billing, and recovery have owners. Control Tower/LZA/custom resources have different supported removal workflows; manual deletion is not decommissioning.
Knowledge check
- Why is a landing zone more than Control Tower or an OU tree?
- When does LZA add value, and what ownership does it add?
- What evidence can justify a custom implementation?
- Why should Log Archive and Security Tooling be separate?
- Which identity responsibilities remain with workload teams?
- How does a control catalog prevent overlap and gaps?
- Why can centralized egress/DNS/KMS become correlated failure?
- What must accompany a cross-account log or backup path?
- Why must build and production deployment authority be separated?
- What makes account vending idempotent and acceptable?
- Which costs scale by account, Region, resource, or configuration change?
- What must a landing-zone RACI include?
- Why do brownfield and M&A accounts enter in waves/quarantine?
- Which workload gates must pass before production migration?
- Why does backup success not prove disaster recovery?
- What makes a platform change safely reversible?
Lesson acceptance
Pass requires a requirement-linked tool-selection ADR; six architecture views and dependency graph; all domain contracts; account/OU and identity design; network/DNS/IPAM and failure tests; immutable evidence architecture; security/control catalog; key/data perimeter; backup/restore; software supply chain; account-vending acceptance; FinOps/capacity; RACI/SLO; eight supplied diagnoses; phased migration with workload gates/rollback; game days; and no-create proof.