AWS 343: Organizational complexity and new solutions
Why this checkpoint matters
Large organizations rarely fail because nobody knows an AWS service name. They fail because authority, identity, network, data, security, finance, and delivery decisions cross team and account boundaries. A new solution can be technically elegant yet impossible to govern or operate.
This checkpoint tests whether you can combine AWS324 and AWS325 rather than repeat them. You receive an incomplete enterprise brief, discover contradictions, design a new solution, prove policy and failure behavior, and defend trade-offs. It is open-book, but unsupported assertions receive no credit.
Outcomes
You will demonstrate that you can:
- turn ambiguous stakeholder statements into measurable requirements;
- design account and OU boundaries without mistaking them for network topology;
- reason correctly about IAM, resource policies, SCPs, permissions boundaries, and session policies;
- select regional and global services from latency, residency, availability, and operating constraints;
- expose control-plane, data-plane, shared-service, and organizational failure domains;
- define ownership, evidence, cost allocation, lifecycle, and acceptance before deployment;
- change a design when evidence invalidates an assumption.
Rules and safety
This is a T0 document exercise. Do not create resources or modify an organization. You may use official documentation, earlier artifacts, local diagram tools, and sanitized read-only evidence. Cite every service limit or behavior that materially affects a decision. Label facts, assumptions, constraints, and unknowns separately.
Submit your own reasoning. A diagram without packet, identity, data, failure, and ownership explanations is not evidence.
Case: Northstar acquisition platform
Northstar has 42 existing AWS accounts in one organization and is acquiring Meridian, which has 18 accounts in another organization. Teams operate in India, Germany, and the United States. The company will launch a document-processing platform for regulated customers.
The initial statements are deliberately incomplete:
- customers upload documents through a public API and receive an asynchronous result;
- Indian records must remain in India, and some German records must remain in the EU;
- the business asks for 99.95 percent monthly availability and “no lost accepted document”;
- workforce access uses separate identity providers during the acquisition;
- application teams need autonomy, but security requires immutable organization-wide evidence;
- both companies use overlapping RFC1918 ranges;
- the platform must onboard ten teams in six months;
- finance wants cost per processed document and separate acquisition costs;
- one executive requests one central network and one production account “for simplicity”;
- operations has no agreed after-hours ownership model.
Your first task is not architecture. It is clarification.
Part 1: requirement discovery
Write at least 18 questions. Cover data classification, legal authority, data residency versus processing location, accepted-request semantics, payload size, throughput, latency percentiles, burst behavior, retention, deletion, tenants, recovery objectives, identity, external integrations, audit retention, deployment frequency, support, and budget.
Convert the answers you choose into a requirement ledger:
| ID | Type | Testable requirement | Source | Confidence | Acceptance evidence |
|---|---|---|---|---|---|
| NFR-01 | Reliability | Define precisely when a document is accepted and maximum permitted loss | Product and risk | Medium until approved | Failure test plus durable-record query |
| SEC-01 | Residency | Indian classified document bytes and replicas remain in approved Indian Regions | Legal | High | Data-flow review, configuration evidence, negative test |
Do not translate “no lost document” directly into a product choice. Define acknowledgement boundary, duplicate behavior, recovery point, retention, and reconciliation first. Record contradictions and the stakeholder who must resolve each one.
Part 2: organization and governance design
Create an account model for platform production, non-production, security tooling, log archive, networking, shared services, sandbox, and acquisition quarantine. Explain why workload lifecycle and control ownership justify each boundary. An OU groups accounts for governance; it does not route packets or grant access.
Build a policy evaluation worksheet for six requests:
- A deployment role creates an approved service in a production account.
- The same role attempts a disallowed Region.
- An incident role reads central logs.
- An application administrator attempts to disable a trail.
- An acquired account requests access to a shared artifact bucket.
- A break-glass role performs a documented emergency action.
For each, evaluate authentication, identity policy, permissions boundary, session policy, SCP/RCP limits where applicable, resource policy, KMS key policy/grants, and service-specific conditions. SCPs restrict available permissions but do not grant them. Identify the explicit deny or missing allow that determines the result.
Design migration of Meridian accounts as states: discover, quarantine, identity review, evidence baseline, network/IP assessment, policy dry run, controlled OU movement, validation, and operational handover. Include exit and rollback criteria. Never assume an account can be moved safely because Organizations permits the API operation.
Part 3: new-solution architecture
Produce five coordinated views:
- Context: users, administrators, external systems, legal boundaries, and trust boundaries.
- Request/data: DNS and edge, authentication, upload, durable acceptance, asynchronous processing, result storage, notification, retention, and deletion.
- Deployment: accounts, Regions, Availability Zones, network paths, endpoints, and shared services.
- Identity/security: workforce and workload identities, authorization points, encryption ownership, evidence flow, and incident access.
- Failure/operations: detection, retry, reconciliation, recovery, owner, and customer-visible effect.
Choose services only after stating the required capability. Compare direct S3 upload with API-proxied upload; SQS with EventBridge or stream processing; Lambda with containers; DynamoDB with Aurora; and Route 53 with Global Accelerator or regional endpoints. Reject candidates based on payload, ordering, consistency, latency, protocol, residency, team skill, cost, or failure semantics.
Your design must explain:
- when a request becomes durable and what the client receives;
- idempotency key, duplicate window, poison-message handling, retry limits, and dead-letter recovery;
- tenant isolation and authorization at every data access;
- encryption in transit and at rest, KMS key ownership, rotation, deletion protection, and cross-account use;
- residency-aware routing and how a wrong-Region request is rejected;
- multi-AZ behavior and whether multi-Region recovery is required per data class;
- observability from customer request ID to processing result without exposing document content;
- deployment safety, schema compatibility, rollback, and operational readiness;
- unit-cost numerator and denominator plus shared-cost allocation.
Part 4: required calculations
Use explicit assumptions and show units. Calculate average and peak requests per second for 12 million documents per month with a 12-times peak factor. Estimate monthly ingested bytes at 8 MiB average, storage after replication, and transfer paths. Estimate concurrency using peak arrival rate multiplied by average processing time. Explain why averages do not size burst queues, downstream quotas, or recovery catch-up.
Construct expected, 50-percent growth, and incident-month cost scenarios. Include requests, compute duration/capacity, storage classes, retrieval, logs, KMS requests, NAT or endpoint processing, inter-AZ/inter-Region transfer, support, backup, and shared platform costs. The goal is a model whose uncertainty is visible, not a guessed total.
Part 5: failure challenges
Respond to each inject with detection, immediate decision, containment, recovery, reconciliation, owner, customer effect, and permanent improvement:
- The primary queue accepts work, but consumers fail for 45 minutes.
- A tenant retries the same 20,000 uploads after a timeout.
- The central inspection path loses one Availability Zone.
- A KMS key policy change blocks processing but not storage writes.
- An acquired account advertises an overlapping route.
- The identity provider is unavailable during a security incident.
- A deployment changes an event schema while old consumers still run.
- A legal reviewer finds that a backup crosses the residency boundary.
Update the architecture and requirements after the failures. Merely describing an alarm is insufficient.
Scoring rubric
| Dimension | Points | Full-credit evidence |
|---|---|---|
| Requirements and contradictions | 15 | Testable ledger, sources, confidence, approvals |
| Organization and authorization | 15 | Defensible account/OU model and six correct evaluations |
| Architecture and service selection | 20 | Five consistent views and rejected alternatives |
| Data, security, and compliance | 15 | End-to-end ownership, residency, encryption, deletion |
| Reliability and failure reasoning | 15 | Eight complete failure responses and corrections |
| Operations and delivery | 10 | Owners, telemetry, release, rollback, lifecycle |
| Capacity and cost | 10 | Unit-correct calculations and scenario sensitivity |
Pass requires 80/100 and at least half the points in every dimension. Automatic failure applies to fabricated evidence, undefined authoritative data, no accountable production owner, unrestricted root use, no audit path, or a residency design contradicted by its backup/replication path.
Required artifacts
Submit: brief clarification log, requirement ledger, conflict/assumption register, stakeholder map, account/OU model, six-request authorization worksheet, acquisition state model, five architecture views, service decision matrix, capacity worksheet, three-scenario cost model, eight failure records, risk/findings register, corrected architecture, and approval record.