AWS 325: Designing new enterprise solutions
Why this lesson matters
A greenfield workload has no application migration constraint, but it is not constraint-free. Enterprise identity, networking, data governance, delivery standards, budgets, skills, regulations, shared platforms, and support models already exist. A service diagram alone is not an architecture. A deployable solution must explain behavior, ownership, failure, evidence, rollout, and retirement.
This lesson converts approved requirements into conceptual, logical, deployment, security, data, operating, and financial views. It uses progressive detail so teams can reject weak options before expensive implementation.
Learning outcomes
By the end, you can:
- establish architecture principles and a system context;
- decompose capabilities without creating unnecessary microservices;
- produce logical and deployable AWS architecture views;
- design identity, network, data, integration, resilience, and observability together;
- calculate initial capacity, quotas, and cost drivers;
- plan infrastructure, application, and data delivery pipelines;
- validate important unknowns with experiments and threat/failure analysis;
- create staged rollout, rollback, and operational acceptance plans.
Architecture evidence map
| Architecture claim | Minimum supporting evidence |
|---|---|
| The design meets user needs | Accepted journey, measurable target, and representative test |
| The design survives failure | Failure-mode analysis, recovery procedure, and exercised result |
| Access is controlled | Trust flow, effective authorization, and allowed/denied test |
| Data is protected and recoverable | Classification, key/access model, backup, and restore evidence |
| Capacity is sufficient | Demand model, quota check, load result, and scaling behavior |
| Operations are sustainable | Named owner, SLO, telemetry, runbook, and support model |
| Cost is acceptable | Usage assumptions, transfer, environments, support, and growth model |
1. Begin with an architecture brief
Carry forward AWS322 requirements and AWS323's option decision. The brief should contain:
- business outcome and success metrics;
- user personas and critical journeys;
- functional scope and exclusions;
- demand, data, security, compliance, availability, RTO, RPO, and performance targets;
- enterprise platform and organization constraints;
- delivery date, budget, skills, and support model;
- assumptions, unknowns, decisions, and risks;
- acceptance evidence and decision owners.
Define principles that guide unresolved detail, such as automate repeatable changes, prefer managed capability when requirements fit, use least privilege and short-lived credentials, isolate failure domains, encrypt and classify data, make operations observable, and choose reversible paths under uncertainty. Principles are not substitutes for decisions.
2. Progress through architecture views
Context view
Show users, external systems, trust boundaries, geographic locations, and the workload as one box. Label protocols, identities, data classes, volume, ownership, and dependency SLOs.
Capability and domain view
Divide business capability by cohesive responsibility and data ownership. A modular monolith may be the right first design when the team and domain are small. Microservices add independent deployment and scaling but also network failure, distributed data, versioning, observability, and organizational overhead.
Logical component view
Show entry, authentication, application behavior, asynchronous work, data stores, caches, search, analytics, notifications, audit, and administration without forcing every component to an AWS product too early.
Deployment view
Map components to accounts, Regions, Availability Zones, VPCs, subnets, services, endpoints, keys, roles, scaling units, and failure domains. Include management planes such as delivery, observability, backup, security, and operations.
Dynamic and failure views
Sequence critical requests and events. Then redraw during AZ loss, dependency timeout, throttling, duplicate event, stale cache, identity-provider outage, deployment failure, and Regional disaster. Architecture behavior during failure matters more than the happy-path icon layout.
3. Identity and authorization
Separate workforce and customer identity from workload identity. Define:
- identity provider and federation;
- authentication strength and session lifecycle;
- authorization model and tenant boundary;
- service roles, resource policies, and cross-account trust;
- secret and certificate ownership/rotation;
- KMS key policy, grants, recovery, and separation of duties;
- privileged operations, break-glass, and audit evidence.
Follow a request from user token to edge, application, downstream service, database, encryption key, logs, and administrative action. A network path is not authorization. A role that can call an API may still be denied by an SCP, resource policy, permissions boundary, session policy, or KMS key policy.
4. Network and edge design
Document DNS, certificates, public/private entry, DDoS and web controls, load balancing/API ingress, VPC/subnet boundaries, routes, security groups, endpoints, egress, inspection, hybrid paths, IPv4/IPv6, and administration.
Use multiple AZs when the requirement demands resilience, but verify each tier actually distributes capacity and state. Decide whether egress is distributed or centralized based on availability, inspection, routing, and cost. Private subnets do not automatically make workloads secure; authorization, outbound controls, patch paths, and telemetry still matter.
Calculate address demand and growth. Identify source IP requirements, DNS behavior, MTU, data-transfer charges, cross-zone/Region paths, and quota dependencies.
5. Compute and integration shape
Choose compute by execution behavior, not fashion. Compare EC2, containers, and serverless for runtime, control, scaling, startup, duration, networking, patching, deployment, observability, quotas, and cost.
Use synchronous calls where the caller needs an immediate result and the dependency can meet the latency/availability budget. Use queues or events to absorb bursts, decouple lifecycle, or run deferred work. Define timeout, bounded retry, exponential backoff, jitter, idempotency, duplicate handling, ordering, poison-message handling, and backpressure.
Avoid long synchronous chains. Their availability and latency compound. Define degradation: can checkout accept an order when recommendations fail? Can a queued notification wait? Which dependency failure must stop the transaction?
6. Data architecture
Start with access patterns, transactions, consistency, scale, retention, query, and recovery rather than product names. For each store define:
- system of record and data owner;
- keys, access patterns, indexes, and expected growth;
- consistency and transaction boundary;
- classification, residency, encryption, and authorization;
- backup, point-in-time recovery, RTO, RPO, and restore testing;
- lifecycle, archive, legal hold, and deletion;
- replication, analytics, search, and cache copies;
- schema/version evolution and data-quality evidence.
Do not let every component write every database. Publish owned interfaces or events. For distributed changes, evaluate saga and transactional-outbox patterns, compensation, and reconciliation. Exactly-once business outcomes usually come from idempotency and state control, not a transport slogan.
7. Reliability and disaster recovery
Turn service objectives into budgets. Identify component and dependency contribution to user-journey availability and latency. Design health checks, timeouts, circuit breaking, bulkheads, scaling, quota headroom, deployment safety, and recovery.
Select backup/restore, pilot light, warm standby, or multi-site only after RTO/RPO and business impact justify it. Cross-Region infrastructure without recoverable data, keys, DNS/traffic controls, dependencies, runbooks, and practiced failover is not DR.
Define fault-injection and recovery tests, failover authority, fencing, data reconciliation, and failback. Protect backup evidence from the same identities and failures as production.
8. Security and threat model
Identify assets, actors, entry points, trust boundaries, abuse cases, and controls. Cover credential theft, tenant escape, injection, data exfiltration, dependency compromise, malicious deployment, denial of service, and operator error.
Layer prevention, detection, response, and recovery. Include CloudTrail and workload audit, configuration evidence, vulnerability and dependency scanning, image/artifact provenance, secrets scanning, WAF/Shield where justified, GuardDuty/Security Hub ownership, and incident procedures. Logging sensitive data is itself a risk.
9. Observability and operations
Define service-level indicators from user journeys: success, latency, correctness, freshness, and durability. Map them to logs, metrics, traces, synthetic checks, business events, alarms, dashboards, runbooks, and owner. Avoid alerting only on CPU.
Specify on-call hours, incident severity, escalation, deployment ownership, patching, capacity, certificate/key rotation, backup restoration, cost review, and dependency communication. Create runbooks for frequent known events and playbooks for investigative incidents.
10. Delivery and infrastructure as code
Separate application, infrastructure, data/schema, and policy changes while coordinating compatibility. Use version control, peer review, static checks, security tests, artifact integrity, environment promotion, approvals for material risk, and immutable evidence.
Prefer backward-compatible database changes: expand, deploy compatible code, migrate/backfill, verify, then contract later. Use canary, linear, blue/green, or rolling deployment based on state and rollback behavior. “Rollback” may require forward repair after irreversible data writes.
11. Capacity, quota, and cost model
Estimate requests, concurrency, event rate, connection count, storage growth, throughput, transfer, logs, backups, and recovery capacity. Include normal, peak, failure, and growth scenarios. Map each to service quotas and scale units.
Model fixed and variable service cost, data transfer, observability, backup, security tools, support, commitments, third-party licenses, environments, experiments, and people. Unit cost per successful order or active user supports later optimization.
12. Safe read-only evidence
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws ec2 describe-regions --query 'Regions[].{Name:RegionName,OptIn:OptInStatus}' --output table
aws service-quotas list-services --max-results 30 --output table
aws wellarchitected list-lenses --output table
aws pricing describe-services --region us-east-1 --max-results 30 --output table
These calls verify catalog or account state, not architecture suitability. Confirm features, quotas, prices, and Region support in current official sources and approved account evidence.
13. Guided workshop: new order platform
Design a new order platform for India and Singapore with mobile/web clients, 8,000 request/second bursts, payment and warehouse integrations, no duplicate charges, 99.95% availability, one-hour RTO, five-minute RPO, six-year order retention, analytics within 15 minutes, and a six-month first release.
Produce:
- architecture brief and principles;
- context and trust-boundary diagram;
- capability/domain decomposition;
- two logical options and selection record;
- account/Region/AZ deployment view;
- identity and authorization flow;
- network/edge/egress flow;
- three critical request/event sequences;
- data ownership and lifecycle model;
- idempotency, consistency, and reconciliation design;
- threat model and control evidence;
- availability budget and failure-mode analysis;
- backup, restore, and DR runbook outline;
- observability and incident model;
- infrastructure/application/data delivery design;
- capacity, quota, and three-scenario cost model;
- experiment and failure-test plan;
- staged release, rollback/repair, and operational acceptance plan.
14. Architecture review failures
Reject diagrams without requirements, single-AZ state under a multi-AZ label, synchronous dependency chains without budgets, shared administrator credentials, hidden public/egress paths, databases selected without access patterns, events without idempotency, backups without restore tests, multi-Region claims without data/failover, and designs without owners or cost.
Cost and cleanup
This workshop creates no resources. Any later POC needs tags, cost cap, expiration, manifest, and verified cleanup. Retain sanitized diagrams and decision evidence.
Knowledge check
- Why is greenfield not constraint-free? Enterprise identity, governance, skills, data, and operations already constrain it.
- When are microservices justified? When domain, team, deployment, scaling, and ownership benefits exceed distributed-system cost.
- What makes DR complete? Recoverable data/keys/dependencies plus traffic control, authority, tests, and failback.
- Why model failure views? Happy-path diagrams hide the behavior that determines resilience.
- Why separate workload and workforce identity? Their trust, lifetime, automation, and audit needs differ.
Lesson acceptance
Submit all 18 artifacts. The design must trace to accepted requirements, expose trust and failure boundaries, assign ownership, quantify capacity/quota/cost, include security and recovery evidence, and provide executable rollout, validation, rollback or forward-repair, and operational acceptance.