AWS 203: SAA-C03 practical architecture review
Why this lesson matters
An architecture diagram is a hypothesis. An architect must connect each requirement to a control, each control to observable evidence, each failure to a recovery owner and each resource to a cost and cleanup owner. This independent review is the SAA phase exit gate: the learner defends the real P09 template and AWS200–201 evidence, including weaknesses, rather than presenting a collection of AWS service icons.
What you will be able to do
By the end, you can:
- trace the P09 client, load-balancing, compute, identity and S3 data paths;
- map requirements to CloudFormation resources and runtime proof;
- distinguish what the template proves from what requires deployment evidence;
- defend security, resilience, performance and cost tradeoffs;
- respond to failure, scaling, encryption, TLS and cleanup challenges;
- record a finding with severity, owner, due date and measurable closure test;
- make a justified pass, conditional pass or not yet ready decision.
Assessment independence and evidence honesty
The reviewer should not be the learner who built P09. If no independent person is available, run a recorded self-review twice: first as presenter, then after a break as a hostile reviewer using the challenge bank below. Never claim a live deployment, failover, scaling event, CloudFormation validation or zero-residual cleanup unless the submitted evidence proves it.
Valid tracks:
- Live track: reviewed template plus deployment, behavior, fault and cleanup evidence from an approved learner-owned account.
- Evidence track: reviewed template plus instructor-provided runtime evidence. Mark every runtime claim as supplied evidence.
- Design-only track: local static validation and defended predictions. This can earn feedback but cannot pass runtime or cleanup criteria.
Required portfolio
| Artifact | Minimum content |
|---|---|
| Requirement traceability matrix | requirement ID, priority, control/resource, runtime proof, owner and gap |
| Current architecture diagram | account/Region/AZ boundaries, public/private addressing, ports, SG direction, identity role, S3 endpoint and request arrows |
| P09 template | exact reviewed template.yaml hash and validation output |
| Healthy baseline | stack outputs, two healthy targets in different AZs, repeated ALB responses and object value |
| Four recovery timelines | compute, network, versioned-data and identity fault; before/failure/repair/retest timestamps |
| Security proof | IMDSv2, encrypted gp3, bucket public block/versioning/encryption, SG-to-SG ingress and scoped role policy |
| Cost record | Region, date, quantities, hours, request/data assumptions, estimate links and uncertainty |
| Cleanup ledger | exact owned inventory before/after, bucket versions/delete markers, stack deletion and delayed billing review |
| Decision record | accepted limitations, production extensions, rejected alternatives and trigger to revisit |
Use the P09 template, deployment and cleanup runbook, and controlled fault cards. Record SHA-256 hashes so the reviewer knows the evidence and template refer to the same revision.
P09 architecture trace
Internet client
-> ALB public DNS :80 (two public subnets, two AZs, ALB SG)
-> target group health/routing :8080
-> app SG permits only source ALB SG
-> ASG desired 2, min 2, max 4; launch template; AL2023; no public IPv4
-> instance role obtains temporary credentials through IMDSv2
-> S3 gateway endpoint and endpoint policy
-> private versioned bucket app/message.txt
CloudWatch <- HealthyHostCount alarm and ASG target-tracking metric
CloudFormation <- ownership, outputs, update behavior and deletion workflow
The baseline deliberately uses HTTP so the learner can prove reachability without owning a domain. That is not acceptable for confidential production traffic. A production proposal must add a DNS name, ACM certificate, HTTPS listener, HTTP redirect or removal, an explicit target-side encryption decision, access logging/retention, WAF decision, and operational observability.
Read-only verification commands
Use a known caller and Region. These commands inspect P09; they do not create resources. Run them from the directory containing the downloaded P09 template.yaml.
set -euo pipefail
export AWS_DEFAULT_REGION="ap-south-1"
stack_name="nw-p09-capstone"
aws sts get-caller-identity --query Arn --output text
aws cloudformation validate-template \
--template-body file://template.yaml
aws cloudformation describe-stacks --stack-name "$stack_name" \
--query 'Stacks[0].{Status:StackStatus,Parameters:Parameters,Outputs:Outputs}' --output json
aws cloudformation list-stack-resources --stack-name "$stack_name" --output table
aws cloudformation describe-stack-events --stack-name "$stack_name" --max-items 30 --output table
aws resourcegroupstaggingapi get-resources \
--tag-filters Key=Project,Values=NitWings-P09 --output json
validate-template checks template syntax and limited structure against the service; it does not prove successful creation, IAM least privilege, runtime behavior or cleanup. Starting drift detection changes control-plane state, so use existing drift evidence unless the account owner separately approves a new detection. The tagging API does not support every resource type and is not a complete cleanup oracle.
Twenty-minute architecture defense
Use no more than 12 slides and leave at least 15 minutes for challenges.
- 0–2 minutes - requirements: users, traffic, availability target, data sensitivity, RPO/RTO, budget and exclusions.
- 2–5 minutes - request trace: DNS/ALB listener, SGs, target group, app process, IMDSv2 credentials, endpoint route/policy, bucket policy/data.
- 5–8 minutes - security: least privilege, temporary credentials, encryption, public exposure, administrative boundaries and remaining TLS/logging gaps.
- 8–11 minutes - resilience: two AZs, desired capacity, target health, replacement ownership, versioned-data recovery and what remains single-Region.
- 11–14 minutes - performance/scaling: target tracking, health grace, ALB behavior, S3 path, expected bottlenecks and load-test limits.
- 14–17 minutes - cost: ALB hours/LCUs, EC2 hours, EBS, public IPv4, requests, logs and data transfer; explain why there is no NAT gateway.
- 17–20 minutes - operations: alarms, four recovery timelines, cleanup ledger, known risks and next production changes.
Reviewer challenge bank
The reviewer chooses at least eight, including two from each SAA domain.
Secure
- Why does “no public IPv4” not by itself prove the workload is private?
- Which policies can still deny the S3 read after the role policy allows it?
- What does IMDSv2 mitigate, and what does it not mitigate?
- Where does encryption in transit start and stop in this baseline?
- Why is an ALB-SG source safer than allowing the VPC CIDR on port 8080?
Resilient
- What happens to user requests while one instance is terminated and replaced?
- Does two-AZ compute protect against a Regional outage? Prove the boundary.
- Why is S3 versioning not the same as a backup or immutable retention?
- Which signal declares a target unhealthy, and who initiates replacement?
- State measurable RPO and RTO for the object-delete exercise from timestamps.
High-performing
- What happens if CPU is low but request latency or queue depth is high?
- Why can six ALB requests fail to show both instances even when both are healthy?
- Which parameters govern scale-out speed and new-target readiness?
- Where could an S3 call on every HTTP request become a bottleneck?
- When would CloudFront, caching or asynchronous processing be justified?
Cost-optimized
- Name every continuously billed component in P09 and its quantity driver.
- Which traffic path avoids NAT gateway processing cost?
- What is the cost consequence of increasing minimum capacity from two to four?
- When could Savings Plans be justified, and why not from one short lab?
- How will you detect a versioned bucket or other residual that survives cleanup?
Scoring rubric
Score each category from 0 to 4. A claim without matching evidence cannot score above 2.
| Category | 0 | 2 | 4 |
|---|---|---|---|
| Requirements and traceability | absent | major requirements mapped | every priority maps to control, proof, owner and gap |
| Security | unsafe or unexplained | baseline controls named | identity/network/data boundaries proved; limitations explicit |
| Resilience and recovery | icons only | failure behavior described | four timestamped recovery paths and failure boundaries defended |
| Performance and scaling | no workload model | scaling components named | metric, threshold, cooldown/readiness, bottleneck and test connected |
| Cost | no estimate | main resources listed | quantities, duration, requests, transfer, uncertainty and owner shown |
| Operations and observability | no signals | alarm/log sources listed | transaction-to-signal timeline and actionable ownership shown |
| IaC and change safety | no artifact | template shown | hash, validation limits, dependency/update/rollback/drift reasoning shown |
| Communication and tradeoffs | assertions | some alternatives | concise decision, rejected options, risks and revisit triggers defended |
| Evidence integrity and cleanup | unverifiable | partial screenshots | reproducible commands, timestamps, redaction and zero-residual ledger |
Maximum: 36 points.
- Pass: at least 30/36, no category below 3, live or supplied runtime evidence, and cleanup criterion passed.
- Conditional pass: 26–29 with no security or cleanup score below 3; findings must close and be retested within seven days.
- Not yet ready: below 26, any fabricated claim, unsafe live action, unrecovered fault, unowned cost, or residual paid resource.
The assessment gate is stricter than merely answering certification questions because the course goal includes practical architecture behavior.
Finding and closure format
Every review challenge that exposes a gap becomes a finding:
Finding ID:
Severity: critical | high | medium | low
Requirement at risk:
Evidence observed:
Root cause or uncertainty:
Owner and due date:
Smallest safe remediation:
Changed retest and measurable pass condition:
Closure evidence and reviewer sign-off:
A statement such as “enable monitoring” is not measurable. “Create an alarm when healthy targets remain below two for two one-minute periods, route it to the named owner, inject one approved target failure and preserve notification/recovery timestamps” is measurable.
Diagnose disagreements
- If diagram and template disagree, the deployed/template revision and hash decide the implemented state; update the diagram or raise a finding.
- If Console and CLI disagree, verify account, Region, role, filters, eventual consistency and timestamp before choosing either view.
- If CloudFormation says complete but the request fails, inspect listener, target health reason, SG source, process and S3/IAM response in request order.
- If recovery succeeded but no one can reproduce it, the evidence is insufficient; repeat with exact commands and UTC timestamps.
- If the estimate and bill differ, reconcile runtime hours, LCU dimensions, public IPv4, requests, log ingestion/retention, data transfer, tax and pricing date.
- If cleanup tagging shows zero, still inspect unsupported/untagged dependencies and the versioned bucket. Never delete by name similarity.
Lesson acceptance
- All required portfolio artifacts exist, agree by revision/hash and contain no secrets or unredacted account IDs.
- The learner completes the 20-minute defense and answers at least eight balanced challenges.
- An independent reviewer scores all nine categories and records rationale.
- Every conditional finding has owner, date, smallest safe repair and changed retest.
- A pass includes runtime/recovery evidence and reproducible zero-residual cleanup; design-only work is labelled and cannot masquerade as deployment proof.
- The final decision is pass, conditional pass or not yet ready; certification scheduling is recommended only after AWS202 and AWS203 both pass.