AWS 322: Collecting business and technical requirements
Why this lesson matters
Architecture begins with a problem, not an AWS service. A request such as “build a fast, secure global platform” cannot be designed, tested, priced, or accepted. The architect must discover who needs what outcome, how the current system behaves, which constraints are mandatory, what failure means, and which evidence proves success.
Requirements are not a shopping list. They are an agreed, traceable description of outcomes and boundaries. Good requirements prevent teams from optimizing the wrong thing and make service selection in AWS323 defensible.
Learning outcomes
By the end, you can:
- conduct discovery with technical and nontechnical stakeholders;
- separate facts, assumptions, constraints, preferences, risks, decisions, and unknowns;
- express functional and measurable nonfunctional requirements;
- inventory users, data, integrations, dependencies, operations, and current evidence;
- expose conflicts and facilitate explicit trade-off decisions;
- define acceptance tests before choosing services;
- build a traceability matrix and control requirement changes;
- reject premature designs that lack evidence.
1. Start with outcomes, not solutions
“Use Lambda,” “move to Kubernetes,” and “deploy in three Regions” are proposed solutions. Ask what outcome each proposal serves. The real need may be irregular demand, deployment independence, or regional recovery. Several designs may satisfy it.
A useful outcome statement is:
For [stakeholder/user], improve [measurable result]
from [baseline] to [target] by [date],
while preserving [mandatory constraints].
Example: “For checkout customers in India and Singapore, increase completed orders from 96% to 99.5% by Q2 while keeping p95 confirmation latency below two seconds, meeting payment-data obligations, and remaining within the approved cost range.” Numbers require evidence and owner approval. Never invent a target to complete a template.
2. Stakeholders and decision rights
Include the sponsor, business owner, product manager, users, security, privacy, legal, finance, operations, support, developers, data owners, network teams, enterprise architecture, vendors, and audit or regulatory representatives as applicable.
Record each stakeholder's interest, impact, knowledge, authority, availability, and communication route. Define who decides scope, risk, cost, security exceptions, and acceptance. A RACI can show responsible, accountable, consulted, and informed roles, but it does not replace named acceptance owners.
Use multiple discovery methods:
- interviews for goals, pain points, exceptions, and sensitive concerns;
- workshops for shared processes, conflicts, and priorities;
- observation to compare documented and actual work;
- document review for contracts, diagrams, incidents, bills, and audits;
- telemetry for demand, latency, errors, utilization, and growth;
- prototypes when behavior cannot be predicted confidently.
Ask open questions first, then quantify: “What happens during enrollment?” followed by “How many enrollments per minute, at which percentile and season?” Repeat back your interpretation and ask the owner to confirm it.
3. Keep evidence types separate
| Type | Meaning | Example |
|---|---|---|
| Fact | Verified by named, dated evidence | Last quarter's peak was 2,100 requests/second |
| Assumption | Temporarily believed and awaiting test | Demand will grow 30% next year |
| Constraint | Mandatory boundary | Customer records must remain in India |
| Preference | Desirable but negotiable | The team prefers PostgreSQL |
| Risk | Uncertain event with impact | Vendor connectivity may miss the RTO |
| Decision | Approved choice with rationale | Accept active-passive recovery |
| Unknown | Missing information with owner | Maximum batch-file size is unknown |
Give every assumption and unknown an owner, validation method, and date. An assumption silently treated as fact becomes architecture debt.
4. Users, journeys, and functional requirements
Identify human users, administrators, devices, applications, partners, batch jobs, and auditors. For each persona capture identity source, location, accessibility, authorization, volume, frequency, and failure impact.
Write functional requirements as observable behavior:
Given [starting state]
when [actor or event performs action]
then [observable result]
and [recorded evidence]
Cover normal, alternate, failure, retry, duplicate, timeout, cancellation, and recovery paths. Describe inputs, validation, processing, outputs, notifications, and audit events. Avoid prescribing user-interface or AWS implementation detail unless it is a genuine constraint.
5. Measurable nonfunctional requirements
“Highly available,” “fast,” “secure,” and “scalable” are not testable. Define the workload boundary, measurement point, period, percentile, target, evidence, and owner.
Availability and resilience
State the measured journey, calculation period, dependency boundary, and error-budget owner. Define failure domains and degraded modes. Capture:
- RTO, the maximum acceptable restoration time;
- RPO, the maximum acceptable data loss measured in time;
- backup retention and restore-test frequency;
- Region, Availability Zone, dependency, and operator failures;
- continuity procedures when technology is unavailable.
RTO is not availability, and a backup is not proven until restoration is tested.
Performance and scale
Record current, normal, peak, burst, and forecast demand. Define requests per second, concurrency, event rate, data size, and growth. Use p95 or p99 latency when tail behavior matters and measure from the user's relevant boundary. Include queue delay, batch deadline, cold-start, throughput, and throttling expectations.
Security, privacy, and compliance
Classify data and record authentication, authorization, privileged access, separation of duties, encryption, key ownership, secrets, network boundaries, audit retention, incident response, vulnerability management, deletion, legal hold, residency, sovereignty, and applicable contracts. “Encrypt everything” is incomplete without data, key, identity, rotation, and recovery ownership.
Operations and support
Define deployment frequency, change windows, rollback time, observability, alert response, escalation, runbooks, patching, capacity reviews, service ownership, support hours, skills, accessibility, and localization.
Cost and commercial limits
Capture budget range, allocation model, forecast horizon, purchase commitments, licenses, support plan, transfer assumptions, and cost authority. Distinguish a hard ceiling from an optimization goal.
Migration and retirement
Define downtime, sequencing, coexistence, data validation, cutover, rollback, training, archive, contractual exit, decommissioning, and evidence that legacy data and billing have ended.
6. Data, integration, and dependency inventory
For each dataset record owner, source, schema, volume, growth, sensitivity, residency, retention, deletion, consistency, access patterns, backup, RTO, RPO, and consumers.
For each integration record producer, consumer, protocol, authentication, endpoint owner, payload, rate, timeout, retry, idempotency, ordering, delivery guarantee, maintenance window, quota, error contract, and test environment.
Draw the present request and data flow. Include DNS, identity provider, certificate authority, network path, third parties, queues, notification systems, observability, people, and manual approvals. A workload cannot be more available than an unexamined critical dependency.
7. Establish the current-state baseline
Collect diagrams, inventories, configuration, incidents, change history, telemetry, quotas, bills, support tickets, security findings, audit reports, contracts, and procedures. Mark source and observation date. A diagram is a hypothesis until verified against configuration and traffic evidence.
Read-only discovery may include:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws organizations describe-organization --output json
aws resourcegroupstaggingapi get-resources --resources-per-page 50 --output json
aws service-quotas list-services --max-results 20 --output table
aws health describe-events \
--filter eventStatusCodes=open upcoming \
--max-results 20 --output table
Organizations and Health may be unavailable to a learning role. Resource Groups Tagging API is not a complete inventory because service support varies. Record access boundaries and use approved Config, inventory, billing, and telemetry evidence. Do not seek administrator access merely to complete a worksheet.
8. Prioritize and resolve conflict
MoSCoW can label Must, Should, Could, and Won't for a defined release, but cannot demote laws, contracts, or safety controls. Add value, risk reduction, urgency, dependency, effort, and cost evidence.
Make conflicts visible: zero data loss versus asynchronous replication, very low latency versus distant residency, or a low fixed budget versus continuous multi-Region capacity. Present options and consequences to the authorized owner. Record the decision, dissent, assumptions, review trigger, and rejected alternatives. Engineers must not silently make a business risk decision.
9. Acceptance and traceability
Each requirement needs an ID, statement, rationale, source, priority, owner, acceptance method, and status. Link it through delivery:
| Requirement | Decision | Control/component | Test | Evidence | Owner |
|---|---|---|---|---|---|
| NFR-AV-01 | Multi-AZ design | Load balancer and targets | AZ impairment | Metrics and test report | Service owner |
| NFR-DR-02 | Point-in-time recovery | Backup policy | Restore isolated copy | Restore log and checksum | Data owner |
| SEC-04 | Least privilege | Roles and policies | Allowed and denied tests | Policy review and events | Security owner |
Acceptance specifies environment, data, load, failure, measurement tool, threshold, duration, and approver. “Dashboard looks healthy” is not evidence. Architecture decision records should contain context, decision, alternatives, consequences, evidence, owner, date, and revisit trigger.
10. Validation and change control
Play requirements back to stakeholders: goals, scope, journeys, data flow, behavior, NFRs, assumptions, conflicts, priorities, acceptance, and open decisions. Obtain explicit approval from accountable owners; meeting attendance is not sign-off.
Version requirements. For every change record requester, reason, and impact on scope, architecture, security, reliability, schedule, cost, tests, migration, and approvals. Update traceability instead of overwriting history.
Design can start when critical outcomes and constraints are accepted, major unknowns have owners and deadlines, dependencies are represented, acceptance is testable, and unresolved decisions are visible. Uncertainty need not be hidden.
11. Failure patterns
| Failure | Why it fails | Correction |
|---|---|---|
| Services chosen in meeting one | Discovery anchors around a preferred answer | Restate outcome and compare later |
| Only sponsor interviewed | Users, operators, data, and controls are missed | Build stakeholder map |
| Generic NFRs copied | Targets are untestable or unaffordable | Add boundary, number, period, evidence |
| Preference treated as constraint | Valid designs disappear | Record type and decision authority |
| Telemetry ignored | Capacity rests on anecdotes | Build dated baseline |
| Conflicts hidden | Engineers inherit business decisions | Present options and obtain decision |
| No test link | Completion becomes subjective | Maintain traceability |
| Scope changes silently | Cost, security, and acceptance become invalid | Version and assess changes |
12. Guided workshop: global customer platform
A retailer asks for a “global, always-on customer platform.” It serves India, Singapore, and the UK. Checkout uses a payment provider; identity uses a corporate provider; warehouse stock arrives through batches and events. Promotion peaks are unknown. Finance has a monthly figure but has not declared it a ceiling. Legal mentions residency without identifying data. Operations supports India business hours. The old system has incidents, undocumented manual steps, and seven years of orders.
Produce 16 artifacts:
- Outcome statement with baseline gaps, target, date, and constraints.
- Scope and out-of-scope list.
- Stakeholder map and decision-rights table.
- Interview guide with open and quantitative follow-ups.
- Fact, assumption, constraint, preference, risk, decision, and unknown register.
- Personas and key journeys including failures.
- Functional requirements with event-style acceptance.
- Demand and performance profile with percentiles and boundary.
- Availability, RTO, RPO, degraded mode, backup, and restore requirements.
- Security, privacy, compliance, residency, retention, and deletion matrix.
- Data catalog and lifecycle map.
- Integration and dependency inventory with failure contracts.
- Operations, migration, support, accessibility, and retirement requirements.
- Cost inputs and budget decision question.
- Conflict log and two decision records with options and consequences.
- Traceability matrix linking requirements, decisions, tests, evidence, and owners.
Do not select final AWS services. Candidates may be recorded for later comparison, but acceptance depends on problem clarity and evidence.
Knowledge check
- Why is “use three Regions” not a requirement?
Answer: It is a proposed solution; discover whether the outcome is latency, residency, availability, or recovery.
- What makes an NFR testable?
Answer: Scope, measurement point, workload, numerical target, period, evidence method, and owner.
- How do RTO and RPO differ?
Answer: RTO limits restoration time; RPO limits acceptable data loss measured in time.
- What happens when latency and residency conflict?
Answer: Present viable options and consequences to the authorized owner and record the decision.
- When may design start?
Answer: When critical outcomes and constraints are accepted, acceptance is measurable, dependencies are represented, and major unknowns have owners and dates.
Cost and cleanup
This T0 workshop creates no AWS resources. Keep artifacts sanitized. Discovery may expose sensitive organization, health, or inventory data, so share only approved excerpts. Record a no-create result.
Lesson acceptance
Submit all 16 artifacts. Every critical requirement must have an ID, source, owner, priority, and acceptance method. Assumptions and unknowns need validation dates. Data and dependencies must be inventoried, conflicts decided or visibly pending, and traceability connected to future architecture decisions and tests.