Lesson 329 · AWS Learning Path

AWS 329: Serverless design principles

· Published · 8 min read

Labelled process diagram for AWS 329: Event and durable payload to Managed integration and stateless compute to Purpose-built state and workflow to Outcome, retry, trace, scale, and cost evidence, with decision...

Why this lesson matters

Serverless means the provider manages servers, availability, and much scaling infrastructure; it does not mean no architecture or operations. Customers still own code, identity, data, quotas, event behavior, observability, resilience, and cost. Event-driven serverless systems can scale quickly, so a small design error can become a retry storm, downstream outage, or unexpected bill equally quickly.

Learning outcomes

By the end, you can:

  • decide when serverless fits and when another compute model is better;
  • design stateless Lambda functions with managed state;
  • distinguish direct, synchronous, asynchronous, queue, and stream invocation;
  • design concurrency, backpressure, idempotency, retries, and failure destinations;
  • choose API, event, queue, workflow, and data services by behavior;
  • secure event paths and minimize blast radius;
  • observe business outcomes and distributed failures;
  • model quota, performance, and cost for a document workflow.

1. Serverless decision boundary

Good candidates include event-driven work, APIs with variable demand, automation, stream/queue consumers, scheduled business tasks, and workflows that fit service limits. Evaluate runtime, duration, startup, package/image, CPU/memory, temporary storage, network, protocol, concurrency, state, compliance, and predictable utilization.

RequirementServerless directionConsider alternatives when
Variable event demandManaged scaling can reduce idle capacityDownstream cannot absorb scaling
Short bounded workLambda can fit wellExecution is long-running or specialized
Managed integrationEvents/queues/workflows reduce custom infrastructureRequired protocol/control is unsupported
Stateless processingExternal managed state enables scaleTight local state or host affinity dominates
Pay per useEfficient at intermittent demandSustained high utilization is cheaper elsewhere

Do not split every line of code into a function. Choose cohesive units that can be deployed, secured, observed, and retried safely.

2. Event-first design

For every trigger record producer, event meaning, schema/version, payload size/classification, frequency/burst, ordering, duplicate possibility, retention, retry owner, consumer, and business deadline.

Prefer events that describe facts and commands that request actions. Store large payloads in S3 or another controlled store and pass a claim-check reference with authorization, integrity, and expiration. Avoid chaining large payloads through multiple services.

Prevent recursive loops. An S3 write that triggers a function which writes into the same matching prefix can scale until concurrency or cost controls stop it. Separate prefixes/buckets or filter events and detect recursion.

3. Invocation behavior

Synchronous

The caller waits for a response and owns retry behavior. Lambda does not automatically retry function errors for direct synchronous invocation. The function may have completed partially before the caller sees timeout, so idempotency still matters.

Asynchronous

Lambda queues the event and can retry function errors. Configure maximum event age, retry attempts, and an on-failure destination or DLQ according to current invocation semantics. Monitor age and destination delivery. A successful invocation does not guarantee the downstream business action succeeded unless application evidence proves it.

SQS event source

Lambda polls and processes batches. Align visibility timeout with function timeout and retry duration. Use partial batch responses when appropriate so successful records are not retried with failed records. Configure redrive, maximum receives, DLQ alarm, quarantine, replay authorization, and idempotency.

Streams

Kinesis and DynamoDB Streams have ordered shards/partitions. A poison record can block progress. Understand batch, iterator age, parallelization, bisect-on-error, retry/age limits, and failure destinations. Ordering scope is not global unless the stream design provides it.

4. Stateless functions and managed state

Treat the execution environment as reusable but disposable. Initialize SDK clients and immutable configuration outside the handler when safe, but never depend on in-memory state for correctness. Do not leak one tenant's mutable data into another invocation.

Use DynamoDB, S3, RDS/Aurora, ElastiCache, queues, or workflows according to access patterns. For relational connections, control concurrency and consider RDS Proxy where it fits; unbounded functions can exhaust database connections before Lambda reaches its own quota.

Store configuration securely, encrypt sensitive data, and use Secrets Manager or Parameter Store according to requirements. Cache secrets only within approved rotation and exposure boundaries.

5. Idempotency

At-least-once processing means duplicates are normal. Choose an idempotency key with business scope, such as tenant plus document version or payment request ID. Atomically create a record before side effects and persist outcome.

Model states: IN_PROGRESS, COMPLETED, FAILED_RETRYABLE, and where necessary FAILED_FINAL. Handle same key with different payload, concurrent claims, function timeout after side effect, stale in-progress state, replay after TTL, and downstream idempotency.

TTL controls storage; it does not guarantee immediate deletion. Retain keys longer than any valid retry, redrive, or replay window.

6. Concurrency and backpressure

Account concurrency is shared unless reserved. Reserved concurrency can guarantee capacity for a function and cap its impact; provisioned concurrency prepares environments for latency-sensitive workloads but costs while configured. Verify current quotas and Regional behavior.

Calculate:

approximate concurrency = requests per second x average duration seconds

This is a starting estimate. Include burst behavior, retries, batch size, downstream capacity, and tail duration. For a database supporting 200 concurrent operations, allowing a function to scale to thousands simply moves failure downstream. Use queues, reserved concurrency, event-source maximum concurrency, rate controls, or workflow limits.

7. Timeout, retry, and error handling

Set upstream timeout longer than the function's useful work but shorter than the user/business deadline. Set downstream HTTP/SDK timeouts explicitly. Classify retryable throttles and transient failures separately from validation, authorization, or permanent business errors.

Use bounded exponential backoff with jitter. Avoid retries at every layer. Track attempts per completed document. A DLQ or failure destination needs owner, alarm, sensitive-data policy, retention, diagnosis, remediation, replay rate, idempotency, and deletion criteria.

8. Orchestration

Step Functions makes sequence, branching, parallel work, waits, retries, catches, callbacks, and execution history explicit. Choose Standard or Express from duration, execution semantics, history, rate, and price requirements. Verify current service behavior before selection.

Lambda durable functions can provide application-centric orchestration with checkpointing and recovery for supported use cases. Compare it with Step Functions rather than assuming either universally replaces the other. Avoid using Lambda sleeps or home-built polling loops for long waits.

For a document workflow, orchestrate validation, malware scanning, text extraction, classification, human review where required, publication, and notification. Define compensation or forward repair for every side effect.

9. Security

Give each function a least-privilege execution role. Resource policies control who can invoke; execution roles control what code can call. Protect event sources, queues, buses, buckets, tables, KMS keys, secrets, destinations, and log access.

Validate input even from AWS services. Prevent event injection and confused-deputy access with source conditions where supported. Keep functions outside a VPC unless they need VPC resources; VPC attachment changes network/egress design but does not itself make a function private. Use endpoints or controlled NAT paths and account for cost.

Separate development and production accounts, protect deployments, scan dependencies/artifacts, sign where required, and avoid secrets or personal data in logs.

10. Observability

Use structured logs with request, correlation, tenant-safe, event, attempt, and outcome fields. Emit business metrics such as documents completed, failed, quarantined, and age-to-completion, plus Lambda errors, duration, throttles, concurrency, async event age, queue age/depth, DLQ, stream iterator age, workflow failure, and destination delivery.

Trace across API, function, queue/event, workflow, and downstream services where supported. Sampling must preserve incident utility without uncontrolled cost or sensitive data.

Alarm on user impact and backlog age, not every individual invocation failure. Runbooks must distinguish code error, malformed event, permissions, KMS, quota, downstream saturation, timeout, and poison item.

11. Performance and cost

Memory affects available CPU and can reduce duration. Benchmark representative payloads across memory settings; the cheapest invocation is not always the smallest memory. Measure cold and warm paths, initialization, dependency calls, package size, and p95/p99.

Cost includes requests, duration, provisioned concurrency, API calls, workflow transitions, queue/event requests, state, logs/traces, data transfer, NAT, retries, duplicate work, and development/operations. Model cost per completed business item at normal, burst, and failure conditions.

12. Read-only inventory

export AWS_DEFAULT_REGION="ap-south-1"
aws lambda list-functions --output table
aws lambda get-account-settings --output json
aws apigatewayv2 get-apis --output table
aws events list-rules --output table
aws sqs list-queues --output table
aws stepfunctions list-state-machines --output table

Inventory does not prove safe retries, least privilege, or business completion.

13. Guided workshop: document processing

Design for uploads from 20 tenants, 500 normal and 5,000 burst documents/minute, 50 MB maximum files, malware scan, extraction, classification, optional human approval, 15-minute completion target, tenant isolation, seven-year final-record retention, 30-day temporary-file retention, and replayable failures.

Produce:

  1. workload and serverless-fit decision;
  2. event catalog and schemas;
  3. payload claim-check design;
  4. function responsibility map;
  5. invocation-mode table;
  6. concurrency/downstream capacity model;
  7. timeout and retry ownership map;
  8. idempotency state design;
  9. SQS visibility, DLQ, and replay policy;
  10. workflow with error/approval paths;
  11. state, retention, and lifecycle model;
  12. identity, resource-policy, KMS, and tenant controls;
  13. observability and business-SLI design;
  14. normal/burst/failure quota model;
  15. performance experiment;
  16. cost-per-completed-document model;
  17. eight failure-injection tests;
  18. deployment, rollback/forward-repair, and acceptance plan.

14. Troubleshooting

SymptomInvestigate
Throttles despite account capacityReserved concurrency, event-source cap, downstream SDK throttle
Duplicate resultIdempotency claim/state, timeout after side effect, replay window
Queue age growsArrival rate, duration, concurrency, visibility, downstream limits
Stream iterator age growsPoison item, shard count, parallelization, function errors
Async event disappearsEvent age/retries, destination permission, application evidence
Cost spikeLoop, fan-out, retries, logs, provisioned concurrency, NAT transfer

Cost and cleanup

This lesson creates no AWS resources. Any later lab needs budget alarms, tags, capped concurrency, synthetic data, a manifest, and cleanup of functions, versions, mappings, APIs, queues/DLQs, rules, workflows, tables, buckets, logs, and alarms.

Knowledge check

  1. Does serverless mean no operations? No; customers own code, events, data, identity, quotas, evidence, and cost.
  2. Why cap concurrency? To protect shared quotas and downstream capacity.
  3. Why use idempotency for SQS? Messages can be delivered more than once.
  4. What should a DLQ include operationally? Alarm, owner, diagnosis, safe replay, retention, and deletion.
  5. Why store large payloads outside events? To respect limits and control data lifecycle and access.

Lesson acceptance

Submit all 18 artifacts. Invocation behavior, duplicate handling, concurrency, downstream protection, timeout/retry ownership, security, observability, quota, cost, and failure tests must be explicit and traceable to business acceptance.

Official sources

Advertisement