Lesson 142 · AWS Learning Path

AWS 142: Lambda concurrency, scaling and retries

· Published · 8 min read

Labelled process diagram for AWS 142: Burst demand to Concurrency control to Function and downstream capacity to Throttle, retry, age, and DLQ evidence, with decision, proof and rejection evidence.

Why this lesson matters

Control account concurrency, reserved concurrency, provisioned concurrency, asynchronous retries, throttles, and downstream pressure as separate mechanisms.

What you will be able to do

By the end, you can:

  • explain lambda concurrency, scaling and retries in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeControl account concurrency, reserved concurrency, provisioned concurrency, asynchronous retries, throttles, and downstream pressure as separate mechanisms.
Scope and boundaryThe learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Lambda concurrency, scaling, and retries.
Evidence of successSuccess means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Lambda concurrency, scaling, and retries.
Cost modelRequests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Safe rejection ruleAvoid unbounded concurrency against a limited database or retrying non-idempotent work without duplicate protection.

How the request flows

+----------------------+
|     Burst demand     |
+----------------------+
           |
           v
+-----------------------+
|  Concurrency control  |
+-----------------------+
           |
           v
+------------------------------------+
|  Function and downstream capacity  |
+------------------------------------+
                  |
                  v
+------------------------------------------+
|  Throttle, retry, age, and DLQ evidence  |
+------------------------------------------+

For Lambda concurrency, scaling, and retries, the important boundary is this: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Lambda concurrency, scaling, and retries. Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Lambda concurrency, scaling, and retries. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.

Architecture decision table

SituationDirectionReason
Requirement matchesUse reserved concurrency to protect a function or the shared account pool and provisioned concurrency only for measured latency needs.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid unbounded concurrency against a limited database or retrying non-idempotent work without duplicate protection.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

Calculate demand before setting controls

Concurrency is simultaneous in-flight work, not requests per second. A first approximation is average requests/second × average duration seconds; capacity planning uses high-percentile duration and burst shape. At 200 requests/s and 0.4 seconds average, expected concurrency is about 80, but a two-second dependency slowdown raises it toward 400 without any traffic increase. That feedback can exhaust database connections and the account concurrency pool.

The Regional account concurrency quota is shared. Reserved concurrency both guarantees up to the configured amount for one function and caps that function, protecting downstream systems and other functions. Setting it to zero intentionally throttles the function. Unreserved functions share the remaining pool, with a portion retained for functions without reservations according to current Lambda rules. Provisioned concurrency pre-initializes a number of environments on a published version or alias to reduce startup latency; it is a paid readiness feature, cannot target $LATEST, can spill into on-demand concurrency and does not independently cap total invocation. It must fit under reserved/account limits.

Scaling behavior depends on invocation source and current Regional limits. SQS event-source mappings poll and scale differently from synchronous APIs and streams; maximum ESM concurrency, reserved function concurrency and queue visibility/redrive settings must be coherent. Newer SQS provisioned poller mode and Lambda Managed Instances have distinct scaling and prices and must be selected explicitly from current documentation - not blended into default Lambda rules.

Retry ownership and backpressure

  • A synchronous caller receives the function error/timeout and decides whether/how to retry.
  • Lambda asynchronous invocation normally queues and retries function errors according to configured attempts/maximum event age, then sends success/failure records to a destination or legacy DLQ as configured.
  • SQS redelivery is controlled by visibility timeout and redrive policy; Lambda polling/batch response controls which messages are acknowledged.
  • Streams retain and retry around ordered shard checkpoints; one poison record can raise iterator age until retry/age/bisect/failure policies advance it.

Retries consume concurrency and can create a storm. Use exponential backoff with jitter, bounded attempts, a total time budget, circuit breaking and a durable idempotency key. Invalid input should be quarantined rather than retried. Queue buffering smooths a producer burst, but backlog age measures user delay and retained messages can still expire.

Protect a 100-connection database by allocating fewer function concurrency units than usable connections after reserving operations and other clients, or by using RDS Proxy/pooling plus measured transactions. Batch processing may reduce connections but increases duration and failure scope. Load test until one defined stop condition - connection utilization, p99 latency, error/throttle rate or queue age - is reached, then set alarms below the failure point.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Use the Console service search and open Lambda, Functions, Configuration, Concurrency and asynchronous invocation; confirm the account and Region before reading the page.
  2. Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
  3. Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
  4. Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws lambda get-account-settings --query 'AccountLimit.{ConcurrentExecutions:ConcurrentExecutions,Unreserved:UnreservedConcurrentExecutions}' --output table
aws lambda list-functions --query 'Functions[].{Name:FunctionName,Reserved:ReservedConcurrentExecutions}' --output table
aws lambda list-function-event-invoke-configs --function-name replace-with-function-name --output json

Expected interpretation

Limits and configured values describe capacity controls. Use ConcurrentExecutions, Throttles, duration, errors, age, downstream metrics, and retry evidence to explain behavior.

Practical work

Given a 2,000-message burst and a database limit of 100 connections, set a concurrency guard, batch plan, retry policy, DLQ or destination, alarm, and load-test stop condition.

Show calculations for batch sizes 1, 10 and 100; a 0.5-second and 5-second handler; and one versus ten messages per database transaction. Reserve at least 20% database connection headroom and justify the final cap. Draw normal flow, throttled Lambda, slow database, poison message, DLQ redrive and replay. Include rollback of reserved/provisioned/ESM concurrency and prove the final configuration inventory.

Diagnose this topic from its own evidence

Use ConcurrentExecutions, ClaimedAccountConcurrency, UnreservedConcurrentExecutions, Throttles, Errors, duration percentiles and provisioned concurrency utilization/spillover together with source queue depth/age or stream iterator age and downstream saturation. Throttles with a flat function cap indicate reserved/ESM concurrency; widespread account throttles indicate the shared pool; growing queue age without throttles but with errors indicates handler failure and poller scale-down; high spillover means provisioned capacity is below simultaneous latency-sensitive demand. A concurrency increase is unsafe when downstream utilization is already at its ceiling.

Cost and cleanup

Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Control account concurrency, reserved concurrency, provisioned concurrency, asynchronous retries, throttles, and downstream pressure as separate mechanisms.

  1. Which scope or ownership boundary must be proved first?

Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Lambda concurrency, scaling, and retries.

  1. What evidence is strong enough to accept the result?

Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Lambda concurrency, scaling, and retries.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid unbounded concurrency against a limited database or retrying non-idempotent work without duplicate protection.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.

Lesson acceptance

Pass when the learner calculates concurrency from rate and duration, distinguishes account/unreserved/reserved/provisioned and ESM controls, and protects a quantified downstream limit. The design must map retry ownership for synchronous, asynchronous, queue and stream paths; implement bounded jittered retries and idempotency; isolate poison records; alarm on throttles, age/lag and dependency saturation; and include load-stop, replay, rollback and full cost evidence. Fail if provisioned concurrency is described as a maximum cap, quota increase is the first answer to downstream overload, or DLQ presence is treated as successful recovery.

Official sources

Advertisement