AWS 152: Troubleshoot permissions, retries, duplicate delivery, concurrency, and dead-letter handling
Why this lesson matters
Diagnose one permission fault, one retry or duplicate fault, one concurrency or age fault, and one dead-letter fault in the AWS151 flow, then delete it in dependency order.
Distributed failures are easiest to understand while the correlation ID, queue receive count, execution history and exact policy denial can still be observed together. This lab introduces one bounded fault at a time into the owned AWS151 stack, preserves the original evidence, repairs only the proven boundary, runs a changed retest and then destroys the complete stack.
What you will be able to do
By the end, you can:
- explain troubleshoot permissions, retries, duplicate delivery, concurrency, and dead-letter handling in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This lesson has an optional live path. Check current prices, obtain the account owner's approval, set a hard timer, use course tags, and complete the stated cleanup. The evidence path is a complete alternative.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Diagnose one permission fault, one retry or duplicate fault, one concurrency or age fault, and one dead-letter fault in the AWS151 flow, then delete it in dependency order. |
| Scope and boundary | The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Serverless break-fix and cleanup lab. |
| Evidence of success | Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Serverless break-fix and cleanup lab. |
| Cost model | The AWS151 timer remains active until all APIs, functions, mappings, queues, topics, rules, workflows, roles, policies, and owned log groups are removed and a delayed billing review is scheduled. |
| Safe rejection rule | Avoid random retries, adding AdministratorAccess, deleting the original evidence, purging unknown queues, or claiming cleanup from one empty page. |
How the request flows
+-------------------------+
| Observed failed order |
+-------------------------+
|
v
+---------------------------------------------+
| Logs, metrics, policy, and queue evidence |
+---------------------------------------------+
|
v
+-------------------------+
| One controlled repair |
+-------------------------+
|
v
+--------------------------------------+
| Changed retest and reverse cleanup |
+--------------------------------------+
For Serverless break-fix and cleanup lab, the important boundary is this: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Serverless break-fix and cleanup lab. Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Serverless break-fix and cleanup lab. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use evidence-led changes to learn distributed failure and duplicate-safe recovery. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid random retries, adding AdministratorAccess, deleting the original evidence, purging unknown queues, or claiming cleanup from one empty page. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Use the Console service search and open CloudWatch Logs and Metrics plus every
nw-p07-resource from AWS151; confirm the account and Region before reading the page. - Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
- Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
- Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws logs describe-log-groups --log-group-name-prefix /aws/lambda/nw-p07- --output table
aws lambda list-event-source-mappings --query 'EventSourceMappings[?contains(FunctionArn, `nw-p07-`)].{Id:UUID,State:State,Last:LastProcessingResult}' --output table
aws sqs list-queues --queue-name-prefix nw-p07- --output table
aws stepfunctions list-executions --state-machine-arn replace-with-nw-p07-state-machine-arn --max-results 10 --output table
Expected interpretation
A pass preserves the original error, proves the failed layer, changes one control, retests with a new correlation ID, rolls back the injected fault, and ends with zero nw-p07- resources.
Practical work
Use the supplied fault cards or instructor stack. For each fault record symptom, timestamp, hypothesis, decisive query, smallest repair, retest, regression check, rollback, and final deletion proof.
Fault cards and cleanup order
Use one fault at a time. Capture the failure before making a change.
| Fault | Evidence to inspect | Smallest repair | Required retest |
|---|---|---|---|
| intake cannot send to SQS | Lambda log, execution-role policy, queue ARN | restore only sqs:SendMessage on the owned queue | new order reaches source queue |
| worker receives a duplicate | message ID, receive count, business idempotency record | enforce the order ID as an idempotency key | replay does not repeat the side effect |
| queue age rises | queue age, visible count, Lambda concurrency and throttles | correct the proven concurrency or downstream limit | backlog drains within the stated bound |
| poison message never reaches DLQ | redrive policy, receive count, visibility, worker errors | correct redrive or failure handling | controlled bad message reaches DLQ |
Fault 1: permission direction
Use the supplied denied CloudTrail/Lambda evidence track by default. It shows intake's execution role missing events:PutEvents for the named bus. Identify why API Gateway invocation permission cannot repair that downstream denial. On an instructor-approved live stack only, replace the owned inline role policy with the supplied denied variant, send a new correlation ID, preserve the error/request ID, and immediately redeploy the unchanged CloudFormation template to restore the policy. Do not attach a managed administrator policy or edit an unowned role.
Required proof: API reached intake; intake failed at PutEvents; no workflow execution exists; the exact execution principal/action/bus ARN were identified; stack redeploy repaired only drift; changed request crossed all hops.
Fault 2: duplicate delivery and idempotency
Submit the same orderId twice with two correlation IDs. EventBridge and workflow may legitimately process both transport events. The first worker attempt conditionally claims the order ID in DynamoDB and publishes to SNS; the second must log duplicate_skipped and acknowledge the SQS message without a second publish. Explain the limitation: a crash after the idempotency claim but before SNS publication leaves a claimed-but-unpublished state. Propose PROCESSING/COMPLETED, lease expiry and outbox/transaction patterns for production.
Fault 3: concurrency and backlog
Set reserved concurrency of the owned worker only to zero for no more than two minutes, send three unique fake orders, and observe Lambda throttles plus SQS visible/oldest-age growth. Restore by deleting the function's reserved-concurrency setting, then prove the backlog drains without exceeding the timer. Do not raise account quota or increase downstream capacity. Record the exact before/fault/after values.
Fault 4: poison message and DLQ
Send one unique fake order with "poison":true. The worker intentionally returns that message ID in batchItemFailures; after the queue's bounded receive count it must appear in nw-p07-orders-dlq. Record source receive count, worker failures and DLQ message attributes without exposing sensitive content. Explain why the worker must not add an idempotency record for an unprocessed poison event. Do not redrive until a hypothetical code/schema correction is documented; delete the learning stack instead.
Teardown owned by CloudFormation
Export/redact required evidence, then delete the stack rather than manually deleting its resources:
set -euo pipefail
export AWS_DEFAULT_REGION="ap-south-1"
export STACK_NAME="nw-p07-serverless-order-flow"
aws cloudformation delete-stack --stack-name "$STACK_NAME"
aws cloudformation wait stack-delete-complete --stack-name "$STACK_NAME"
if aws cloudformation describe-stacks --stack-name "$STACK_NAME" >/dev/null 2>&1; then
echo "ERROR: stack still exists" >&2
exit 1
fi
The final describe-stacks pattern is supplemental; a failed command can also mean wrong Region/account/permission. Prove cleanup using the same caller/Region and scoped inventories for API Gateway APIs, Lambda functions/mappings, SQS queues, SNS topics, EventBridge buses/rules, Step Functions state machines, DynamoDB tables, IAM roles and log groups. Inspect CloudFormation deleted/failed events and retained-resource settings. Schedule a next-day Cost Explorer/billing review because billing data can lag.
Delete in reverse dependency order after the changed retests: API routes/integrations and API, event-source mapping, EventBridge targets/rules/bus, Step Functions executions and state machine, SNS subscriptions/topic, source queue and DLQ, Lambda functions, owned log groups after evidence retention, then inline policies and roles. Re-run every nw-p07- inventory from AWS151. A delete response without the empty after-inventory is not proof.
Diagnose this topic from its own evidence
For every fault use: expected state → observed timestamp/symptom → one correlation/business/message/execution ID → first missing boundary → hypothesis → decisive query → one repair → changed-ID retest → regression check → rollback. Preserve the original error before redeploying. API status alone cannot locate an asynchronous failure; correlate API access log, function log, EventBridge metric, workflow history, SQS age/receive count, DynamoDB conditional result, worker log and SNS metric.
If stack deletion fails, inspect stack events and the exact logical resource. Do not manually remove random dependencies. Correct drift/dependency only for the named stack, retry deletion, and verify every global/Regional inventory. IAM roles and log groups are commonly missed because they appear outside the application service consoles.
Cost and cleanup
The AWS151 timer remains active until all APIs, functions, mappings, queues, topics, rules, workflows, roles, policies, and owned log groups are removed and a delayed billing review is scheduled.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Diagnose one permission fault, one retry or duplicate fault, one concurrency or age fault, and one dead-letter fault in the AWS151 flow, then delete it in dependency order.
- Which scope or ownership boundary must be proved first?
Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Serverless break-fix and cleanup lab.
- What evidence is strong enough to accept the result?
Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Serverless break-fix and cleanup lab.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid random retries, adding AdministratorAccess, deleting the original evidence, purging unknown queues, or claiming cleanup from one empty page.
- Which cost dimensions and retained resources need an owner?
Expected direction: The AWS151 timer remains active until all APIs, functions, mappings, queues, topics, rules, workflows, roles, policies, and owned log groups are removed and a delayed billing review is scheduled.
Lesson acceptance
Pass when all four faults are diagnosed from preserved evidence; permission direction is correct; duplicate transport produces one business publication; controlled zero concurrency creates and then drains measurable backlog; poison reaches the DLQ after the configured count; and no broad access, purge or unknown deletion occurs. Final acceptance requires DELETE_COMPLETE, zero scoped nw-p07- resources across every listed service, retained evidence with secrets/account identifiers redacted, timer stopped and delayed billing review scheduled. Fail if any component is retained, a delete API response replaces inventory proof, AdministratorAccess is used, or a retry is claimed safe without idempotency evidence.