Lesson 152 · AWS Learning Path

AWS 152: Troubleshoot permissions, retries, duplicate delivery, concurrency, and dead-letter handling

· Published · 10 min read

Labelled process diagram for AWS 152: Observed failed order to Logs, metrics, policy, and queue evidence to One controlled repair to Changed retest and reverse cleanup, with decision, proof and rejection evidence.

Why this lesson matters

Diagnose one permission fault, one retry or duplicate fault, one concurrency or age fault, and one dead-letter fault in the AWS151 flow, then delete it in dependency order.

Distributed failures are easiest to understand while the correlation ID, queue receive count, execution history and exact policy denial can still be observed together. This lab introduces one bounded fault at a time into the owned AWS151 stack, preserves the original evidence, repairs only the proven boundary, runs a changed retest and then destroys the complete stack.

What you will be able to do

By the end, you can:

  • explain troubleshoot permissions, retries, duplicate delivery, concurrency, and dead-letter handling in plain language;
  • locate the current service controls in the AWS Management Console;
  • run the matching CloudShell or AWS CLI queries and explain every important field;
  • draw the identity, network, data, failure, and monitoring path;
  • choose the service from requirements and reject it when those requirements are absent;
  • diagnose a failed or misleading result from evidence;
  • state the cost owner and prove cleanup or a no-create result.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This lesson has an optional live path. Check current prices, obtain the account owner's approval, set a hard timer, use course tags, and complete the stated cleanup. The evidence path is a complete alternative.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeDiagnose one permission fault, one retry or duplicate fault, one concurrency or age fault, and one dead-letter fault in the AWS151 flow, then delete it in dependency order.
Scope and boundaryThe learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Serverless break-fix and cleanup lab.
Evidence of successSuccess means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Serverless break-fix and cleanup lab.
Cost modelThe AWS151 timer remains active until all APIs, functions, mappings, queues, topics, rules, workflows, roles, policies, and owned log groups are removed and a delayed billing review is scheduled.
Safe rejection ruleAvoid random retries, adding AdministratorAccess, deleting the original evidence, purging unknown queues, or claiming cleanup from one empty page.

How the request flows

+-------------------------+
|  Observed failed order  |
+-------------------------+
            |
            v
+---------------------------------------------+
|  Logs, metrics, policy, and queue evidence  |
+---------------------------------------------+
                      |
                      v
+-------------------------+
|  One controlled repair  |
+-------------------------+
            |
            v
+--------------------------------------+
|  Changed retest and reverse cleanup  |
+--------------------------------------+

For Serverless break-fix and cleanup lab, the important boundary is this: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Serverless break-fix and cleanup lab. Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Serverless break-fix and cleanup lab. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.

Architecture decision table

SituationDirectionReason
Requirement matchesUse evidence-led changes to learn distributed failure and duplicate-safe recovery.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid random retries, adding AdministratorAccess, deleting the original evidence, purging unknown queues, or claiming cleanup from one empty page.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Use the Console service search and open CloudWatch Logs and Metrics plus every nw-p07- resource from AWS151; confirm the account and Region before reading the page.
  2. Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
  3. Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
  4. Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws logs describe-log-groups --log-group-name-prefix /aws/lambda/nw-p07- --output table
aws lambda list-event-source-mappings --query 'EventSourceMappings[?contains(FunctionArn, `nw-p07-`)].{Id:UUID,State:State,Last:LastProcessingResult}' --output table
aws sqs list-queues --queue-name-prefix nw-p07- --output table
aws stepfunctions list-executions --state-machine-arn replace-with-nw-p07-state-machine-arn --max-results 10 --output table

Expected interpretation

A pass preserves the original error, proves the failed layer, changes one control, retests with a new correlation ID, rolls back the injected fault, and ends with zero nw-p07- resources.

Practical work

Use the supplied fault cards or instructor stack. For each fault record symptom, timestamp, hypothesis, decisive query, smallest repair, retest, regression check, rollback, and final deletion proof.

Fault cards and cleanup order

Use one fault at a time. Capture the failure before making a change.

FaultEvidence to inspectSmallest repairRequired retest
intake cannot send to SQSLambda log, execution-role policy, queue ARNrestore only sqs:SendMessage on the owned queuenew order reaches source queue
worker receives a duplicatemessage ID, receive count, business idempotency recordenforce the order ID as an idempotency keyreplay does not repeat the side effect
queue age risesqueue age, visible count, Lambda concurrency and throttlescorrect the proven concurrency or downstream limitbacklog drains within the stated bound
poison message never reaches DLQredrive policy, receive count, visibility, worker errorscorrect redrive or failure handlingcontrolled bad message reaches DLQ

Fault 1: permission direction

Use the supplied denied CloudTrail/Lambda evidence track by default. It shows intake's execution role missing events:PutEvents for the named bus. Identify why API Gateway invocation permission cannot repair that downstream denial. On an instructor-approved live stack only, replace the owned inline role policy with the supplied denied variant, send a new correlation ID, preserve the error/request ID, and immediately redeploy the unchanged CloudFormation template to restore the policy. Do not attach a managed administrator policy or edit an unowned role.

Required proof: API reached intake; intake failed at PutEvents; no workflow execution exists; the exact execution principal/action/bus ARN were identified; stack redeploy repaired only drift; changed request crossed all hops.

Fault 2: duplicate delivery and idempotency

Submit the same orderId twice with two correlation IDs. EventBridge and workflow may legitimately process both transport events. The first worker attempt conditionally claims the order ID in DynamoDB and publishes to SNS; the second must log duplicate_skipped and acknowledge the SQS message without a second publish. Explain the limitation: a crash after the idempotency claim but before SNS publication leaves a claimed-but-unpublished state. Propose PROCESSING/COMPLETED, lease expiry and outbox/transaction patterns for production.

Fault 3: concurrency and backlog

Set reserved concurrency of the owned worker only to zero for no more than two minutes, send three unique fake orders, and observe Lambda throttles plus SQS visible/oldest-age growth. Restore by deleting the function's reserved-concurrency setting, then prove the backlog drains without exceeding the timer. Do not raise account quota or increase downstream capacity. Record the exact before/fault/after values.

Fault 4: poison message and DLQ

Send one unique fake order with "poison":true. The worker intentionally returns that message ID in batchItemFailures; after the queue's bounded receive count it must appear in nw-p07-orders-dlq. Record source receive count, worker failures and DLQ message attributes without exposing sensitive content. Explain why the worker must not add an idempotency record for an unprocessed poison event. Do not redrive until a hypothetical code/schema correction is documented; delete the learning stack instead.

Teardown owned by CloudFormation

Export/redact required evidence, then delete the stack rather than manually deleting its resources:

set -euo pipefail
export AWS_DEFAULT_REGION="ap-south-1"
export STACK_NAME="nw-p07-serverless-order-flow"
aws cloudformation delete-stack --stack-name "$STACK_NAME"
aws cloudformation wait stack-delete-complete --stack-name "$STACK_NAME"
if aws cloudformation describe-stacks --stack-name "$STACK_NAME" >/dev/null 2>&1; then
  echo "ERROR: stack still exists" >&2
  exit 1
fi

The final describe-stacks pattern is supplemental; a failed command can also mean wrong Region/account/permission. Prove cleanup using the same caller/Region and scoped inventories for API Gateway APIs, Lambda functions/mappings, SQS queues, SNS topics, EventBridge buses/rules, Step Functions state machines, DynamoDB tables, IAM roles and log groups. Inspect CloudFormation deleted/failed events and retained-resource settings. Schedule a next-day Cost Explorer/billing review because billing data can lag.

Delete in reverse dependency order after the changed retests: API routes/integrations and API, event-source mapping, EventBridge targets/rules/bus, Step Functions executions and state machine, SNS subscriptions/topic, source queue and DLQ, Lambda functions, owned log groups after evidence retention, then inline policies and roles. Re-run every nw-p07- inventory from AWS151. A delete response without the empty after-inventory is not proof.

Diagnose this topic from its own evidence

For every fault use: expected state → observed timestamp/symptom → one correlation/business/message/execution ID → first missing boundary → hypothesis → decisive query → one repair → changed-ID retest → regression check → rollback. Preserve the original error before redeploying. API status alone cannot locate an asynchronous failure; correlate API access log, function log, EventBridge metric, workflow history, SQS age/receive count, DynamoDB conditional result, worker log and SNS metric.

If stack deletion fails, inspect stack events and the exact logical resource. Do not manually remove random dependencies. Correct drift/dependency only for the named stack, retry deletion, and verify every global/Regional inventory. IAM roles and log groups are commonly missed because they appear outside the application service consoles.

Cost and cleanup

The AWS151 timer remains active until all APIs, functions, mappings, queues, topics, rules, workflows, roles, policies, and owned log groups are removed and a delayed billing review is scheduled.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Diagnose one permission fault, one retry or duplicate fault, one concurrency or age fault, and one dead-letter fault in the AWS151 flow, then delete it in dependency order.

  1. Which scope or ownership boundary must be proved first?

Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Serverless break-fix and cleanup lab.

  1. What evidence is strong enough to accept the result?

Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Serverless break-fix and cleanup lab.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid random retries, adding AdministratorAccess, deleting the original evidence, purging unknown queues, or claiming cleanup from one empty page.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: The AWS151 timer remains active until all APIs, functions, mappings, queues, topics, rules, workflows, roles, policies, and owned log groups are removed and a delayed billing review is scheduled.

Lesson acceptance

Pass when all four faults are diagnosed from preserved evidence; permission direction is correct; duplicate transport produces one business publication; controlled zero concurrency creates and then drains measurable backlog; poison reaches the DLQ after the configured count; and no broad access, purge or unknown deletion occurs. Final acceptance requires DELETE_COMPLETE, zero scoped nw-p07- resources across every listed service, retained evidence with secrets/account identifiers redacted, timer stopped and delayed billing review scheduled. Fail if any component is retained, a delete API response replaces inventory proof, AdministratorAccess is used, or a retry is claimed safe without idempotency evidence.

Official sources

Advertisement