Lesson 230 · AWS Learning Path

AWS 230: Remediate an EventBridge event with Lambda or Systems Manager

· Published · 12 min read

Labelled process diagram for AWS 230: Narrow operational event to EventBridge pattern and target role to Lambda or Automation runbook to Idempotent result, logs, retry, and DLQ, with decision, proof and rejection...

Why this lesson matters

Event-driven remediation can shorten an outage, but a broad pattern or non-idempotent action can repeat damage at machine speed. “The rule matched” is only the beginning: you must prove event origin, schema, target authorization, delivery, execution, duplicate handling, retries, dead-letter capture, resulting state, observability, and cleanup.

This lab deliberately starts with a record-only action. A matching custom event invokes Lambda, which validates the contract and conditionally records one business request in DynamoDB. It has no permission to modify a workload. Only after this safety gate is understood should a team replace the recorder with a reviewed, reversible remediation.

Outcomes

You will be able to:

  • distinguish an event, event bus, rule, pattern, target, permission, retry

policy, target DLQ, and target execution;

  • predict EventBridge pattern matching, including omitted fields and exact

string behavior;

  • separate EventBridge delivery failure from Lambda execution failure;
  • use a business request token for semantic idempotency;
  • prevent loops using a narrow contract, state predicate, bounded action, and

concurrency/cost alarms;

  • choose Lambda, Systems Manager Automation, Step Functions, or human approval;
  • test positive, negative, duplicate, malformed, throttled, retry, and DLQ paths;
  • prove least privilege, result state, cost ownership, and exact cleanup.

Safety boundary

  • Local artifact review and test-event-pattern are read-only. Deployment,

enabling a rule, publishing an event, changing concurrency, reading/deleting a DLQ message, and stack deletion are mutations requiring owner approval.

  • Use an isolated learning account and ap-south-1; never root or production.
  • The stack deploys its rule disabled by default.
  • The Lambda can write only its log stream and one owned DynamoDB table. It has

no EC2, IAM, network, S3, SSM mutation, or events:PutEvents permission.

  • Test data contains no secret or personal information. Redact account IDs/ARNs.
  • Do not turn this recorder into workload remediation during the exercise.

Event-driven remediation from first principles

approved producer
  PutEvents -> default event bus
                 |
                 v
 exact metadata + detail pattern
                 |
     unmatched --+--> no target invocation
                 |
              matched
                 v
 EventBridge target delivery --(bounded retry)--> Lambda resource policy
                 |                                  |
       exhausted/immediate failure                  v
                 +--> standard SQS DLQ       validate full contract
                                                    |
                                      conditional DynamoDB PutItem
                                          |                    |
                                      first token          same token
                                          v                    v
                                       recorded      duplicate suppressed

An EventBridge event envelope includes fields such as version, id, detail-type, source, account, time, region, resources, and detail. The producer owns semantic truth inside detail; EventBridge does not infer that resourceId is safe merely because JSON is valid.

Components and ownership

ComponentResponsibilityNot proof of
Produceremits correct schema and stable business tokentarget success
Event busaccepts/routes events in its account/Regiondurable business queue
Rule patternselects event fields/valuesfull schema validation
Target permissionlets EventBridge invoke targettarget's downstream authority
Retry policyretries target delivery within age/attempt boundsexactly-once execution
Target DLQretains undelivered target eventsautomatic replay or remediation
Lambdavalidates and performs bounded logicexactly-once invocation
DynamoDB conditionaccepts first business token onlyinstant TTL deletion
Logs/metrics/alarmsexpose behavior and failuresrepaired workload state

Pattern semantics that architects must know

  • Fields omitted from a pattern are effectively unconstrained. A pattern with

only source can match future detail types you never reviewed.

  • Values are arrays of alternatives. Strings match character-for-character by

default; do not assume case folding.

  • Numeric comparison uses JSON number representation rules that can surprise

designs treating equivalent-looking values identically.

  • prefix, suffix, anything-but, exists, numeric, IP, wildcard and $or

operators are useful but widen review scope.

  • Duplicate JSON keys are unsafe; current processing uses only a final value in

cases documented by AWS. Reject duplicate keys before deployment.

  • Pattern match is not JSON Schema validation. The Lambda validates the complete

contract again because producers and patterns can evolve independently.

This lab requires exact account, Region, source, detail type, schema version, environment, action, and resource ID, plus existence of requestToken. The Lambda then validates the token's format. Defense in depth protects against direct Lambda invocation and future pattern edits.

Lambda versus Systems Manager Automation

RequirementPreferReason
Millisecond/second custom validation and API logicLambdacode, SDK, concurrency, versions/aliases
Auditable operational steps, approvals, waits, branchesAutomationrunbook execution/step evidence
Long workflow with service integrations and compensationStep Functionsexplicit state machine and execution history
High-impact/ambiguous decisionHuman approval/ticketautomation lacks safe deterministic authority

An EventBridge rule invokes Lambda through the function's resource-based policy; the Lambda execution role controls what function code can do. For an Automation target, EventBridge normally assumes a target role allowed to call scoped ssm:StartAutomationExecution; the runbook's AutomationAssumeRole controls its steps, and the target may need tightly scoped iam:PassRole. These are separate identities. An approval never grants missing permission.

Download the reviewed artifact

The primary stack contains a rule, Lambda/function role/log group, encrypted on-demand DynamoDB idempotency table with TTL, encrypted standard SQS target DLQ/queue policy, Lambda invoke permission, and two alarms. Deletion policies are intentionally Delete for disposable lab data.

Gate 1: static and offline review

Replace 000000000000 in both full-event fixtures and pattern.json with the approved account. Do not alter the source/detail contract.

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --output json
aws events test-event-pattern \
  --event-pattern file://pattern.json --event file://test-event.json
aws events test-event-pattern \
  --event-pattern file://pattern.json --event file://negative-event.json

Require Result: true then Result: false. Also inspect the template and prove:

  • rule default is DISABLED and target retry is age 300 seconds/2 attempts;
  • target DLQ is Standard, encrypted, retained 14 days, and its queue policy

allows only events.amazonaws.com from this rule/account;

  • Lambda permission is restricted by rule ARN and source account;
  • function role grants only owned log writes and table PutItem;
  • reserved concurrency is 2, timeout 10 seconds, log retention 7 days;
  • conditional key is requestToken; TTL is one day and is not immediate cleanup;
  • function does not publish an event, preventing a direct recursive loop;
  • alarms cover EventBridge failed invocation and visible DLQ messages.

test-event-pattern=true proves selection only. It does not invoke the target, exercise permissions, or execute Lambda.

Gate 2: deploy disabled

Validate against AWS's current CloudFormation schema, then create a change set:

stack_name="nw-p11-eventbridge-remediation"
aws cloudformation validate-template --template-body file://template.yaml
aws cloudformation deploy --stack-name "$stack_name" \
  --template-file template.yaml \
  --parameter-overrides EnableRule=false \
  --capabilities CAPABILITY_NAMED_IAM \
  --tags Project=NitWings-P11 Environment=lab Cleanup=stack-owned \
  --no-fail-on-empty-changeset
aws cloudformation describe-stack-events --stack-name "$stack_name" \
  --max-items 50 --output table
aws events describe-rule --name nw-p11-remediation-record --output json

Require CREATE_COMPLETE and State=DISABLED. CAPABILITY_NAMED_IAM is needed because the template creates a named role; it is acknowledgment, not permission elevation. Save exact outputs:

aws cloudformation describe-stacks --stack-name "$stack_name" \
  --query 'Stacks[0].Outputs' --output table
aws events list-targets-by-rule --rule nw-p11-remediation-record --output json
aws lambda get-policy --function-name nw-p11-remediation-recorder --output json
aws iam get-role-policy --role-name "nw-p11-remediation-lambda-$AWS_DEFAULT_REGION" \
  --policy-name WriteOwnedLogAndIdempotencyRecord --output json
aws sqs get-queue-attributes --queue-url "$(aws cloudformation describe-stacks \
  --stack-name "$stack_name" --query 'Stacks[0].Outputs[?OutputKey==`DlqUrl`].OutputValue' \
  --output text)" --attribute-names Policy SqsManagedSseEnabled MessageRetentionPeriod

Inspect decoded policy statements; names alone do not prove least privilege.

Gate 3: enable and prove positive delivery

Enable through source-controlled stack parameters, not an unrecorded console edit:

aws cloudformation deploy --stack-name "$stack_name" \
  --template-file template.yaml --parameter-overrides EnableRule=true \
  --capabilities CAPABILITY_NAMED_IAM \
  --tags Project=NitWings-P11 Environment=lab Cleanup=stack-owned \
  --no-fail-on-empty-changeset
aws events describe-rule --name nw-p11-remediation-record \
  --query '{State:State,Pattern:EventPattern}' --output json
aws events put-events --entries file://put-events-entry.json

FailedEntryCount=0 proves bus ingestion, not target success. After allowing asynchronous delivery, prove the record and structured log:

aws dynamodb get-item --table-name nw-p11-remediation-idempotency \
  --key '{"requestToken":{"S":"p11-live-request-0001"}}' \
  --consistent-read --output json
aws logs tail /aws/lambda/nw-p11-remediation-recorder --since 10m --format short
aws cloudwatch get-metric-data --metric-data-queries file://metric-data-queries.json \
  --start-time "$(date -u -d '15 minutes ago' +%FT%TZ)" \
  --end-time "$(date -u +%FT%TZ)"

The item must say RECORDED_NO_WORKLOAD_CHANGE; the log must say recorded. Use the EventBridge rule's Invocations, FailedInvocations, InvocationsSentToDLQ, and InvocationsFailedToBeSentToDLQ metrics. A zero failure metric without an invocation datapoint is not success.

Gate 4: negative, malformed, and duplicate tests

The negative fixture must not match and therefore must create no target record. Directly invoking Lambda with malformed data should raise event contract rejected; this tests target validation but bypasses EventBridge and must be labelled as such.

Publish put-events-entry.json a second time. EventBridge assigns a new event ID, but the business request token remains the same. Require:

  • still exactly one DynamoDB item for that token;
  • the second structured log outcome duplicate_suppressed;
  • no workload mutation and no DLQ message.

Exactly-once EventBridge delivery is not assumed. The conditional write makes this one business action idempotent. A TTL deletion can occur later than its expiry timestamp, so the token retention window must exceed realistic retry and replay windows.

Gate 5: retry and target DLQ failure injection

The target DLQ handles EventBridge target-delivery failures. It is different from Lambda asynchronous invocation destinations/DLQs, and different again from an SQS event-source queue's redrive policy.

For an isolated approved test, temporarily set reserved concurrency to zero, publish a new request token, and observe throttled delivery, bounded retries, failure metrics, alarm state, and eventual DLQ message after the configured age or attempts are exhausted. Restore concurrency to 2 immediately afterward, run CloudFormation drift detection, and require the function to return IN_SYNC.

aws lambda put-function-concurrency --function-name nw-p11-remediation-recorder \
  --reserved-concurrent-executions 0
# Change only requestToken in a private copy, publish it, and monitor for <= 10 minutes.
aws lambda put-function-concurrency --function-name nw-p11-remediation-recorder \
  --reserved-concurrent-executions 2
drift_id="$(aws cloudformation detect-stack-drift --stack-name "$stack_name" \
  --query StackDriftDetectionId --output text)"
aws cloudformation describe-stack-drift-detection-status \
  --stack-drift-detection-id "$drift_id" --output json

Read the DLQ with its exact output URL. Inspect body plus message attributes such as error code/message, retry attempts, exhausted condition, rule ARN, and target ARN. Do not delete it until captured and diagnosed. EventBridge can send certain non-retriable delivery failures directly to the DLQ. A DLQ retains evidence; it does not replay automatically.

Systems Manager Automation alternative

The supplied schema 0.3 runbook validates three parameters and returns APPROVED_RECORD_ONLY; it makes no AWS API mutation. Review it before optional creation:

aws ssm create-document --name nw-p11-remediation-decision \
  --document-type Automation --document-format YAML \
  --content file://automation-decision.yaml
aws ssm start-automation-execution --document-name nw-p11-remediation-decision \
  --parameters ResourceId=lab-resource,Action=record,RequestToken=p11-manual-0001

For a direct EventBridge Automation target, create a distinct target role trusted by events.amazonaws.com, scope ssm:StartAutomationExecution to the exact automation definition/version, transform event fields into each required parameter, and configure retry/DLQ. If the runbook uses an Automation role, constrain iam:PassRole. Pin reviewed document versions; $DEFAULT can change.

Choose Automation when step-level history, approvals, waits, branches, or operational runbook reuse matter. It still needs idempotency and explicit compensation; Automation is not a database transaction and does not roll back arbitrary completed side effects automatically.

Loop prevention and production guardrails

Use multiple controls, not one hopeful filter:

  1. Match exact source/detail type/version/environment/action/resource state.
  2. Ensure remediation output cannot satisfy its own input pattern.
  3. Re-read actual state immediately before mutation; exit if already compliant.
  4. Use a stable idempotency key and bounded retention window.
  5. Limit IAM resources, actions, tags, account, and Region.
  6. Bound retry, event age, concurrency, batch size, and target count.
  7. Alarm on upper invocation rate, failure, throttling, DLQ, and cost.
  8. Provide disable-rule, reserved-concurrency-zero, and human escalation controls.
  9. Prefer reversible changes and model compensation/verification.

EventBridge allows multiple targets on a rule, but fan-out multiplies independent delivery/result paths. Use one target here and SNS/Step Functions when their semantics better express the design.

Troubleshooting matrix

SymptomEvidenceLikely boundary
Pattern test falseexact event/pattern JSONcase, nesting, account, Region, type/value mismatch
PutEvents failed entryper-entry error code/messageproducer permission, malformed/oversized entry, bus
Ingestion succeeds, no invocationrule state/pattern and MatchedEvents/Invocationsdisabled/wrong bus or unmatched event
FailedInvocations risesLambda policy, target ARN, CloudTrail, DLQ attributesEventBridge could not deliver target
Lambda Invocations rises, Errors risesfunction logs and error typetarget delivered; function validation/code/downstream failed
Duplicate recordstable key/condition and producer tokenwrong idempotency key or non-conditional operation
Duplicate suppressed unexpectedlytoken ownership and retentionproducer reused business token or window too long
DLQ empty during failurequeue policy, retry age/attempts, DLQ failure metricstill retrying or cannot send to DLQ
Alarm stays OKmetric Region/dimension/period/missing-datawrong metric identity or no datapoint
Recursive invocationsemitted event and patternremediation output rematches rule; disable immediately

Always classify: ingestion, matching, delivery, function invocation, function logic, downstream API, result verification, or observability. Retrying the wrong layer can multiply duplicates.

Cost model

Cost dimensions include custom event ingestion, rule delivery, Lambda requests and duration, CloudWatch Logs ingestion/storage/queries, DynamoDB writes/storage, SQS requests/storage, alarms, Automation steps, KMS if selected, and any real remediated resource. Failed loops can increase cost and throttling quickly. Check current ap-south-1 prices, Free Tier eligibility, retention, and budgets.

Cleanup and negative inventory

Disable first and wait for in-flight delivery. Capture table item, logs, metrics, alarms, Lambda policy, queue policy, and any DLQ evidence. Delete only messages whose evidence has been retained. If you created the optional document, delete the exact owned name after executions finish.

aws cloudformation deploy --stack-name "$stack_name" --template-file template.yaml \
  --parameter-overrides EnableRule=false --capabilities CAPABILITY_NAMED_IAM \
  --no-fail-on-empty-changeset
aws events describe-rule --name nw-p11-remediation-record --query State --output text
aws cloudformation delete-stack --stack-name "$stack_name"
aws cloudformation wait stack-delete-complete --stack-name "$stack_name"
aws cloudformation list-stacks --stack-status-filter DELETE_COMPLETE \
  --query 'StackSummaries[?StackName==`nw-p11-eventbridge-remediation`].[StackName,StackStatus]' \
  --output table
aws events list-rules --name-prefix nw-p11-remediation --output table
aws lambda list-functions --query 'Functions[?FunctionName==`nw-p11-remediation-recorder`]' --output json
aws dynamodb describe-table --table-name nw-p11-remediation-idempotency
aws sqs list-queues --queue-name-prefix nw-p11-remediation-target-dlq
aws logs describe-log-groups --log-group-name-prefix /aws/lambda/nw-p11-remediation-recorder

For APIs that return not-found errors, record those exact expected errors. For prefix listings, require no exact match. Also verify roles, alarms, permissions, rule targets, and optional Automation document are absent. Stack deletion alone is not broad negative inventory.

No-create evidence path

Complete pattern positive/negative reasoning, resource/IAM inventory, producer- to-result diagram, Lambda contract walkthrough, duplicate race explanation, retry/DLQ timeline, Automation identity comparison, five troubleshooting cases, cost worksheet, and exact cleanup query set. Label this as static evidence; it does not prove live delivery or permissions.

Lesson acceptance

Submit redacted account/Region, pattern true/false results, template/IAM review, disabled deployment evidence, positive ingestion plus target/result evidence, negative and duplicate outcomes, retry/DLQ evidence (or fully reasoned no-create track), alarms, restored configuration, and exact cleanup inventory. Reject a submission that changes a workload, uses a broad pattern/admin role, equates PutEvents success with remediation, lacks an idempotency condition, deletes DLQ evidence prematurely, or leaves the rule/role/table/queue/logs behind.

Knowledge check

  1. Does FailedEntryCount=0 prove Lambda ran? No; it proves event-bus entry

acceptance. Rule and target evidence are separate.

  1. Why key by request token, not EventBridge ID? A retried business request

can be republished with another event ID; the semantic token remains stable.

  1. Does a target DLQ capture every Lambda code error? Not necessarily. It

covers EventBridge target-delivery failure; Lambda async execution has its own retry/failure controls.

  1. What permissions invoke Lambda? The function resource policy grants this

rule/account invocation; the execution role governs Lambda's downstream calls.

  1. Why can a narrow pattern still be insufficient? Pattern matching is not

complete schema/business validation, so the target validates again.

  1. When prefer Automation? For auditable operational steps, approvals,

waits/branches, and runbook reuse - with separate target/runbook roles.

Official sources

Advertisement