AWS 230: Remediate an EventBridge event with Lambda or Systems Manager
Why this lesson matters
Event-driven remediation can shorten an outage, but a broad pattern or non-idempotent action can repeat damage at machine speed. “The rule matched” is only the beginning: you must prove event origin, schema, target authorization, delivery, execution, duplicate handling, retries, dead-letter capture, resulting state, observability, and cleanup.
This lab deliberately starts with a record-only action. A matching custom event invokes Lambda, which validates the contract and conditionally records one business request in DynamoDB. It has no permission to modify a workload. Only after this safety gate is understood should a team replace the recorder with a reviewed, reversible remediation.
Outcomes
You will be able to:
- distinguish an event, event bus, rule, pattern, target, permission, retry
policy, target DLQ, and target execution;
- predict EventBridge pattern matching, including omitted fields and exact
string behavior;
- separate EventBridge delivery failure from Lambda execution failure;
- use a business request token for semantic idempotency;
- prevent loops using a narrow contract, state predicate, bounded action, and
concurrency/cost alarms;
- choose Lambda, Systems Manager Automation, Step Functions, or human approval;
- test positive, negative, duplicate, malformed, throttled, retry, and DLQ paths;
- prove least privilege, result state, cost ownership, and exact cleanup.
Safety boundary
- Local artifact review and
test-event-patternare read-only. Deployment,
enabling a rule, publishing an event, changing concurrency, reading/deleting a DLQ message, and stack deletion are mutations requiring owner approval.
- Use an isolated learning account and
ap-south-1; never root or production. - The stack deploys its rule disabled by default.
- The Lambda can write only its log stream and one owned DynamoDB table. It has
no EC2, IAM, network, S3, SSM mutation, or events:PutEvents permission.
- Test data contains no secret or personal information. Redact account IDs/ARNs.
- Do not turn this recorder into workload remediation during the exercise.
Event-driven remediation from first principles
approved producer
PutEvents -> default event bus
|
v
exact metadata + detail pattern
|
unmatched --+--> no target invocation
|
matched
v
EventBridge target delivery --(bounded retry)--> Lambda resource policy
| |
exhausted/immediate failure v
+--> standard SQS DLQ validate full contract
|
conditional DynamoDB PutItem
| |
first token same token
v v
recorded duplicate suppressed
An EventBridge event envelope includes fields such as version, id, detail-type, source, account, time, region, resources, and detail. The producer owns semantic truth inside detail; EventBridge does not infer that resourceId is safe merely because JSON is valid.
Components and ownership
| Component | Responsibility | Not proof of |
|---|---|---|
| Producer | emits correct schema and stable business token | target success |
| Event bus | accepts/routes events in its account/Region | durable business queue |
| Rule pattern | selects event fields/values | full schema validation |
| Target permission | lets EventBridge invoke target | target's downstream authority |
| Retry policy | retries target delivery within age/attempt bounds | exactly-once execution |
| Target DLQ | retains undelivered target events | automatic replay or remediation |
| Lambda | validates and performs bounded logic | exactly-once invocation |
| DynamoDB condition | accepts first business token only | instant TTL deletion |
| Logs/metrics/alarms | expose behavior and failures | repaired workload state |
Pattern semantics that architects must know
- Fields omitted from a pattern are effectively unconstrained. A pattern with
only source can match future detail types you never reviewed.
- Values are arrays of alternatives. Strings match character-for-character by
default; do not assume case folding.
- Numeric comparison uses JSON number representation rules that can surprise
designs treating equivalent-looking values identically.
prefix,suffix,anything-but,exists, numeric, IP, wildcard and$or
operators are useful but widen review scope.
- Duplicate JSON keys are unsafe; current processing uses only a final value in
cases documented by AWS. Reject duplicate keys before deployment.
- Pattern match is not JSON Schema validation. The Lambda validates the complete
contract again because producers and patterns can evolve independently.
This lab requires exact account, Region, source, detail type, schema version, environment, action, and resource ID, plus existence of requestToken. The Lambda then validates the token's format. Defense in depth protects against direct Lambda invocation and future pattern edits.
Lambda versus Systems Manager Automation
| Requirement | Prefer | Reason |
|---|---|---|
| Millisecond/second custom validation and API logic | Lambda | code, SDK, concurrency, versions/aliases |
| Auditable operational steps, approvals, waits, branches | Automation | runbook execution/step evidence |
| Long workflow with service integrations and compensation | Step Functions | explicit state machine and execution history |
| High-impact/ambiguous decision | Human approval/ticket | automation lacks safe deterministic authority |
An EventBridge rule invokes Lambda through the function's resource-based policy; the Lambda execution role controls what function code can do. For an Automation target, EventBridge normally assumes a target role allowed to call scoped ssm:StartAutomationExecution; the runbook's AutomationAssumeRole controls its steps, and the target may need tightly scoped iam:PassRole. These are separate identities. An approval never grants missing permission.
Download the reviewed artifact
- Complete artifact archive
- CloudFormation template
- Pattern fixture
- Positive event
- Negative event
- Repeatable live entry
- Metric query definitions
- Read-only Automation alternative
- Artifact README
The primary stack contains a rule, Lambda/function role/log group, encrypted on-demand DynamoDB idempotency table with TTL, encrypted standard SQS target DLQ/queue policy, Lambda invoke permission, and two alarms. Deletion policies are intentionally Delete for disposable lab data.
Gate 1: static and offline review
Replace 000000000000 in both full-event fixtures and pattern.json with the approved account. Do not alter the source/detail contract.
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --output json
aws events test-event-pattern \
--event-pattern file://pattern.json --event file://test-event.json
aws events test-event-pattern \
--event-pattern file://pattern.json --event file://negative-event.json
Require Result: true then Result: false. Also inspect the template and prove:
- rule default is
DISABLEDand target retry is age 300 seconds/2 attempts; - target DLQ is Standard, encrypted, retained 14 days, and its queue policy
allows only events.amazonaws.com from this rule/account;
- Lambda permission is restricted by rule ARN and source account;
- function role grants only owned log writes and table
PutItem; - reserved concurrency is 2, timeout 10 seconds, log retention 7 days;
- conditional key is
requestToken; TTL is one day and is not immediate cleanup; - function does not publish an event, preventing a direct recursive loop;
- alarms cover EventBridge failed invocation and visible DLQ messages.
test-event-pattern=true proves selection only. It does not invoke the target, exercise permissions, or execute Lambda.
Gate 2: deploy disabled
Validate against AWS's current CloudFormation schema, then create a change set:
stack_name="nw-p11-eventbridge-remediation"
aws cloudformation validate-template --template-body file://template.yaml
aws cloudformation deploy --stack-name "$stack_name" \
--template-file template.yaml \
--parameter-overrides EnableRule=false \
--capabilities CAPABILITY_NAMED_IAM \
--tags Project=NitWings-P11 Environment=lab Cleanup=stack-owned \
--no-fail-on-empty-changeset
aws cloudformation describe-stack-events --stack-name "$stack_name" \
--max-items 50 --output table
aws events describe-rule --name nw-p11-remediation-record --output json
Require CREATE_COMPLETE and State=DISABLED. CAPABILITY_NAMED_IAM is needed because the template creates a named role; it is acknowledgment, not permission elevation. Save exact outputs:
aws cloudformation describe-stacks --stack-name "$stack_name" \
--query 'Stacks[0].Outputs' --output table
aws events list-targets-by-rule --rule nw-p11-remediation-record --output json
aws lambda get-policy --function-name nw-p11-remediation-recorder --output json
aws iam get-role-policy --role-name "nw-p11-remediation-lambda-$AWS_DEFAULT_REGION" \
--policy-name WriteOwnedLogAndIdempotencyRecord --output json
aws sqs get-queue-attributes --queue-url "$(aws cloudformation describe-stacks \
--stack-name "$stack_name" --query 'Stacks[0].Outputs[?OutputKey==`DlqUrl`].OutputValue' \
--output text)" --attribute-names Policy SqsManagedSseEnabled MessageRetentionPeriod
Inspect decoded policy statements; names alone do not prove least privilege.
Gate 3: enable and prove positive delivery
Enable through source-controlled stack parameters, not an unrecorded console edit:
aws cloudformation deploy --stack-name "$stack_name" \
--template-file template.yaml --parameter-overrides EnableRule=true \
--capabilities CAPABILITY_NAMED_IAM \
--tags Project=NitWings-P11 Environment=lab Cleanup=stack-owned \
--no-fail-on-empty-changeset
aws events describe-rule --name nw-p11-remediation-record \
--query '{State:State,Pattern:EventPattern}' --output json
aws events put-events --entries file://put-events-entry.json
FailedEntryCount=0 proves bus ingestion, not target success. After allowing asynchronous delivery, prove the record and structured log:
aws dynamodb get-item --table-name nw-p11-remediation-idempotency \
--key '{"requestToken":{"S":"p11-live-request-0001"}}' \
--consistent-read --output json
aws logs tail /aws/lambda/nw-p11-remediation-recorder --since 10m --format short
aws cloudwatch get-metric-data --metric-data-queries file://metric-data-queries.json \
--start-time "$(date -u -d '15 minutes ago' +%FT%TZ)" \
--end-time "$(date -u +%FT%TZ)"
The item must say RECORDED_NO_WORKLOAD_CHANGE; the log must say recorded. Use the EventBridge rule's Invocations, FailedInvocations, InvocationsSentToDLQ, and InvocationsFailedToBeSentToDLQ metrics. A zero failure metric without an invocation datapoint is not success.
Gate 4: negative, malformed, and duplicate tests
The negative fixture must not match and therefore must create no target record. Directly invoking Lambda with malformed data should raise event contract rejected; this tests target validation but bypasses EventBridge and must be labelled as such.
Publish put-events-entry.json a second time. EventBridge assigns a new event ID, but the business request token remains the same. Require:
- still exactly one DynamoDB item for that token;
- the second structured log outcome
duplicate_suppressed; - no workload mutation and no DLQ message.
Exactly-once EventBridge delivery is not assumed. The conditional write makes this one business action idempotent. A TTL deletion can occur later than its expiry timestamp, so the token retention window must exceed realistic retry and replay windows.
Gate 5: retry and target DLQ failure injection
The target DLQ handles EventBridge target-delivery failures. It is different from Lambda asynchronous invocation destinations/DLQs, and different again from an SQS event-source queue's redrive policy.
For an isolated approved test, temporarily set reserved concurrency to zero, publish a new request token, and observe throttled delivery, bounded retries, failure metrics, alarm state, and eventual DLQ message after the configured age or attempts are exhausted. Restore concurrency to 2 immediately afterward, run CloudFormation drift detection, and require the function to return IN_SYNC.
aws lambda put-function-concurrency --function-name nw-p11-remediation-recorder \
--reserved-concurrent-executions 0
# Change only requestToken in a private copy, publish it, and monitor for <= 10 minutes.
aws lambda put-function-concurrency --function-name nw-p11-remediation-recorder \
--reserved-concurrent-executions 2
drift_id="$(aws cloudformation detect-stack-drift --stack-name "$stack_name" \
--query StackDriftDetectionId --output text)"
aws cloudformation describe-stack-drift-detection-status \
--stack-drift-detection-id "$drift_id" --output json
Read the DLQ with its exact output URL. Inspect body plus message attributes such as error code/message, retry attempts, exhausted condition, rule ARN, and target ARN. Do not delete it until captured and diagnosed. EventBridge can send certain non-retriable delivery failures directly to the DLQ. A DLQ retains evidence; it does not replay automatically.
Systems Manager Automation alternative
The supplied schema 0.3 runbook validates three parameters and returns APPROVED_RECORD_ONLY; it makes no AWS API mutation. Review it before optional creation:
aws ssm create-document --name nw-p11-remediation-decision \
--document-type Automation --document-format YAML \
--content file://automation-decision.yaml
aws ssm start-automation-execution --document-name nw-p11-remediation-decision \
--parameters ResourceId=lab-resource,Action=record,RequestToken=p11-manual-0001
For a direct EventBridge Automation target, create a distinct target role trusted by events.amazonaws.com, scope ssm:StartAutomationExecution to the exact automation definition/version, transform event fields into each required parameter, and configure retry/DLQ. If the runbook uses an Automation role, constrain iam:PassRole. Pin reviewed document versions; $DEFAULT can change.
Choose Automation when step-level history, approvals, waits, branches, or operational runbook reuse matter. It still needs idempotency and explicit compensation; Automation is not a database transaction and does not roll back arbitrary completed side effects automatically.
Loop prevention and production guardrails
Use multiple controls, not one hopeful filter:
- Match exact source/detail type/version/environment/action/resource state.
- Ensure remediation output cannot satisfy its own input pattern.
- Re-read actual state immediately before mutation; exit if already compliant.
- Use a stable idempotency key and bounded retention window.
- Limit IAM resources, actions, tags, account, and Region.
- Bound retry, event age, concurrency, batch size, and target count.
- Alarm on upper invocation rate, failure, throttling, DLQ, and cost.
- Provide disable-rule, reserved-concurrency-zero, and human escalation controls.
- Prefer reversible changes and model compensation/verification.
EventBridge allows multiple targets on a rule, but fan-out multiplies independent delivery/result paths. Use one target here and SNS/Step Functions when their semantics better express the design.
Troubleshooting matrix
| Symptom | Evidence | Likely boundary |
|---|---|---|
| Pattern test false | exact event/pattern JSON | case, nesting, account, Region, type/value mismatch |
| PutEvents failed entry | per-entry error code/message | producer permission, malformed/oversized entry, bus |
| Ingestion succeeds, no invocation | rule state/pattern and MatchedEvents/Invocations | disabled/wrong bus or unmatched event |
| FailedInvocations rises | Lambda policy, target ARN, CloudTrail, DLQ attributes | EventBridge could not deliver target |
| Lambda Invocations rises, Errors rises | function logs and error type | target delivered; function validation/code/downstream failed |
| Duplicate records | table key/condition and producer token | wrong idempotency key or non-conditional operation |
| Duplicate suppressed unexpectedly | token ownership and retention | producer reused business token or window too long |
| DLQ empty during failure | queue policy, retry age/attempts, DLQ failure metric | still retrying or cannot send to DLQ |
| Alarm stays OK | metric Region/dimension/period/missing-data | wrong metric identity or no datapoint |
| Recursive invocations | emitted event and pattern | remediation output rematches rule; disable immediately |
Always classify: ingestion, matching, delivery, function invocation, function logic, downstream API, result verification, or observability. Retrying the wrong layer can multiply duplicates.
Cost model
Cost dimensions include custom event ingestion, rule delivery, Lambda requests and duration, CloudWatch Logs ingestion/storage/queries, DynamoDB writes/storage, SQS requests/storage, alarms, Automation steps, KMS if selected, and any real remediated resource. Failed loops can increase cost and throttling quickly. Check current ap-south-1 prices, Free Tier eligibility, retention, and budgets.
Cleanup and negative inventory
Disable first and wait for in-flight delivery. Capture table item, logs, metrics, alarms, Lambda policy, queue policy, and any DLQ evidence. Delete only messages whose evidence has been retained. If you created the optional document, delete the exact owned name after executions finish.
aws cloudformation deploy --stack-name "$stack_name" --template-file template.yaml \
--parameter-overrides EnableRule=false --capabilities CAPABILITY_NAMED_IAM \
--no-fail-on-empty-changeset
aws events describe-rule --name nw-p11-remediation-record --query State --output text
aws cloudformation delete-stack --stack-name "$stack_name"
aws cloudformation wait stack-delete-complete --stack-name "$stack_name"
aws cloudformation list-stacks --stack-status-filter DELETE_COMPLETE \
--query 'StackSummaries[?StackName==`nw-p11-eventbridge-remediation`].[StackName,StackStatus]' \
--output table
aws events list-rules --name-prefix nw-p11-remediation --output table
aws lambda list-functions --query 'Functions[?FunctionName==`nw-p11-remediation-recorder`]' --output json
aws dynamodb describe-table --table-name nw-p11-remediation-idempotency
aws sqs list-queues --queue-name-prefix nw-p11-remediation-target-dlq
aws logs describe-log-groups --log-group-name-prefix /aws/lambda/nw-p11-remediation-recorder
For APIs that return not-found errors, record those exact expected errors. For prefix listings, require no exact match. Also verify roles, alarms, permissions, rule targets, and optional Automation document are absent. Stack deletion alone is not broad negative inventory.
No-create evidence path
Complete pattern positive/negative reasoning, resource/IAM inventory, producer- to-result diagram, Lambda contract walkthrough, duplicate race explanation, retry/DLQ timeline, Automation identity comparison, five troubleshooting cases, cost worksheet, and exact cleanup query set. Label this as static evidence; it does not prove live delivery or permissions.
Lesson acceptance
Submit redacted account/Region, pattern true/false results, template/IAM review, disabled deployment evidence, positive ingestion plus target/result evidence, negative and duplicate outcomes, retry/DLQ evidence (or fully reasoned no-create track), alarms, restored configuration, and exact cleanup inventory. Reject a submission that changes a workload, uses a broad pattern/admin role, equates PutEvents success with remediation, lacks an idempotency condition, deletes DLQ evidence prematurely, or leaves the rule/role/table/queue/logs behind.
Knowledge check
- Does
FailedEntryCount=0prove Lambda ran? No; it proves event-bus entry
acceptance. Rule and target evidence are separate.
- Why key by request token, not EventBridge ID? A retried business request
can be republished with another event ID; the semantic token remains stable.
- Does a target DLQ capture every Lambda code error? Not necessarily. It
covers EventBridge target-delivery failure; Lambda async execution has its own retry/failure controls.
- What permissions invoke Lambda? The function resource policy grants this
rule/account invocation; the execution role governs Lambda's downstream calls.
- Why can a narrow pattern still be insufficient? Pattern matching is not
complete schema/business validation, so the target validates again.
- When prefer Automation? For auditable operational steps, approvals,
waits/branches, and runbook reuse - with separate target/runbook roles.