AWS 401: Systems Manager Automation and Lambda remediation
Why this lesson matters
Automated remediation can reduce exposure and recovery time or rapidly amplify a bad diagnosis. Systems Manager Automation suits auditable multi-step AWS operations with approvals/rate controls; Lambda suits short event-driven code. Both require exact targets, preconditions, least privilege, idempotency, bounded concurrency, verification, rollback/compensation and an owner.
Selection model
| Need | Automation runbook | Lambda |
|---|---|---|
| Multi-step AWS workflow and outputs | Native document steps/branches | Code/state must be designed |
| Fleet targets and rate controls | Built-in targets, concurrency/errors | Event source/concurrency custom control |
| Human approval/change governance | Approval steps/Change Manager integration | External workflow required |
| Long operation | Automation execution model | Lambda runtime limit; orchestrate elsewhere |
| Complex library/logic | executeScript bounded or external | Natural fit within runtime/package limits |
| Rollback/compensation | Explicit steps/runbook | Explicit code/workflow |
Prefer no automatic mutation when diagnosis confidence, blast radius, data effect or rollback is uncertain. Initial automation may collect evidence, tag/quarantine, open an OpsItem and request approval rather than repair.
Automation runbook design
Version custom documents and execute an explicit approved version, not silently changing default. Define typed parameters with allowed patterns/values, precondition assertions, target parameter, AutomationAssumeRole, timeouts, retries, onFailure, outputs, verification and cleanup. Never accept an arbitrary role/resource/script parameter from an untrusted event.
At scale, use exact tags/resource groups plus maximum concurrency and maximum errors. Understand percentage rounding and that already-running executions can finish after threshold. Start with one canary and error threshold zero. Change Calendar/Change Manager can enforce windows and approval where required; emergency bypass needs separate authority/audit.
Lambda handlers validate schema/account/Region/resource, fetch current state, compute desired delta, use an idempotency record/conditional operation, apply one bounded change, verify result and emit evidence. Configure reserved concurrency, retries/DLQ/destination, timeout, memory, VPC only when necessary, secret handling and structured logs. Avoid recursive events by marking automation origin and matching only noncompliant state.
Identity, safety and lifecycle
Separate detector, router, starter, execution role, target resource role, approver and break-glass operator. The starter may pass only approved roles/documents/functions. Resource policies, KMS, SCP/boundary and cross-account trust remain part of authorization.
Every remediation contract states trigger, confidence, scope, desired state, invariants, idempotency key, maximum targets/concurrency, stop conditions, verification, rollback/compensation, nonreversible effects, timeout, escalation, evidence, owner and expiry. Revalidate after service/config changes. Disable noisy or unsafe automation without disabling detection.
Maintain a state machine for each target: observed noncompliant, evidence collected, eligible, approved, action started, verified, compensated/restored, exception or escalated. Store execution and target IDs so retries resume or stop safely. Verification must re-read authoritative service state and, where customer-facing, run a user/security outcome test. Tagging an execution successful or receiving HTTP 200 from an API is not enough.
Roll automation out like production software: unit/contract tests, policy checks, sandbox failure injection, one target, small percentage, bounded wave, then broad fleet. Record version adoption and stop older vulnerable runbooks/functions. Monitor trigger rate, eligibility rejection, success, verification failure, rollback/compensation, duration, concurrency, throttling, DLQ age, exception age and customer/security outcome.
Read-only inspection and workshop
aws ssm describe-document --name RUNBOOK --document-version VERSION --region ap-south-1
aws ssm describe-automation-executions --filters Key=DocumentNamePrefix,Values=RUNBOOK --region ap-south-1
aws ssm get-automation-execution --automation-execution-id EXECUTION_ID --region ap-south-1
aws lambda get-function-configuration --function-name FUNCTION --region ap-south-1
aws lambda get-function-event-invoke-config --function-name FUNCTION --region ap-south-1
Design one runbook to quarantine an overly permissive security-group rule and one Lambda to tag/open a ticket for an unencrypted-resource finding. Use fictional resources. Include evidence-only/dry-run, exact target, condition check, canary, approval threshold, idempotency, concurrency/error limits, verification, restore/exception process and notifications.
Test 20 failures: event spoofed, account/Region wrong, tag selects production, stale event after manual fix, runbook default changed, starter PassRole broad, execution role deny, SCP/KMS deny, concurrency too high, error threshold misunderstood, approval times out, calendar closed, step timeout, partial state, rollback fails, Lambda duplicate, recursive trigger, DLQ absent, verification checks control plane only, and exception immediately remediated again.
Cost and acceptance
Price Automation steps/executions/current pricing, Lambda requests/duration, EventBridge, DynamoDB idempotency, logs/metrics, Config/Security Hub, notifications and operational risk. This lesson creates nothing.
Submit decision matrix, both automation contracts, runbook/function pseudocode, IAM chains, rate/stop math, idempotency store, verification/rollback, 20 failures, cost and governance. Pass requires bounded least privilege, canary-first execution, no recursive loop, durable deduplication and measured user/security outcome.