AWS 222: Drift detection and custom resources
Why this lesson matters
Drift is a difference between declared and actual supported properties - not automatically a security incident, and not proof that an unchecked resource is compliant. A custom resource extends CloudFormation with provider code, but that code becomes part of create, update, rollback and delete correctness. This lesson teaches both boundaries without deploying provider code.
What you will be able to do
By the end, you can:
- distinguish stack drift-detection status, stack drift status, resource drift status and property difference type;
- interpret
IN_SYNC,DRIFTED,NOT_CHECKED,UNKNOWN,MODIFIEDandDELETED; - explain explicit-property, unsupported-property, nested-stack and permission limitations;
- choose revert actual state, update template, import, accept exception or replace;
- explain current drift-aware change sets as a three-way comparison;
- trace custom-resource Create/Update/Delete requests and required responses;
- design stable physical IDs, idempotency, timeout, retries, least privilege and safe deletion;
- choose a registry resource type instead of ad hoc custom logic when lifecycle modeling is needed.
Drift has two status dimensions
| Evidence | Meaning |
|---|---|
Detection operation DETECTION_IN_PROGRESS/COMPLETE/FAILED | whether comparison finished, not whether resources match |
Stack IN_SYNC/DRIFTED/NOT_CHECKED/UNKNOWN | aggregate drift conclusion |
Resource IN_SYNC/MODIFIED/DELETED/NOT_CHECKED | supported resource comparison result |
Property difference ADD/REMOVE/NOT_EQUAL | shape of an actual-versus-expected difference |
CloudFormation compares expected values from the deployed template/parameters with actual values returned by a resource provider. It checks only supported resource types and properties explicitly declared in the template. Service defaults omitted from the template are generally outside the comparison. Some write-only, sensitive, normalized or service-managed values cannot be compared reliably.
An IN_SYNC stack can contain NOT_CHECKED resources, and a stack with no drift-capable resources can still be reported IN_SYNC. Drift is not a substitute for AWS Config, Security Hub, policy-as-code, runtime tests or data validation.
Nested and fleet drift boundaries
Running drift detection on a root stack does not recursively detect drift inside nested stacks; initiate it on each nested stack. StackSet drift is a fleet operation that aggregates stack-instance/target results and must use reviewed account/Region concurrency and failure tolerance. One successful StackSet operation does not prove every target remains in sync later.
Starting drift detection is a control-plane action. It reads target configurations and needs CloudFormation drift actions plus read permissions for each resource type. It does not change the target resource, but it is not “just a local read,” so use only an approved stack.
P11 drift scenarios
Assume P11 was deployed with its explicit properties. Analyze these supplied cases without making changes:
| Actual change | Likely result | Decision questions |
|---|---|---|
| Environment tag manually changed | MODIFIED / NOT_EQUAL when tag is tracked | emergency exception or unauthorized change? revert or update desired state? |
| Public access block weakened | MODIFIED | restore immediately through reviewed desired-state workflow; investigate actor |
| log group deleted manually | DELETED | restore required? retained evidence lost? update/incident path |
| service applies an omitted default | normally not compared | should the security-relevant default be explicit? |
| KMS alias points to another key | KMS key drift has documented limitations | verify key/policy directly; do not claim in-sync encryption from drift alone |
| child stack resource changes | root detection does not recurse | detect on child explicitly |
Do not “fix drift” by deleting a resource or rerunning an update blindly. First preserve actual/expected values, actor/time, business intent, data and replacement risk.
Safe detection and polling
If an owner approves detection on an existing stack:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
owned_stack="exact-approved-stack-name"
detection_id="$(aws cloudformation detect-stack-drift \
--stack-name "$owned_stack" --query StackDriftDetectionId --output text)"
while :; do
status="$(aws cloudformation describe-stack-drift-detection-status \
--stack-drift-detection-id "$detection_id" \
--query DetectionStatus --output text)"
printf '%s %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$status"
case "$status" in
DETECTION_COMPLETE) break ;;
DETECTION_FAILED) exit 1 ;;
esac
sleep 10
done
aws cloudformation describe-stack-resource-drifts \
--stack-name "$owned_stack" \
--stack-resource-drift-status-filters MODIFIED DELETED NOT_CHECKED \
--output json
Keep the operation ID. A failed operation is not an in-sync result; inspect DetectionStatusReason, permissions, resource support and stack status.
Reconciliation decisions
| Situation | Preferred direction |
|---|---|
| unauthorized/accidental actual change | restore actual state through reviewed CloudFormation update or approved direct correction |
| legitimate emergency change | encode it in template after review, or revert when emergency ends |
| existing unowned resource should become managed | use supported resource import with complete template and retention protection |
| intentional exception | document owner, reason, expiration, compensating control and verification |
| immutable/replacement property drift | plan backup, migration, cutover and rollback; do not force update |
| unsupported/unchecked property | verify through service API/configuration control |
Traditional change sets compare previous deployed state with desired state and can unintentionally revert out-of-band changes without clearly presenting actual state. Current drift-aware change sets use actual, previous and desired state for supported types, show drift details/ignored properties, preserve recognized AWS-managed changes, and attempt to converge actual state to desired state. They still have unsupported/write-only/immutable and cross-resource limitations. Review BeforeValueFrom, drift fields, replacement and ignored-property reasons before execution.
Why custom resources are different
A custom resource declares a ServiceToken for a Lambda function or SNS topic. CloudFormation sends a request and waits for the provider to send a response to the request's presigned S3 ResponseURL. If the provider cannot reach that URL or never responds, the stack waits until ServiceTimeout and fails. Default timeout is 3600 seconds; set a shorter realistic value rather than allowing a coding error to block an hour.
stack operation
|
v
Create / Update / Delete request -> provider
| |
| +-> external API / idempotency record
| |
+<---- SUCCESS or FAILED PUT to presigned ResponseURL
|
v
physical ID + non-secret Data
Request and response contract
A request includes:
RequestType, uniqueRequestId,StackId,LogicalResourceId;ResponseURLand providerServiceToken;- current
ResourceProperties; - on Update,
OldResourceProperties; - on Update/Delete, the prior
PhysicalResourceId.
The provider must return Status, Reason, PhysicalResourceId, StackId, RequestId, LogicalResourceId, optional NoEcho, and optional Data. Response identifiers must match the request. Do not log the presigned response URL or secret properties.
Correct lifecycle design
Create
- validate bounded properties;
- derive an idempotency key from stack/request/logical identity;
- check whether the intended external object already exists and is owned;
- create once, tag/record ownership, verify;
- return a stable physical ID and non-secret attributes.
Update
- compare old/current properties;
- update in place when identity is stable;
- return the same physical ID for in-place modification;
- return a different physical ID only when replacement is intended - CloudFormation then sends Delete for the old one.
Delete
- treat already-absent as successful;
- delete only the exact owned physical ID;
- respect retain/data-protection policy;
- tolerate retries and partial prior work;
- always send a response, including after a handled absence.
At-least-once/retry behavior means handlers must be idempotent. A Lambda timeout does not prove the external API stopped; retry must reconcile before creating again. Use bounded retries/backoff for dependencies and leave enough time to send a FAILED response.
Prefer modeled resource types for durable extensions
Before writing a custom resource, check AWS resource types and the CloudFormation registry. A private/public registry resource type has a schema and standardized create/read/update/delete/list handlers and can support drift when provisionable. Hooks validate operations, and modules package reusable configurations. A simple Lambda custom resource remains suitable for narrow one-off provisioning gaps, but it has weaker native read/drift modeling and more lifecycle code for you to own.
Custom-resource review exercise
Review this provider design: “Create a DNS entry in an external system and return its ID.” Produce:
- schema for hostname, record type, value, TTL and ownership token;
- stable physical ID from external zone + record identity;
- exact IAM/secrets/network permissions;
- Create duplicate/retry behavior;
- Update in-place versus replacement rules;
- Delete already-absent and ownership-mismatch behavior;
- 300-second service timeout and shorter provider timeout;
- redacted structured logs and correlation fields;
- alarms/DLQ or destination appropriate to invocation model;
- unit tests for Create twice, Update twice, Delete twice, partial failure and response failure.
Reject any design that returns success before external verification or generates a random physical ID on every retry.
Diagnose failures
| Symptom | Prove first | Correction |
|---|---|---|
| drift operation fails | detection reason and service read permission | add exact read action or document unsupported state |
| unexpected drift | explicit expected value, normalization and actual service response | verify genuine difference before remediation |
| nested root says in sync | child detection timestamps | detect each child |
| custom resource waits | provider invocation/log and response-URL egress | repair response path; shorten future timeout |
| duplicate external objects | request IDs/idempotency record/physical ID | reconcile and make Create retry-safe |
| update triggers old Delete | physical ID changed | return stable ID unless replacement intended |
| stack delete fails | Delete invocation and provider response | make already-absent delete successful; never skip ownership validation |
| secret exposed | logs, Reason, Data, outputs | rotate, redact and redesign secret flow |
Cost and cleanup
Drift detection itself does not provision P11 resources, but provider calls, Lambda duration/logs, SNS delivery, external APIs, registry extensions and retained resources can cost money. Frequent fleet drift scans also create operational/API load.
This lesson deploys no provider. If an approved drift operation was started, wait for its terminal result and retain the operation ID/evidence; there is no resource to delete. Do not remediate the inspected stack unless separately authorized.
Knowledge check
- Does
NOT_CHECKEDmean compliant?
No; CloudFormation did not compare that resource/property.
- Are omitted service defaults checked?
Drift generally compares explicitly declared properties.
- Does root drift recurse through nested stacks?
No; detect each nested stack separately.
- Why must a custom resource preserve physical ID on in-place update?
A changed ID signals replacement and triggers Delete for the old resource.
- What happens if no response reaches
ResponseURL?
CloudFormation waits until the configured/default timeout and fails the operation.
Lesson acceptance
- Drift operation, stack, resource and property statuses are interpreted separately.
- P11 scenarios include supported, unchecked, deleted, security and nested cases.
- Each drift case has an owner-approved reconciliation decision and replacement/data check.
- Traditional and drift-aware change sets are distinguished with current limitations.
- Custom Create/Update/Delete and response fields are mapped completely.
- Provider design passes duplicate/retry/partial-failure/delete-idempotency and secret-redaction review.