Lesson 222 · AWS Learning Path

AWS 222: Drift detection and custom resources

· Published · 8 min read

Labelled process diagram for AWS 222: Desired template state to Actual resource and drift detection to Difference or custom provider event to Reconciliation and verified stack state, with decision, proof and...

Why this lesson matters

Drift is a difference between declared and actual supported properties - not automatically a security incident, and not proof that an unchecked resource is compliant. A custom resource extends CloudFormation with provider code, but that code becomes part of create, update, rollback and delete correctness. This lesson teaches both boundaries without deploying provider code.

What you will be able to do

By the end, you can:

  • distinguish stack drift-detection status, stack drift status, resource drift status and property difference type;
  • interpret IN_SYNC, DRIFTED, NOT_CHECKED, UNKNOWN, MODIFIED and DELETED;
  • explain explicit-property, unsupported-property, nested-stack and permission limitations;
  • choose revert actual state, update template, import, accept exception or replace;
  • explain current drift-aware change sets as a three-way comparison;
  • trace custom-resource Create/Update/Delete requests and required responses;
  • design stable physical IDs, idempotency, timeout, retries, least privilege and safe deletion;
  • choose a registry resource type instead of ad hoc custom logic when lifecycle modeling is needed.

Drift has two status dimensions

EvidenceMeaning
Detection operation DETECTION_IN_PROGRESS/COMPLETE/FAILEDwhether comparison finished, not whether resources match
Stack IN_SYNC/DRIFTED/NOT_CHECKED/UNKNOWNaggregate drift conclusion
Resource IN_SYNC/MODIFIED/DELETED/NOT_CHECKEDsupported resource comparison result
Property difference ADD/REMOVE/NOT_EQUALshape of an actual-versus-expected difference

CloudFormation compares expected values from the deployed template/parameters with actual values returned by a resource provider. It checks only supported resource types and properties explicitly declared in the template. Service defaults omitted from the template are generally outside the comparison. Some write-only, sensitive, normalized or service-managed values cannot be compared reliably.

An IN_SYNC stack can contain NOT_CHECKED resources, and a stack with no drift-capable resources can still be reported IN_SYNC. Drift is not a substitute for AWS Config, Security Hub, policy-as-code, runtime tests or data validation.

Nested and fleet drift boundaries

Running drift detection on a root stack does not recursively detect drift inside nested stacks; initiate it on each nested stack. StackSet drift is a fleet operation that aggregates stack-instance/target results and must use reviewed account/Region concurrency and failure tolerance. One successful StackSet operation does not prove every target remains in sync later.

Starting drift detection is a control-plane action. It reads target configurations and needs CloudFormation drift actions plus read permissions for each resource type. It does not change the target resource, but it is not “just a local read,” so use only an approved stack.

P11 drift scenarios

Assume P11 was deployed with its explicit properties. Analyze these supplied cases without making changes:

Actual changeLikely resultDecision questions
Environment tag manually changedMODIFIED / NOT_EQUAL when tag is trackedemergency exception or unauthorized change? revert or update desired state?
Public access block weakenedMODIFIEDrestore immediately through reviewed desired-state workflow; investigate actor
log group deleted manuallyDELETEDrestore required? retained evidence lost? update/incident path
service applies an omitted defaultnormally not comparedshould the security-relevant default be explicit?
KMS alias points to another keyKMS key drift has documented limitationsverify key/policy directly; do not claim in-sync encryption from drift alone
child stack resource changesroot detection does not recursedetect on child explicitly

Do not “fix drift” by deleting a resource or rerunning an update blindly. First preserve actual/expected values, actor/time, business intent, data and replacement risk.

Safe detection and polling

If an owner approves detection on an existing stack:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
owned_stack="exact-approved-stack-name"
detection_id="$(aws cloudformation detect-stack-drift \
  --stack-name "$owned_stack" --query StackDriftDetectionId --output text)"

while :; do
  status="$(aws cloudformation describe-stack-drift-detection-status \
    --stack-drift-detection-id "$detection_id" \
    --query DetectionStatus --output text)"
  printf '%s %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$status"
  case "$status" in
    DETECTION_COMPLETE) break ;;
    DETECTION_FAILED) exit 1 ;;
  esac
  sleep 10
done

aws cloudformation describe-stack-resource-drifts \
  --stack-name "$owned_stack" \
  --stack-resource-drift-status-filters MODIFIED DELETED NOT_CHECKED \
  --output json

Keep the operation ID. A failed operation is not an in-sync result; inspect DetectionStatusReason, permissions, resource support and stack status.

Reconciliation decisions

SituationPreferred direction
unauthorized/accidental actual changerestore actual state through reviewed CloudFormation update or approved direct correction
legitimate emergency changeencode it in template after review, or revert when emergency ends
existing unowned resource should become manageduse supported resource import with complete template and retention protection
intentional exceptiondocument owner, reason, expiration, compensating control and verification
immutable/replacement property driftplan backup, migration, cutover and rollback; do not force update
unsupported/unchecked propertyverify through service API/configuration control

Traditional change sets compare previous deployed state with desired state and can unintentionally revert out-of-band changes without clearly presenting actual state. Current drift-aware change sets use actual, previous and desired state for supported types, show drift details/ignored properties, preserve recognized AWS-managed changes, and attempt to converge actual state to desired state. They still have unsupported/write-only/immutable and cross-resource limitations. Review BeforeValueFrom, drift fields, replacement and ignored-property reasons before execution.

Why custom resources are different

A custom resource declares a ServiceToken for a Lambda function or SNS topic. CloudFormation sends a request and waits for the provider to send a response to the request's presigned S3 ResponseURL. If the provider cannot reach that URL or never responds, the stack waits until ServiceTimeout and fails. Default timeout is 3600 seconds; set a shorter realistic value rather than allowing a coding error to block an hour.

stack operation
     |
     v
Create / Update / Delete request -> provider
     |                              |
     |                              +-> external API / idempotency record
     |                              |
     +<---- SUCCESS or FAILED PUT to presigned ResponseURL
                    |
                    v
       physical ID + non-secret Data

Request and response contract

A request includes:

  • RequestType, unique RequestId, StackId, LogicalResourceId;
  • ResponseURL and provider ServiceToken;
  • current ResourceProperties;
  • on Update, OldResourceProperties;
  • on Update/Delete, the prior PhysicalResourceId.

The provider must return Status, Reason, PhysicalResourceId, StackId, RequestId, LogicalResourceId, optional NoEcho, and optional Data. Response identifiers must match the request. Do not log the presigned response URL or secret properties.

Correct lifecycle design

Create

  1. validate bounded properties;
  2. derive an idempotency key from stack/request/logical identity;
  3. check whether the intended external object already exists and is owned;
  4. create once, tag/record ownership, verify;
  5. return a stable physical ID and non-secret attributes.

Update

  1. compare old/current properties;
  2. update in place when identity is stable;
  3. return the same physical ID for in-place modification;
  4. return a different physical ID only when replacement is intended - CloudFormation then sends Delete for the old one.

Delete

  1. treat already-absent as successful;
  2. delete only the exact owned physical ID;
  3. respect retain/data-protection policy;
  4. tolerate retries and partial prior work;
  5. always send a response, including after a handled absence.

At-least-once/retry behavior means handlers must be idempotent. A Lambda timeout does not prove the external API stopped; retry must reconcile before creating again. Use bounded retries/backoff for dependencies and leave enough time to send a FAILED response.

Prefer modeled resource types for durable extensions

Before writing a custom resource, check AWS resource types and the CloudFormation registry. A private/public registry resource type has a schema and standardized create/read/update/delete/list handlers and can support drift when provisionable. Hooks validate operations, and modules package reusable configurations. A simple Lambda custom resource remains suitable for narrow one-off provisioning gaps, but it has weaker native read/drift modeling and more lifecycle code for you to own.

Custom-resource review exercise

Review this provider design: “Create a DNS entry in an external system and return its ID.” Produce:

  • schema for hostname, record type, value, TTL and ownership token;
  • stable physical ID from external zone + record identity;
  • exact IAM/secrets/network permissions;
  • Create duplicate/retry behavior;
  • Update in-place versus replacement rules;
  • Delete already-absent and ownership-mismatch behavior;
  • 300-second service timeout and shorter provider timeout;
  • redacted structured logs and correlation fields;
  • alarms/DLQ or destination appropriate to invocation model;
  • unit tests for Create twice, Update twice, Delete twice, partial failure and response failure.

Reject any design that returns success before external verification or generates a random physical ID on every retry.

Diagnose failures

SymptomProve firstCorrection
drift operation failsdetection reason and service read permissionadd exact read action or document unsupported state
unexpected driftexplicit expected value, normalization and actual service responseverify genuine difference before remediation
nested root says in syncchild detection timestampsdetect each child
custom resource waitsprovider invocation/log and response-URL egressrepair response path; shorten future timeout
duplicate external objectsrequest IDs/idempotency record/physical IDreconcile and make Create retry-safe
update triggers old Deletephysical ID changedreturn stable ID unless replacement intended
stack delete failsDelete invocation and provider responsemake already-absent delete successful; never skip ownership validation
secret exposedlogs, Reason, Data, outputsrotate, redact and redesign secret flow

Cost and cleanup

Drift detection itself does not provision P11 resources, but provider calls, Lambda duration/logs, SNS delivery, external APIs, registry extensions and retained resources can cost money. Frequent fleet drift scans also create operational/API load.

This lesson deploys no provider. If an approved drift operation was started, wait for its terminal result and retain the operation ID/evidence; there is no resource to delete. Do not remediate the inspected stack unless separately authorized.

Knowledge check

  1. Does NOT_CHECKED mean compliant?

No; CloudFormation did not compare that resource/property.

  1. Are omitted service defaults checked?

Drift generally compares explicitly declared properties.

  1. Does root drift recurse through nested stacks?

No; detect each nested stack separately.

  1. Why must a custom resource preserve physical ID on in-place update?

A changed ID signals replacement and triggers Delete for the old resource.

  1. What happens if no response reaches ResponseURL?

CloudFormation waits until the configured/default timeout and fails the operation.

Lesson acceptance

  • Drift operation, stack, resource and property statuses are interpreted separately.
  • P11 scenarios include supported, unchecked, deleted, security and nested cases.
  • Each drift case has an owner-approved reconciliation decision and replacement/data check.
  • Traditional and drift-aware change sets are distinguished with current limitations.
  • Custom Create/Update/Delete and response fields are mapped completely.
  • Provider design passes duplicate/retry/partial-failure/delete-idempotency and secret-redaction review.

Official sources

Advertisement