AWS 375: Drift detection, desired-state remediation, and configuration ownership
Why this lesson matters
Drift means actual resource configuration differs from the state declared to CloudFormation. Detection is useful only when an owner can distinguish harmful change, approved emergency change, AWS-managed behavior, unsupported visibility, and stale desired state, then restore one authoritative owner safely.
Drift model and limits
CloudFormation compares explicitly declared supported properties against actual values. Defaults omitted from a template might not be tracked. Unsupported resource types/properties are NOT_CHECKED; a stack with no supported resources can still appear IN_SYNC. Nested stacks require their own checks. Lambda code content and other opaque artifacts may not be reconstructable from a property comparison.
| Status | Meaning | Required interpretation |
|---|---|---|
IN_SYNC | Checked properties match | Not proof of full security or runtime correctness |
MODIFIED | At least one checked property differs | Identify actor, intent, and operational effect |
DELETED | Managed physical resource is absent | Do not recreate before checking data and dependencies |
NOT_CHECKED | Type/property not evaluated | Add another control and owner |
| Detection failed | CloudFormation could not evaluate | Diagnose permission, state, quota, or provider |
Drift-aware change sets can compare actual, previous deployment, and desired state and propose reverting supported drift. They preserve some AWS-managed/external-tag behavior and have unsupported or immutable-property limits. Review every BeforeValueFrom, ignored property, replacement, cross-stack attachment, and write-only-property fallback before execution.
Ownership and remediation decisions
Classify each difference:
- Unauthorized or accidental: contain access, preserve evidence, then revert through the authoritative pipeline.
- Emergency approved: decide whether to encode actual state in source or restore original state after the incident.
- Desired state is wrong: update reviewed source and deploy, rather than overwriting the correct live fix.
- Another controller owns the property: remove conflicting ownership or explicitly ignore/document the boundary.
- Unmanaged resource: import when supported and accurate, deliberately leave externally owned, or retire it.
- False/limited visibility: document unsupported property and add Config, service API, policy, or runtime evidence.
Never run a blanket remediation over a production emergency change. Record physical/logical ID, property, previous/actual/desired values, actor/event, ticket, customer effect, security implication, data/replacement risk, chosen owner, approval, and verification.
Prioritize identity, public access, encryption, logging, backup/retention, routing, and destructive replacement drift above cosmetic tags. A security-significant finding needs incident triage in parallel with infrastructure reconciliation. If immediate containment changes live state again, preserve both event chains and update the desired-state decision after containment. Do not allow the desire for a clean drift report to override forensic or availability requirements.
Competing controllers create oscillation: CloudFormation, Auto Scaling, Kubernetes, security auto-remediation, a console operator, and a separate IaC stack may continuously overwrite one another. Define property-level ownership. AWS-managed scaling values and incident-response controls need explicit exceptions that do not become permanent undocumented drift.
Detection and evidence
aws cloudformation detect-stack-drift --stack-name STACK --region ap-south-1
aws cloudformation describe-stack-drift-detection-status --stack-drift-detection-id ID --region ap-south-1
aws cloudformation describe-stack-resource-drifts --stack-name STACK --stack-resource-drift-status-filters MODIFIED DELETED --region ap-south-1
aws cloudformation describe-change-set --stack-name STACK --change-set-name DRIFT_CHANGE --region ap-south-1
aws cloudtrail lookup-events --lookup-attributes AttributeKey=ResourceName,AttributeValue=RESOURCE --region ap-south-1
Detection is asynchronous; wait for completion and record timestamp. Compare processed template, parameters, current service API, CloudTrail/configuration history, deployment execution, and runtime health. CloudTrail may not identify old changes outside retention. Redact account IDs, ARNs, tags, parameter values, network details, and customer data.
Workshop and game day
Given a stack with 20 resources, build a coverage matrix listing drift support, explicit tracked properties, alternate detection, and owner. Analyze supplied findings for security-group rule, bucket policy, Auto Scaling desired capacity, deleted alarm, changed database retention, Lambda code, nested stack, and external IAM attachment.
For each choose revert drift, accept and update source, import, exception, retire, or investigate. Produce a drift-aware change-set review and a canary remediation sequence. Inject 14 cases: stale check, unsupported type, default omitted, emergency rule, malicious policy, AWS-managed scaling, replacement required, missing physical resource, nested drift, cross-stack attachment, provider failure, role denial, concurrent deployment, and remediation alarm. Prove final desired/actual agreement and user health.
Cost, security, and acceptance
Include detection/API activity, Config/security tooling, logs, retained evidence, replacement overlap, outage risk, and engineering time. This lesson creates nothing. Schedule detection by criticality and event, not merely a noisy universal interval; alert on aged unresolved drift and failed checks.
Submit coverage and ownership matrices, eight classifications, actor timeline, remediation decision record, drift-aware review, canary/rollback plan, fourteen failures, cost, and residual blind spots. Pass requires no claim that IN_SYNC means secure, no unmanaged overwrite, one owner per property, and verified post-remediation runtime health.