AWS 372: CloudFormation stack policies, termination protection, failure options, and rollback
Why this lesson matters
CloudFormation coordinates resource state, but different controls protect different operations. Stack policy, termination protection, deletion/update-replace policies, failure behavior, rollback triggers, and service-level backups are not interchangeable.
Protection control map
| Control | Protects against | Does not protect against |
|---|---|---|
| Stack policy | Unauthorized updates to selected logical resources | Stack deletion or direct service changes |
| Termination protection | Deleting the stack through CloudFormation | Resource replacement during update |
DeletionPolicy | Resource disposition when removed/deleted | Every replacement event |
UpdateReplacePolicy | Old physical resource during replacement | Data consistency or ongoing cost |
| Failure option | Automatic rollback behavior | Logical design errors and immutable updates |
| Rollback trigger | Alarm during operation/monitoring window | Unmonitored business failure |
| Service backup/replication | Recoverable data | Correct template or application behavior |
A stack policy defaults to allowing updates unless statements deny them. Once set, it cannot be deleted, only replaced. Temporary update-policy overrides apply only to that update, so review their exact target and restore protection. Logical IDs and resource types must match intended resources.
Termination protection is a stack property. For nested stacks it is controlled at the root; nested stacks are not independently deleted. It prevents delete-stack, not updates that replace or remove a resource.
Failure and rollback states
Default create failure normally rolls back created resources. Preserving successfully provisioned resources leaves successful resources and failed state for diagnosis and later retry/update/rollback, but has conditions and unsupported immutable-update cases. It can increase cost and leave partial architecture exposed, so pre-authorize ownership and cleanup.
Read stack events from oldest causal failure forward while accounting for parallel dependencies. UPDATE_ROLLBACK_FAILED means rollback could not complete, often because a resource changed or permission/capacity is missing. Repair the cause and continue rollback; skip resources only with documented consequence because template and reality may diverge.
Rollback triggers monitor specified CloudWatch alarms during the update and monitoring period. Missing/deleted alarms or ALARM state can cause rollback. Define missing-data treatment, dimensions, baseline health, and business signal. CloudFormation rollback restores managed resource configuration where possible, not external writes or data semantics.
Before a risky update, capture the current template, parameters without exposing secrets, stack policy, resource inventory, physical identifiers, change set, role, quotas, alarm baseline, backup recovery point, and tested rollback authority. During recovery, do not delete a failed stack merely to remove an error status. Determine which resources succeeded, which rolled back, which were retained, and whether callers already wrote data. Reconcile the final physical state to one approved template and document every retained exception.
Data retention semantics
DeletionPolicy: Retain leaves the physical resource when its template resource is deleted; CloudFormation stops managing it and billing continues. Snapshot applies only to supported resources. RetainExceptOnCreate can clean a newly created unused resource if the creating operation rolls back while retaining it in later removal cases. Verify current service support.
UpdateReplacePolicy decides what happens to the old physical resource after replacement. Retaining an old database or volume is not a backup strategy without catalog, encryption/key retention, tested restore, access control, expiry, and owner. Secrets, snapshots, buckets, and databases need explicit lifecycle decisions.
Read-only inspection and workshop
aws cloudformation describe-stacks --stack-name STACK --region ap-south-1
aws cloudformation get-stack-policy --stack-name STACK --region ap-south-1
aws cloudformation describe-stack-events --stack-name STACK --region ap-south-1
aws cloudformation get-template --stack-name STACK --template-stage Original --region ap-south-1
aws cloudformation detect-stack-drift --stack-name STACK --region ap-south-1
Create a local two-tier template with database, storage, compute, alarms, and nested boundary. Add least-permissive stack-policy targets, termination protection runbook, deletion/update-replace policy for every stateful resource, failure option, rollback trigger, backup, and decommission plan. Draw create, update, replace, remove-from-template, stack-delete, failed-create, failed-update, and failed-rollback outcomes.
Diagnose 12 scenarios: stack-policy deny, overly broad override, termination assumption, replacement surprise, retained-resource orphan, snapshot KMS loss, alarm already red, alarm missing, preserve-partial exposure, quota failure, rollback role denial, and external state change. Record event, physical/logical identity, data risk, cost owner, repair, and final reconciliation.
Cost, security, and acceptance
CloudFormation itself may have no direct stack fee for ordinary resource types, but managed resources, retained data, snapshots, logs, alarms, custom resources, third-party types, and failed partial deployments cost money. Use a service role with least privilege, protect template/parameters, use dynamic references appropriately, and never expose NoEcho values through outputs or logs.
Submit the control matrix, annotated template, eight-operation disposition table, event diagnosis, policy/override review, alarm semantics, backup/restore evidence plan, twelve failures, and cost/decommission ownership. Pass requires no claim that one control substitutes for another and explicit recovery for both infrastructure and data.