AWS 221: Change sets, nested stacks and StackSets
Why this lesson matters
These three features solve different scale problems. A change set previews one stack operation. A nested stack divides one application hierarchy into owned components. A StackSet orchestrates related stack instances across accounts and Regions. Treating them as interchangeable causes unreviewed replacement, tightly coupled templates, or an organization-wide blast radius.
What you will be able to do
By the end, you can:
- read change-set Add/Modify/Remove, replacement and scope evidence;
- explain what a change set cannot predict;
- design nested-stack inputs, outputs, template storage, update and failure ownership;
- distinguish StackSet, stack instance, operation and per-target stack;
- choose self-managed or Organizations service-managed permissions;
- calculate account concurrency, Regional order and failure tolerance;
- explain strict versus soft failure-tolerance concurrency;
- plan canary OUs/accounts/Regions, drift review, rollback and retained-stack ownership.
Three separate control planes
| Feature | Scope | Primary purpose | Execution unit |
|---|---|---|---|
| Change set | one stack, optionally its nested hierarchy | preview a proposed create/update/import | resource change |
| Nested stack | one parent hierarchy in one account/Region | component reuse and lifecycle composition | child AWS::CloudFormation::Stack |
| StackSet | one StackSet home Region to target account/Region pairs | governed fleet deployment | stack instance operation |
A StackSet is Regional, even though it targets many Regions. A stack instance is the StackSet's reference to a target stack in one account and one Region. Its status is not the same thing as the target stack's resource health.
Change sets: a preview, not a guarantee
A change set compares submitted template/parameters with current stack state and describes proposed resource actions:
Add,Modify,Remove,Import, or dynamic changes;- logical/physical resource identity;
- changed properties and evaluation mode;
- replacement value
True,False, orConditional; - scope such as properties, metadata, tags or creation policy;
- nested child change sets when inclusion was requested.
Conditional means CloudFormation can determine replacement only during execution after runtime values resolve. Even False does not prove no interruption at the application layer. A change set does not prove target API permissions, quotas, unique names, hooks/runtime success, custom-resource behavior, data migration, service health, or that a dynamic reference still resolves.
Never execute from a screenshot. Record change-set ID, stack ID, creator/caller, creation time, status, execution status, template hash, parameters, capabilities and every replacement/removal. Recreate stale change sets after the template or environment changes, and delete abandoned ones.
P11 change-review exercise
Without deploying, make three local copies of P11 and predict the change set:
| Proposed edit | Expected action if stack existed | Review concern |
|---|---|---|
LogRetentionDays: 7 -> 14 | modify log group in place | extra retained storage cost |
CreateIngestionAlarm: false -> true | add alarm | charge and missing-data behavior |
| change bucket physical name expression | replace bucket | old bucket retained; new global name; application cutover |
remove ApplicationLogGroup | remove from stack, physical group retained | unmanaged retained data/cost |
| disable public-access block | modify | security rejection even if technically executable |
The review decision is not merely “change set succeeded.” It is approve, reject, or require migration/backup/rollback evidence.
Read an existing change set safely
Choose an instructor-supplied or owned stack/change set. These commands do not create or execute one:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
owned_stack="exact-approved-stack-name"
aws cloudformation list-change-sets --stack-name "$owned_stack" --output table
owned_change_set="exact-approved-change-set-name-or-id"
aws cloudformation describe-change-set \
--stack-name "$owned_stack" --change-set-name "$owned_change_set" \
--include-property-values --output json
If the CLI version does not support --include-property-values, omit it and record the tooling difference. No change set is an acceptable no-create result; analyze the P11 predictions instead.
Nested stacks
A parent declares a child with AWS::CloudFormation::Stack and a reachable TemplateURL. The parent passes parameters; it reads child outputs using Fn::GetAtt ChildStack.Outputs.OutputName. This creates lifecycle coupling: parent create/update/delete operations orchestrate the child.
Use nesting when a component:
- belongs to the same application lifecycle;
- has a clear owner and bounded input/output contract;
- is reused enough to justify versioned template packaging;
- stays within depth, resource and operation quotas.
Do not use deep nesting to hide complexity. A separately operated shared network, security platform or database may need its own stack/interface because its owner and release cadence differ. Store child templates in an approved versioned artifact location, pin the intended object/version through the delivery process, and review artifact integrity.
For CLI-created change sets, nested inclusion is not automatic: use --include-nested-stacks when creating the root change set. Execute and delete from the root. Review every generated child change set. Updating a nested stack directly can desynchronize ownership; normally update through the root.
StackSet permission models
| Model | Administration | Target authorization | Appropriate when |
|---|---|---|---|
| Self-managed | administrator account | explicitly configured administration/execution roles | accounts are outside one Organization or custom trust is required |
| Service-managed | Organizations management account or delegated administrator | trusted access creates/manages required access | targets are OUs/accounts in one AWS Organization |
Service-managed StackSets can automatically deploy to accounts added to targeted OUs and remove stack instances when accounts leave, according to automatic-deployment and retain/removal choices. A delegated administrator can have broad StackSet reach; AWS notes that this cannot be restricted to selected OUs through the delegation itself, so separation of duties and preventive controls matter.
Never put a Region-unique/global resource with the same fixed name into every target Region. IAM, S3 names, Route 53 and other scopes need explicit conflict analysis.
Operation preferences and blast radius
Define:
- target OUs/accounts and Regions;
- Region order and sequential versus parallel Region concurrency;
- failure tolerance count/percentage per Region;
- maximum concurrent accounts count/percentage;
- strict or soft failure-tolerance concurrency;
- managed execution, operation queueing, canary wave and stop/rollback authority.
With strict failure tolerance, actual account concurrency starts no higher than the lower of maximum concurrency and failure tolerance + 1, then reduces as failures occur. With soft failure tolerance, configured maximum concurrency is maintained despite failures already observed, so more targets can be affected before queued work stops. Percentage values round down to whole accounts.
For a first wave, choose one non-production account, one Region, maximum concurrency 1, failure tolerance 0, sequential Regions, and strict mode. Promote only after target-stack resources, application behavior, logs, cost and drift are accepted.
Read-only StackSet inspection
aws cloudformation list-stack-sets --status ACTIVE \
--query 'Summaries[].{Name:StackSetName,Model:PermissionModel,Status:Status,Auto:AutoDeployment}' \
--output table
owned_stack_set="exact-approved-stack-set-name"
aws cloudformation describe-stack-set --stack-set-name "$owned_stack_set" \
--query 'StackSet.{Name:StackSetName,Model:PermissionModel,Status:Status,ManagedExecution:ManagedExecution,AutoDeployment:AutoDeployment,Drift:DriftStatus}' \
--output json
aws cloudformation list-stack-instances --stack-set-name "$owned_stack_set" \
--query 'Summaries[].{Account:Account,Region:Region,Status:Status,Detailed:StackInstanceStatus.DetailedStatus,Drift:DriftStatus}' \
--output table
aws cloudformation list-stack-set-operations --stack-set-name "$owned_stack_set" \
--max-results 20 --output table
When using a service-managed StackSet from a delegated administrator, commands require the appropriate --call-as DELEGATED_ADMIN; do not add it when acting as self/management account by assumption.
Central logging design exercise
Design - but do not deploy - a StackSet that configures account-local log destinations:
- security/platform team owns the template and StackSet home Region;
- target is one canary OU before wider OUs;
- each target Region uses a Regional destination that forwards to the approved central boundary;
- names include account/Region where service scope requires uniqueness;
- parameters separate organization-wide invariants from target overrides;
- operation starts strict, one account, zero tolerated failures, sequential Regions;
- acceptance proves target stack, destination policy/KMS permissions, real fake event delivery, cost tags and drift;
- rollback defines whether stack instances/resources are deleted or retained and who owns retained stacks.
“Operation SUCCEEDED” is not end-to-end log-delivery evidence.
Diagnose failures
| Symptom | Prove first | Correction |
|---|---|---|
| Change set has no changes | submitted template hash/parameters and current stack | accept no-op or submit intended artifact; do not execute stale set |
Replacement is Conditional | changed property and runtime references | stop until migration/retention evidence exists |
| Nested preview misses children | root change set's include-nested setting | recreate root preview with nested inclusion |
| Child update failed | root and child earliest failed events | repair child cause through parent workflow |
| StackSet operation failed | per-Region/per-account results and first target error | stop promotion; repair canary identity/quota/conflict |
| Many targets fail quickly | soft/high concurrency settings | reduce blast radius for next reviewed operation |
| Stack instance says OUTDATED | latest operation and parameter overrides | reconcile deliberately; do not delete/recreate blindly |
| Removed instance still runs | retain-stacks choice | transfer owner and manage target stack directly |
Cost and cleanup
Change-set storage/preview is not the main bill; target resources, S3 template artifacts, logs, KMS/API use and retained resources are. Fleet deployment multiplies cost by account × Region. Nested stacks can also retain child data after parent deletion.
This lesson creates nothing. Do not execute or delete an existing change set, child stack, StackSet operation or stack instance. Remove local P11 exercise copies if unneeded. Document the exact owner of every hypothetical retained target stack.
Knowledge check
- Does a change set guarantee an update succeeds?
No; permissions, quotas, runtime values and service behavior are evaluated during execution.
- How do CLI users include nested changes?
Create the root change set with --include-nested-stacks.
- What is a stack instance?
A StackSet's reference to one target stack in one account and Region.
- How does strict concurrency constrain blast radius?
It ties actual concurrency to failure tolerance + 1 and reduces it after failures.
- What happens when stacks are retained during StackSet removal?
They detach and remain managed directly in target accounts/Regions.
Lesson acceptance
- Five P11 changes are classified by action, replacement, security and data-retention risk.
- Change-set limitations and required approval evidence are explicit.
- Nested-stack ownership, artifact, input/output, update and failure boundaries are designed.
- Self-managed/service-managed permission models and delegated-admin risk are distinguished.
- Canary targets, Region order, strict/soft concurrency, maximum concurrency and failure tolerance are calculated.
- Read-only stack/change-set/StackSet evidence is captured or a documented no-resource result is supplied.