AWS 374: StackSets administration, delegated administration, concurrency, and failure tolerance
Why this lesson matters
StackSets applies one template across accounts and Regions. That multiplication makes identity, organizational scope, parameters, Region order, concurrency, failure tolerance, auto-deployment, drift, and account-removal policy architecture decisions rather than checkboxes.
Permission and scope models
| Model | Identity path | Best fit | Main risk |
|---|---|---|---|
| Self-managed | Administrator role assumes execution role in each target | Accounts outside one Organization or custom role control | Role bootstrap/trust drift |
| Service-managed | Organizations trusted access creates/manages execution roles | OU/account targeting in one Organization | Broad organizational authority |
| Delegated administrator | Member account operates service-managed StackSets | Separation from management account | Delegate can deploy organization-wide; OU restriction is not a safety boundary |
A StackSet is Regional control-plane state even though its stack instances span Regions. Record administration account, home Region, permission model, execution role, target OUs/accounts, account filters, deployment Regions, parameters, and template version. Service-managed StackSets cannot target outside the Organization and have feature restrictions such as unsupported nested stacks/macros in documented contexts.
Trusted access and delegated administration are high-impact organizational changes. Use dedicated roles, change approval, CloudTrail monitoring, SCP guardrails where suitable, and emergency revocation. When acting as a delegate, API/CLI operations require the delegated call mode.
Operation mathematics
Failure tolerance is the number or percentage of failed accounts allowed per Region before the operation stops. Percentage values round down. Maximum concurrent accounts also resolves by count or percentage. Under strict failure tolerance, actual concurrency is constrained relative to tolerance and decreases after failures. Soft mode maintains requested concurrency despite failures, increasing speed and potential blast radius.
Example: 53 accounts, maximum concurrency 20 percent, and failure tolerance 5 percent. Calculate documented rounding before release, then choose a canary OU or small account list first. Region concurrency may be sequential or parallel; sequential respects Region order and stops later Regions when tolerance is exceeded. Put low-risk/canary Regions first only when that ordering reflects real dependencies and data residency.
| Rollout phase | Accounts/Regions | Preference | Gate |
|---|---|---|---|
| Laboratory | 1 nonproduction account, 1 Region | concurrency 1, tolerance 0 | Template and runtime tests |
| Canary | Representative small OU | strict, low concurrency | Drift, security, business health |
| Wave | Bounded OUs and sequential Regions | measured tolerance | Error budget and operator approval |
| Broad | Remaining approved scope | justified mode | Continued alarms and audit |
Auto-deployment, removal, and drift
Service-managed auto-deployment can add stack instances when accounts enter targeted OUs and remove them when accounts leave. RetainStacksOnAccountRemoval decides whether stacks/resources stay but leave StackSet management. Retention can create unmanaged cost/security drift; deletion can destroy needed state. Classify every resource and account lifecycle first.
Moving an account between OUs can queue delete/create operations. Organizational events, operation queueing, suspended/closed accounts, opt-in Regions, quota, and role propagation all affect timing. StackSet dependencies may order eligible automatic deployments, but do not replace application dependency analysis.
Drift detection is an explicit operation, not continuous assurance. A StackSet summary can be NOT_CHECKED; inspect each stack instance and resource result. Drift remediation through update may overwrite an emergency fix. Record why drift exists, owner, security effect, approved desired state, and change window.
Read-only inspection
aws cloudformation list-stack-sets --status ACTIVE --region ap-south-1
aws cloudformation describe-stack-set --stack-set-name STACK_SET --region ap-south-1
aws cloudformation list-stack-instances --stack-set-name STACK_SET --region ap-south-1
aws cloudformation list-stack-set-operations --stack-set-name STACK_SET --region ap-south-1
aws cloudformation list-stack-set-operation-results --stack-set-name STACK_SET --operation-id OPERATION --region ap-south-1
For delegated administration add --call-as DELEGATED_ADMIN where required. Correlate operation, account, Region, status reason, OU/filter target, parameters, template, role, drift status, and CloudTrail actor. Redact organization/account IDs, ARNs, email, tags, and parameters.
Workshop and failure game day
Design a baseline rollout to 120 accounts in four OUs and three Regions. Exclude a regulated account set, define parameters by environment, calculate count/percentage rounding, sequence laboratory/canary/waves, choose strict or soft mode, state stop/continue rules, and define account join/move/leave behavior.
Test: wrong home Region, delegate lacks registration, target filter includes wrong accounts, execution role drift, opt-in Region disabled, quota, parameter mismatch, template replacement, percentage rounding surprise, strict concurrency slowdown, soft-mode broad failure, first-Region tolerance exceeded, account move queue, retained orphan, delete with state, drift overwrite, suspended account, and operation conflict. For each preserve operation and instance evidence and define remediation without rerunning successful instances blindly.
Cost and acceptance
Price resources multiplied by accounts/Regions, CloudFormation third-party handlers if used, logs, drift operations, Config/security monitoring, cross-Region transfer, retained stacks, and failed partial deployment. This no-create lesson has no cleanup.
Submit permission/trust map, target inventory, rollout math, four-phase plan, parameter governance, auto-deployment/removal table, drift procedure, 18 failures, monitoring, cost forecast, rollback/forward-fix plan, and decommission. Pass requires bounded canaries, verified rounding, least privilege, explicit retained-state ownership, and account/Region-level evidence.