AWS 417: SDLC automation and configuration management
Purpose and exam boundary
This checkpoint assesses DOP-C02 Domain 1 SDLC Automation and Domain 2 Configuration Management and IaC. It does not reteach service menus. You must select controls from requirements, follow source to runtime identity, reason through failure/rollback and reject plausible distractors using evidence.
The certification guide weights SDLC Automation at 22 percent and Configuration Management/IaC at 17 percent of scored content. Weighting guides study effort but does not replace practical competence. Services and exact features evolve; reasoning should survive a product-name change.
Scenario packet
A company deploys a stateful API to three accounts and two Regions. Teams rebuild artifacts per environment, share deployment roles, manually edit stacks during incidents, use mutable image tags and cannot explain drift. Releases must support hourly nonproduction delivery, weekly production windows, separation of duties, no downtime for compatible changes, measured rollback and complete evidence.
| Evidence supplied | Learner must determine |
|---|---|
| Repository/workflow and branch rules | Trigger, review and provenance weaknesses |
| Build logs, SBOMs and image manifests | Whether bytes are reproducible and identical |
| Pipeline/role policies | Stage isolation, cross-account trust and escalation |
| CloudFormation/CDK outputs | Change set, replacement and dependency impact |
| Deployment histories/alarms | Strategy, bake, rollback and data compatibility |
| Config/drift/emergency records | Desired-state ownership and safe reconciliation |
Assessment tasks
- Draw the source, build, artifact, promotion, deployment and runtime identity chain. Mark every trust boundary and mutable reference.
- Redesign roles for source, build, signing, artifact, pipeline, deployment, stack execution and runtime. Prove constrained trust and PassRole.
- Select build/test/security stages and their order. Explain what blocks, warns or requires an expiring exception.
- Design one immutable artifact promoted across accounts/Regions, including digest, signature, SBOM, provenance, KMS and retention.
- Compare all-at-once, rolling, blue/green and canary for this workload. Calculate overlap capacity and choose alarms, bake and rollback.
- Review a supplied change set containing update, replacement and deletion. Identify state loss, ordering, policy and recovery risks.
- Decide CloudFormation nesting/module/CDK boundaries, deterministic context/assets and bootstrap trust. Explain generated-template inspection.
- Classify drift by owner and risk. Choose reconcile source, import/refactor, revert emergency change, replace or accept temporary exception.
- Design multi-account/Region StackSet or pipeline rollout with failure tolerance, concurrency, canaries and evidence.
- Produce rollback and roll-forward decisions for application, configuration, database schema and infrastructure changes.
Quantitative and troubleshooting section
Calculate canary traffic and expected error counts, deployment capacity with minimum healthy percentage/maximum percentage, StackSet failure tolerance/concurrency, artifact storage/transfer and rollback RTO. State assumptions and reject mathematically valid plans that violate quotas, availability or independence.
Diagnose eight incidents from supplied evidence: source revision mismatch, cache poisoning/stale dependency, artifact KMS denial, cross-account AssumeRole failure, CloudFormation rollback stuck on retained resource, CDK asset/bootstrap mismatch, StackSet partial rollout and runtime alarm after apparently successful deployment. For each identify first failed control, strongest evidence, safe next action, prevention and rollback limit.
Architecture decisions
Write six short ADRs: CI service/runner isolation, artifact store and promotion, deployment strategy, IaC composition tool, organization rollout mechanism and drift/emergency ownership. Each includes requirements, assumptions, at least two viable alternatives, selection, rejection reasons, security/availability/cost, failure behavior and reversal.
Do not select a service because it is named in a question. For example, CodePipeline can orchestrate but does not make a build reproducible; CloudFormation can converge declared resources but does not own every runtime/data change; CDK improves abstraction but generated CloudFormation still requires review; immutable infrastructure reduces drift but does not remove state migration.
Practical evidence dossier
Using supplied execution artifacts, create a release dossier with commit/review, build environment, tests, digest/signature/SBOM, scan/policy results, promotion proof, change set, approval, deployment events, runtime digest, alarms/SLIs, rollback readiness and CloudTrail IDs. Identify any gap that prevents release.
Then modify the design for three constraints: pipeline control Region unavailable, artifact destination key disabled and database migration not backward compatible. Explain degraded operation, stop condition, recovery and evidence. Do not claim multi-Region simply because artifacts are replicated.
Scoring and remediation
| Area | Points | Automatic failure condition |
|---|---|---|
| Identity/provenance and roles | 20 | Static key or unbounded production role |
| Pipeline gates and evidence | 20 | Mutable artifact or bypassed failure |
| Deployment/recovery | 20 | No user-level verification/rollback analysis |
| IaC/change/drift | 20 | Destructive change not identified |
| Troubleshooting/decisions | 20 | Answers unsupported by supplied evidence |
Pass at 80/100 with no automatic failure and at least 60 percent in each area. For every miss, cite the earlier lesson, restate the rule, solve a variant and add evidence to the learner portfolio. Memorized definitions without a defensible design do not pass.
Submission acceptance
Submit diagrams, ten task answers, calculations, eight incident reports, six ADRs, release dossier, three constraint variants and a personal remediation map. Acceptance requires exact artifact identity, policy-layer reasoning, safe data-aware recovery, quantitative rollout and evidence-backed rejection of alternatives.