AWS 382: IaC linting, policy tests, security scans, change evidence, and promotion gates
Why this lesson matters
No single scanner proves infrastructure safe. Reliable promotion layers parser/schema validation, linting, tests, policy, security and secret checks, dependency/artifact review, change evidence, cost/quota analysis, runtime tests, approvals, and expiring exceptions. Each gate needs scope, owner, severity, evidence, and failure behavior.
Layered gate model
| Gate | Finds | Cannot prove alone |
|---|---|---|
| Parse/schema/transform | Invalid syntax, type/property errors | Architecture intent or runtime behavior |
| Linter | Provider-specific mistakes and quality issues | Organizational policy completeness |
| Unit/assertion | Expected generated resources/relationships | Real service integration |
| Policy as code | Deterministic organization rules | Unknown threats or business correctness |
| Security scan | Known risky configurations | Exploitability and all generated/live state |
| Secret/dependency scan | Credentials and vulnerable dependencies | Rotation or runtime exposure eliminated |
| Diff/change set | Proposed additions/modifications/removals | Runtime success and all drift effects |
| Cost/quota/capacity | Financial and scale risk | Actual demand/user outcome |
| Integration/resilience | Behavior in an environment | Every production failure mode |
Run cheap deterministic gates early and environment-dependent gates later. Validate source and processed/generated templates because transforms, CDK/SAM/modules/macros can add resources or IAM. Pin tool/rule versions and capture configuration plus database timestamp so results are reproducible.
Policy, findings, and waivers
Rules need stable identifiers, rationale, severity, scope, examples, remediation, owner, and tests for pass/fail/boundary cases. Start critical rules in report mode only when required to measure impact, then enforce through a dated rollout. Avoid simplistic rules that force insecure workarounds.
Normalize findings from tools without erasing source evidence. Deduplicate by resource/rule/context, classify true/false positive, and set severity plus deadline from risk and environment. A waiver contains rule/finding, exact resource/scope, business reason, compensating control, accountable approver, creation/expiry, review trigger, and evidence. Expired waivers fail closed or require explicit renewal; permanent suppression without owner is hidden policy failure.
Protect gate configuration, runner images, plugins/rules, baselines, artifact storage, and credentials as supply-chain assets. Pull-request code is untrusted: do not expose production secrets or privileged roles to arbitrary forks. Sign or hash reports and bind them to commit, generated template, artifact, pipeline execution, and target environment.
Change evidence and promotion
A promotion dossier should include source review, commit/signature, dependency lock, tool versions, source/generated template hashes, test reports, policy/security findings, waivers, secret scan, SBOM/provenance, change set with replacements/deletions/IAM, drift, cost delta, quota/capacity, backup/rollback readiness, runtime tests, approvals, deployment identity, alarms, and cleanup.
Retain the dossier according to audit and incident needs in an access-controlled, encrypted store with immutable versioning or equivalent tamper evidence. Define who can write, approve, read, expire, and delete evidence. Reports can contain resource names, paths, vulnerabilities, and architecture details, so redact exports and never make a build dashboard public. Test that an auditor can reconstruct one historical release after tools and environments have changed.
Use explicit thresholds: which severity blocks; whether missing/unreadable output blocks; maximum report age; baseline policy; replacement approval; allowed cost increase; quota margin; and post-deployment SLI. Scanner crash, timeout, malformed report, or skipped target must not be interpreted as a pass.
Promotion across environments should use the same immutable source/artifact and versioned configuration. Re-run environment-specific policy and change evidence because accounts, Regions, quotas, live state, and IAM differ. An approval applies to one execution and evidence set, not every future retry.
Local workshop
Create a deliberately flawed CloudFormation or SAM template, then build a local gate script/CI design that runs YAML/JSON parse, cfn-lint, aws cloudformation validate-template where approved, unit assertions, CloudFormation Guard rules/tests, secret scan, dependency/SBOM scan, generated-template security scan, and change-set review. Use pinned containers/packages and produce machine-readable plus human reports.
cfn-lint template.yaml
cfn-guard validate --rules rules.guard --data template.yaml
cfn-guard test --rules-file rules.guard --test-data tests.yaml
aws cloudformation validate-template --template-body file://template.yaml
Detect public storage, unencrypted data, wildcard IAM, missing logs, unrestricted ingress, missing backup, mutable image, unsafe deletion, and absent ownership tags. Add justified exception examples and verify expiry enforcement.
Failure game day and acceptance
Inject 20 cases: malformed YAML, unsupported property, transform hides IAM, linter warning ignored, unit snapshot stale, policy rule bug, scanner database stale, secret in Git history, dependency confusion, malicious runner image, fork obtains credential, report missing, parser treats error as empty pass, finding dedup hides critical context, baseline grows, waiver too broad, waiver expired, change set replacement overlooked, cost/quota omitted, and production drift invalidates approval.
Price CI compute, licensed scanners, artifact/report storage, KMS, test environments, logs, change-set resources, security operations, false-positive work, and delayed delivery. Submit gate architecture, rule catalog/tests, threat model, normalized finding format, waiver workflow, complete dossier, 20 failures, cost, and metrics. Pass requires generated-output scanning, fail-closed evidence, immutable binding, expiring exceptions, environment-specific review, and measured runtime outcome.