Lesson 317 · AWS Learning Path

AWS 317: Security pillar

· Published · 9 min read

Labelled process diagram for AWS 317: Workload, assets, and threats to Layered security review to Prioritized control improvement to Prevention, detection, response, and recovery evidence, with decision, proof and...

Why this lesson matters

The Security pillar protects data, systems, and assets while delivering business value through risk assessment and mitigation. Current AWS Well-Architected best-practice areas cover security foundations, identity and access management, detection, infrastructure protection, data protection, incident response, and application security.

Security is not one firewall, KMS key, or compliance report. It is a chain from governance and identity through software, network, data, telemetry, response, and recovery. A resource can be encrypted and still be exposed by a broad role. A private subnet can run vulnerable code. A passing scan can miss a stolen session.

This lesson treats every “secure” claim as a statement requiring scoped evidence, negative tests, and an owner.

What you will be able to do

By the end, you can:

  • apply the AWS shared-responsibility model to a workload;
  • build account, organization, Region, and security-governance foundations;
  • design workforce, workload, device, and emergency identities;
  • evaluate authentication, authorization, federation, session, and secret controls;
  • create a threat model and attack-path analysis;
  • design logging, detection, triage, and evidence retention;
  • protect network, compute, containers, APIs, and dependencies;
  • classify and protect data throughout its lifecycle;
  • integrate application-security and supply-chain controls;
  • plan and exercise incident containment, eradication, recovery, and learning; and
  • produce a defensible Security pillar review.

Before you start

  • Use one fictional or approved workload. T0 means no security-control changes.
  • Never include secrets, exploit details, personal data, account IDs, incident evidence, or internal topology in shared work.
  • Security testing needs authorization, scope, time window, safety limits, and response contacts.
  • Compliance evidence is not proof that current controls resist the workload's actual threats.

1. Establish security foundations

Foundations include accountable leadership, risk appetite, policies, control objectives, account/organization structure, inventory, ownership, standards, exception handling, threat intelligence, and continuous assurance.

business asset and threat
 -> control objective
 -> preventive/detective/responsive/recovery controls
 -> implementation evidence
 -> negative test
 -> monitoring and owner

Use multiple accounts to isolate production, non-production, security/log archive, networking, and workloads according to risk. Apply Organizations, SCPs, delegated administration, Control Tower or equivalent governance deliberately. SCPs limit maximum permissions; they do not grant access.

Restrict Regions when legally/operationally appropriate while preserving global-service and recovery dependencies. Maintain complete resource/software/data/identity inventories and ownership. Unknown assets are unmanaged attack surface.

Exceptions require risk, compensating controls, owner, approval, expiry, and review. Permanent “temporary” exceptions are control failures.

2. Apply shared responsibility precisely

AWS secures the cloud infrastructure; customers secure identities, data, configurations, workloads, and service-specific responsibilities. Responsibility changes by service:

  • EC2 customers patch guest OS and application;
  • managed databases shift engine/infrastructure tasks but leave users, network, schema, data, and configuration;
  • serverless removes server management but leaves code, dependencies, identity, input, and data;
  • SaaS-like services still require tenant/user/access/data/lifecycle controls.

For every component write “AWS owns,” “customer owns,” and “shared/dependent.” Never use “managed” to mean no security work.

3. Design identity and access management

Identity types:

IdentityPreferred pattern
WorkforceFederation through IAM Identity Center or approved IdP, MFA, short sessions
WorkloadIAM role and temporary credentials
AWS serviceService role/service-linked role with scoped trust
DevicePurpose-built device identity, not workforce keys
External tenant/customerApplication identity mapped to resource authorization
EmergencyProtected break-glass role, tested and alarmed

Authorization needs principal, action, resource, condition, session context, resource policy, permission boundary, SCP, and explicit deny evaluation. Least privilege is an evidence-driven process: start from required actions/resources/conditions, test denial, monitor use, and refine.

Trust policies control who can assume roles. Permission policies control what the session can do. Protect iam:PassRole, role chaining, external IDs, OIDC claims, session tags, and confused-deputy source conditions.

Eliminate long-lived access keys where possible. Rotate unavoidable secrets, store them in Secrets Manager/Parameter Store as appropriate, restrict retrieval, and monitor access. A secret copied into CI logs or container layers remains compromised after removal from source.

4. Build a threat model

Identify:

  1. business assets and unacceptable outcomes;
  2. users, administrators, workloads, vendors, and attackers;
  3. trust boundaries and data flows;
  4. entry points and external dependencies;
  5. threats such as spoofing, tampering, repudiation, disclosure, denial, privilege escalation, fraud, and supply-chain compromise;
  6. existing controls and detection;
  7. residual likelihood/impact; and
  8. treatment, owner, and verification.

Trace attack paths. Example:

phished developer -> source token -> pipeline role -> artifact tampering
 -> production workload role -> customer data

Controls must break or detect multiple steps: phishing-resistant MFA, scoped source token, protected branches, isolated runners, artifact signing, deployment approval, least-privilege runtime role, data-layer authorization, and anomaly detection.

5. Detect and investigate

Centralize CloudTrail organization trails, AWS Config, security findings, VPC/DNS/load-balancer/application logs, identity-provider events, endpoint/container signals, and data access logs according to risk. Protect log integrity with separate account, restrictive access, encryption, retention, and deletion controls.

Security Hub aggregates/normalizes findings; GuardDuty detects selected threats; Detective and query tools aid investigation; Macie discovers selected S3 sensitive data; Inspector assesses supported workloads. None covers every service or proves compromise.

For each detection define signal source, threat, severity, owner, false-positive handling, enrichment, runbook, response SLA, evidence retention, and test. Alert on disabling or tampering with security controls.

Time synchronization and request/user/session correlation are vital. Avoid logging credentials, authorization headers, secrets, full tokens, or unnecessary personal content.

6. Protect infrastructure

Layer controls:

  • VPC/subnet/route boundaries based on flows, not “public/private” labels;
  • security groups for stateful resource access and NACLs only where stateless subnet controls add value;
  • WAF, Shield, rate limits, bot/abuse controls at applicable edges;
  • Systems Manager instead of broad inbound administration;
  • patched/hardened immutable images, EDR where required;
  • container image scanning, non-root runtime, read-only filesystem, seccomp/capability and admission controls;
  • serverless input validation, dependency scanning, concurrency limits, scoped roles;
  • VPC endpoints/private access where requirements justify them;
  • egress allowlisting/filtering and DNS monitoring; and
  • segmentation that limits blast radius without blocking recovery.

Private network placement does not authenticate the caller. TLS does not authorize a request. WAF does not fix vulnerable business logic.

7. Protect data lifecycle

Classify data before selecting controls. Track create/collect, ingest, process, store, share, archive, restore, and delete.

For each class define allowed purpose, locations, identities, encryption, key ownership, tokenization/masking, logging, backup, retention, legal hold, residency, sharing, and deletion verification.

KMS controls key usage, not application authorization. Design key policy, grants, rotation, separation, multi-Region behavior, deletion protection, and recovery. Test loss/deny scenarios. A scheduled key deletion can make data unrecoverable.

Encrypt in transit and validate peer identity. Manage certificates through approved issuance, renewal, monitoring, revocation, and private-key protection.

Backups require isolation, immutability/lock where justified, malware/corruption-aware recovery, least privilege, restore testing, and separate deletion authority. Encryption of a backup does not prove restorability.

8. Secure applications and supply chain

Security begins with requirements and threat modeling. Apply secure coding, code review, SAST/DAST/SCA/secret scanning as suitable, dependency pinning, SBOM, artifact signing/verification, protected branches, isolated build roles, and progressive deployment.

Validate input at the trust boundary and encode output for its context. Use parameterized database access. Apply object/function-level authorization on every request. Protect sessions, CSRF, SSRF, deserialization, upload processing, and business workflows.

Do not automatically deploy every scanner finding or suppress it indefinitely. Triage exploitability and exposure, fix according to severity/SLA, verify, and manage exceptions.

9. Prepare incident response

Plan roles, legal/privacy/communications involvement, evidence handling, severity, authority, contacts, isolation methods, and alternate secure communications.

prepare -> detect/analyze -> contain -> eradicate
 -> recover -> validate -> communicate -> learn

Prebuild response roles, clean-room accounts, forensic storage, snapshots, quarantine groups, credential-revocation procedures, and tested automation. Keep response capability independent of the identity/network suspected compromised.

Containment can harm availability or evidence. Define when to isolate, revoke, block, snapshot, fail over, or keep observing. Record every action and hash/export evidence according to policy.

Recovery uses known-clean artifacts and restored trust, not merely restarting. Rotate affected credentials, patch root cause, validate data/integrity and business outcome, increase monitoring, then remove temporary controls carefully.

10. Read-only evidence

aws organizations describe-organization
aws cloudtrail describe-trails --output table
aws configservice describe-configuration-recorders --output table
aws securityhub describe-hub
aws guardduty list-detectors
aws accessanalyzer list-analyzers --output table
aws kms list-keys --limit 20 --output table
aws inspector2 batch-get-account-status --account-ids REDACTED

Resource existence is not configuration quality. Confirm scope, Region, delegated admin, coverage, status, destinations, findings workflow, retention, and tests.

11. Diagnose from evidence

SymptomEvidenceResponse
User denied unexpectedlyprincipal/session, policy evaluation, SCP/boundary/resource policy, conditionsCorrect exact layer; do not add administrator access.
Role usable by wrong accounttrust policy, external ID/source conditions, CloudTrail assume eventsRestrict trust, revoke sessions if needed, investigate.
Public-data alert but bucket policy looks privateACL/access point/object ownership/CDN/presigned pathTrace every access path and contain exposure.
Findings exist for monthsownership, severity/SLA, ticket state, exception expiryAssign accountable treatment and verify closure.
Logs absent during incidenttrail/log source scope, permissions, tampering, retentionPreserve remaining evidence and repair independent logging.
KMS decrypt failskey policy, IAM, grant, encryption context, key/Region/stateRepair least-privilege path; never replace key blindly.
Patch caused outageartifact/tests/deployment/rollback evidenceRoll back safely, correct compatibility, maintain risk visibility.

12. Security review workshop

Review a fictional internet API handling customer documents. Submit:

  1. assets/classification/unacceptable outcomes;
  2. shared-responsibility matrix;
  3. account/organization/Region foundations;
  4. workforce/workload/emergency identity map;
  5. authorization and trust-policy tests;
  6. threat model and five attack paths;
  7. detection/logging coverage;
  8. network/compute/container/API controls;
  9. data/key/certificate lifecycle;
  10. application/supply-chain controls;
  11. backup/restore/ransomware boundaries;
  12. incident roles and decision authority;
  13. three runbooks and a game day;
  14. seven diagnostic cases;
  15. risk-ranked improvement plan; and
  16. evidence-backed Well-Architected answers.

Cost and cleanup

Security costs include logging/storage/query, findings services, WAF/Shield/firewalls, KMS/Secrets, scanning, endpoint tooling, response accounts, backups, testing, staff, and remediation. Optimize based on risk and signal value, not by disabling evidence.

T0 creates nothing. Any later test must remove test identities, findings, snapshots, logs, rules, endpoints, keys/grants, and synthetic data only under approved retention and exact ownership.

Knowledge check

  1. Seven areas? Foundations, IAM, detection, infrastructure, data, incident response, application security.
  2. SCP effect? Limits maximum permission; it grants nothing.
  3. Trust versus permission policy? Who may assume versus what the session may do.
  4. Why is private subnet insufficient? Identity, vulnerable code, egress, and data authorization remain.
  5. What does KMS not solve? Who the application should allow to access plaintext.
  6. Why centralize protected logs? Compromise of a workload should not erase its evidence.
  7. Why test restore? A backup object does not prove recoverable trustworthy service.
  8. What proves a control? Scoped configuration plus negative test, monitoring, ownership, and response.

Lesson acceptance

Pass when every material threat maps to preventive, detective, response, and recovery evidence with owners and negative tests. Fail if encryption/firewalls/compliance badges substitute for authorization, broad roles remain, logging can be erased by the workload, sensitive-data lifecycle is unknown, supply chain is untrusted, or incident recovery cannot restore clean trust.

Official sources

Advertisement