Lesson 371 · AWS Learning Path

AWS 371: Troubleshoot source, build, artifact, permission, deploy, and rollback failures

· Published · 4 min read

Labelled process diagram for AWS 371: Versioned intent to Automated validation to Controlled AWS change to Observed result and retained evidence, with decision, proof and rejection evidence.

Why this lesson matters

A pipeline failure message usually reports where orchestration stopped, not the root cause. This lesson diagnoses source, trigger, build, test, artifact, KMS, IAM, network, approval, deployment, hook, health, rollback, concurrency, and quota failures from time-ordered evidence.

The diagnostic method

Freeze retries first. Record account, Region, pipeline version, execution ID/mode, source revision, artifact digest, action/deployment/build IDs, target release, first failing timestamp, and customer effect. Preserve logs before retention removes them. Build a timeline in UTC and start with the earliest causal failure, not the loudest downstream error.

LayerEvidenceFrequent cause
Trigger/sourceEvent, connection, source revisionFilter, branch, token, duplicate event
OrchestrationExecution/action stateSuperseded/queued limit, wrong pipeline version
Build/testBuild phases and reportsDependency, environment, test, timeout
ArtifactName, digest, S3 version, metadataMissing file, mutation, wrong output mapping
AuthorizationCloudTrail denial and policy contextTrust, role, KMS key, bucket, SCP, boundary
NetworkDNS, route, endpoint, flow evidencePrivate build cannot reach dependency
DeploymentTarget events/hooks/healthCapacity, bad config, failed assertion
RecoveryPrevious identity and user SLIMissing artifact or incompatible state

Do not begin by adding administrator access. For AccessDenied, identify caller, API, resource, action, Region, session policy, identity policy, resource policy, permissions boundary, SCP, endpoint policy, KMS key policy/grant, and relevant conditions. A same-account IAM allow cannot overcome an explicit deny; cross-account access normally needs both sides.

Execution and artifact reasoning

Pipeline execution modes change causality. Superseded executions can stop older work, queued executions wait, and parallel executions can overlap. Match each action to its execution and pipeline definition version. A manual retry may consume a newer artifact or environment unless identity is pinned.

Trace artifacts as a graph: output namespace/name from one action must equal the next input. Compare source commit, S3 object version, ECR image digest, ZIP contents, manifest, task/AppSpec filename, KMS key, and target release. A successful upload proves storage, not completeness or authenticity.

For CodeBuild, read phase contexts from download through finalization. Distinguish command exit, report failure, out-of-memory, timeout, Docker privilege, VPC DNS/route/NAT/endpoint, quota, and artifact upload. Secret retrieval and artifact upload can fail after compilation succeeds.

Deployment and rollback

Follow target-specific evidence: Lambda alias/version and alarm; ECS service/task/target/hook; EC2 deployment instance lifecycle; Auto Scaling refresh; EKS controller/Pod/event. Verify health through the intended traffic path. A hook may time out after completing an external write, so handlers need idempotency.

Rollback can itself fail because the previous artifact was deleted, KMS or IAM changed, old capacity is unavailable, or data is no longer compatible. Decide among stop, retry, rollback, and forward fix from customer exposure and state compatibility. Never retry a non-idempotent action blindly.

Read-only evidence commands

aws codepipeline get-pipeline-state --name PIPELINE --region ap-south-1
aws codepipeline list-action-executions --pipeline-name PIPELINE --filter pipelineExecutionId=EXECUTION --region ap-south-1
aws codebuild batch-get-builds --ids BUILD_ID --region ap-south-1
aws deploy get-deployment --deployment-id DEPLOYMENT_ID --region ap-south-1
aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=Decrypt --region ap-south-1
aws service-quotas list-service-quotas --service-code codebuild --region ap-south-1

Capture exact reason fields and timestamps. Redact account IDs, ARNs, repository/artifact locations, tokens, parameters, environment values, logs, and customer data.

Failure lab

Given six supplied evidence packs, diagnose: trigger filter excludes commit; CodeBuild role lacks KMS decrypt; build in private subnet lacks egress; output artifact omits AppSpec; deployment hook reports success for the wrong endpoint; rollback cannot read an old artifact. For each produce a timeline, first causal event, ruled-out alternatives, smallest safe repair, retry safety decision, customer effect, prevention, and verification query.

Then classify 20 additional symptoms including queued limit, superseded execution, webhook duplication, dependency checksum, test-report path, disk/memory exhaustion, S3 region mismatch, key-policy deny, role trust, pass-role, SCP, endpoint policy, image pull, target capacity, health matcher, alarm missing data, hook timeout, schema incompatibility, quota, and cleanup failure.

Cost and acceptance

Troubleshooting cost includes repeated build minutes, target overlap, NAT/data transfer, logs/traces, retained artifacts, provisioned test environments, and engineer/customer time. This evidence track creates nothing. An approved live lab must stop retries and remove all tagged resources after exporting evidence.

Submit the six complete diagnoses, 20-case matrix, artifact lineage, policy evaluation for one denial, target evidence, rollback decision tree, repair validation, and cost. Pass requires time-ordered proof, no privilege broadening as diagnosis, no artifact ambiguity, idempotent retry reasoning, and verified user recovery.

Official sources

Advertisement