Lesson 079 · AWS Learning Path

AWS 079: Diagnose address, route, security, metadata, storage, and status-check failures

· Published · 7 min read

A regional VPC address space is divided across three zones with reserved growth and a non-overlapping connected network

The problem

A team uses Diagnose address, route, security, metadata, storage, and status-check failures as a label or Console setting without proving the identity, scope, behavior, failure boundary, cost, or operational result. The configuration appears complete but the real requirement remains untested.

Final outcome

The learner will inspect, explain, test, and document Diagnose address, route, security, metadata, storage, and status-check failures. The submission must connect the Console and CLI view to the same AWS control, state what the evidence proves, diagnose one likely failure, and make a requirement-based architecture decision.

Learning outcomes

You will be able to:

  1. explain break-fix scope;
  2. explain one change at a time;
  3. explain packet path;
  4. explain metadata path;
  5. explain storage path;
  1. distinguish successful configuration from successful workload behavior;
  2. preserve redacted evidence and complete the stated cleanup or retention decision.

Mental model

requirement
   -> identity and permission
   -> account, Region, and resource scope
   -> configuration or request
   -> observable state
   -> workload result
   -> failure evidence
   -> cleanup or controlled retention

Never begin with a create button or a copied command. Start with the result that must be proven and the boundary that must remain protected.

Core facts

ConceptWhat it means in practice
Break-fix scopeDiagnose supplied faults in address assignment, route association, security group, NACL, IMDSv2 use, EBS mount persistence, and EC2 status evidence.
One change at a timeCapture baseline, state hypothesis, make the smallest reversible correction, retest the original symptom, and record rollback.
Packet pathEvaluate DNS, source address, route, gateway, SG, NACL, destination listener, and return path in order.
Metadata pathIMDSv2 needs a token, correct endpoint options, supported hop limit, and host/application access. Never print role credentials.
Storage pathMatch volume, AZ, attachment, device, filesystem, UUID, mount and fstab. Never format during diagnosis.
Health pathSeparate system, instance, attached-EBS, agent, service, and application health.

How to reason about it

The instructor introduces only faults that can be recovered without data destruction. The learner receives symptom and success criteria, not the fix. Root, broad inbound access, policy expansion, metadata credential display, and filesystem formatting are prohibited.

If a fault cannot be reproduced, record current state and evidence instead of changing random controls. Cleanup restores the approved P04 state or proceeds directly to AWS 080.

Use the following decision table as a starting point, then change the answer when the scenario changes.

RequirementPreferred directionWhy
No public IPv4Inspect launch and ENI address assignmentIGW route alone does not assign an address.
Flow rejectedRead SG and NACL evidenceStateful and stateless controls differ.
Tokenless metadata failsExpected with IMDSv2 requiredUse token flow rather than weakening options.
Mount missing after rebootInspect UUID and fstabDo not recreate filesystem.

Prerequisites, permissions, Region, and cost

Run this lesson as the approved non-root course identity. Begin with aws sts get-caller-identity, confirm the private account record, and set the fixed project Region before any regional query. IAM resources are global within an account, while STS endpoints and the services reached by an identity can be regional.

Use only the read or change actions required for this lesson. An AccessDenied result is evidence to analyse, not permission to switch to root or attach AdministratorAccess. Record the action, resource, principal type, request context, and smallest justified correction.

The listed cost tier is T1 - free or near-free with immediate cleanup. Free Tier eligibility and credits are account-specific. Before a mutating lab, identify every resource that can charge, estimate its duration, start a timer, and write the reverse cleanup order. A budget reports cost after billing data arrives and is not a real-time stop control.

AWS Management Console method

  1. Open EC2 instance Networking, Security, Storage, Status checks, and Systems Manager managed-node state.
  2. Open VPC route tables, subnet associations, NACLs, and Flow Logs or Reachability Analyzer evidence available to the lab.
  3. Change only the instructor-approved faulty value, retest, and record before/after state.

For every step record the service page, selected account and Region, exact object, visible state, and why that state matters. A Console label or green status is not enough unless it is tied to the final workload outcome.

AWS CLI or API evidence

Collect one read-only snapshot across network, metadata options, storage, and status.

aws ec2 describe-instances --region ap-south-1 --filters Name=tag:Name,Values=nw-p04-web-1 --query 'Reservations[].Instances[].{State:State.Name,Subnet:SubnetId,Public:PublicIpAddress,SG:SecurityGroups,Metadata:MetadataOptions}'
aws ec2 describe-instance-status --region ap-south-1 --include-all-instances

Use IDs privately to join route, security, and storage evidence. No single output proves application success.

Before running the command, replace every placeholder, explain each option, and decide whether the operation is read-only or mutating. Capture the exit code immediately. Redact account IDs, ARNs, public addresses, request identifiers, and personal data before sharing.

The CLI and Console are clients of AWS APIs. Matching state across them increases confidence, but neither substitutes for data-plane or application verification.

Practical work

Start the retained nw-p04-web-1 only for this timed exercise. Diagnose at least four instructor-supplied faults, including one network, one metadata or identity, one storage, and one health-state fault. For each submit symptom, timestamp, evidence, hypothesis, one correction, retest, and rollback. Restore the approved baseline and stop the instance when evidence is complete, unless AWS 080 starts immediately. Do not proceed if the target or data-safety boundary is uncertain.

The evidence package must include:

  • non-root principal type, account verified privately, and Region;
  • exact intended and observed state;
  • one Console observation and the matching CLI or API result;
  • one successful result and one denied, failed, or counterexample result;
  • what each result does not prove;
  • cost state and elapsed lab time;
  • cleanup evidence or a named retained-state owner and next lesson.

Verification standard

Use three levels of proof:

  1. Control plane: the object or policy exists with the intended configuration.
  2. Data plane or behavior: the request, packet, session, storage path, or application does what the requirement states.
  3. Operations: monitoring, failure diagnosis, cost, ownership, and cleanup are known.

A control-plane response can precede final readiness. Use waiters or state polling where supported, then test the actual behavior. If a request times out, do not assume it failed. Inspect state and use documented idempotency before retrying a mutation.

Troubleshooting method

SymptomEvidence firstSmallest safe response
command cannot authenticatecredential source, expiry, caller preflightrestore approved temporary login
access is deniedaction, resource, principal, all policy layerscorrect only the missing or conflicting control
object appears missingaccount, Region, filters, pagination, permissionalign scope before creating anything
configured state exists but behavior failsroute, identity, dependency, logs, service statetest the next boundary in the path
cleanup is blockeddependency inventory and owning serviceremove dependants in reviewed reverse order

Keep the original symptom and timestamp. State one hypothesis, make one reversible change, repeat the original test, and record rollback. Never open a management port to the world, expose credentials, disable TLS verification, format an unknown disk, or add broad permissions as a generic fix.

Architecture and certification decisions

Professional-level questions provide competing valid features. Identify the requirement that decides between them: human or workload identity, same-account or cross-account access, regional or zonal scope, stateful or stateless filtering, durable or ephemeral data, latency, RTO/RPO, cost, or operational ownership.

Explain why the selected option fits and why each plausible alternative fails one stated requirement. Do not rely on feature memorization or reproduce protected certification questions.

Knowledge check

  1. Should SSH be opened for diagnosis?

Expected direction: No. Use Session Manager and path evidence.

  1. Should IMDSv1 be enabled when tokenless curl fails?

Expected direction: No. Use the correct IMDSv2 token sequence.

  1. Should mkfs repair a missing mount?

Expected direction: No. It can destroy data.

  1. How many variables change per hypothesis?

Expected direction: One smallest reversible change.

Cost, cleanup, and retained state

Retain only the state explicitly required by the next P04 lesson and record it in the private resource ledger.

Cleanup evidence includes the final state query, not only a successful delete response. Search related resources, other Regions used by the lab, retained storage, public IPv4 addresses, logging destinations, and service-managed dependencies. Schedule a later billing review because cost data can lag.

Completion gate

Pass when the practical artifact explains the real problem, matches Console and CLI evidence, answers every knowledge check, diagnoses one failure without broadening access, records cost, and proves cleanup or approved retention. The learner must defend one decision orally and name the requirement that would change it.

Official sources

Advertisement