AWS 354: Build a tested repository that calls AWS read-only APIs
Why this lab matters
Real automation is more than one successful SDK call. A trustworthy repository has a documented purpose, explicit identity and Region boundaries, reproducible dependencies, modular code, pagination, safe errors, deterministic tests, review controls, CI, redaction, release evidence, and an operator runbook.
You will assemble AWS348-AWS353 into a read-only account inventory tool. It lists enabled Regions, selected VPC metadata, and S3 bucket names while redacting account identity by default. It creates, updates, and deletes no AWS resource.
Acceptance architecture
operator/CI identity -> CLI parser -> account/role guard
-> Boto3 clients -> paginated read APIs -> normalized model
-> deterministic JSON -> redacted log/evidence
Git commit -> review/checks -> annotated tag -> source archive digest
The repository must work in two modes:
- offline mode: all tests run with modeled stubs and no AWS credentials/network;
- authorized live mode: an approved temporary identity runs only the documented read operations.
Required repository structure
aws-read-inventory/
├── .gitignore
├── README.md
├── SECURITY.md
├── OPERATIONS.md
├── pyproject.toml
├── requirements-lock.txt
├── src/aws_read_inventory/
│ ├── __init__.py
│ ├── cli.py
│ ├── collector.py
│ ├── errors.py
│ └── model.py
├── tests/
│ ├── test_cli.py
│ ├── test_collector.py
│ └── test_redaction.py
└── evidence/
└── README.md
Do not commit .venv, credentials, .env, AWS configuration, account IDs, live output, caches, or coverage HTML.
Step 1: define the contract
The CLI interface is:
aws-read-inventory --region REGION [--profile PROFILE]
[--expected-account HASHED_OR_APPROVED_VALUE]
[--max-vpcs N] [--format json|text]
Define exit codes: 0 complete success, 2 input/configuration, 3 identity guard, 4 AWS service/authorization, 5 network/SDK, 6 incomplete pagination/normalization. stdout contains only requested result; stderr contains redacted logs.
The output schema includes tool schema version, collection timestamp, requested Region, redacted account field, caller type, enabled Regions, VPC IDs replaced by stable per-run labels unless sensitive evidence is approved, CIDRs redacted by default, bucket names redacted or hashed by default, completion state, and source operation names. Never call an incomplete result complete.
Step 2: least-privilege API inventory
The application uses:
sts:GetCallerIdentityfor runtime identity;ec2:DescribeRegionsfor enabled Region names;ec2:DescribeVpcsin the selected Region with paginator support;s3:ListAllMyBucketsfor bucket metadata returned byListBuckets.
GetCallerIdentity has special permission behavior but still exposes account/ARN data that must be redacted. Read APIs can reveal sensitive topology and names. Restrict who can run the tool and where evidence is stored.
Create an IAM policy design scoped as tightly as each action supports; document actions that do not support resource-level restriction. Add an explicit deny boundary for mutating actions at a higher governance layer if the lab identity is dedicated. Do not attach a new role merely for training unless the account owner separately authorizes it.
Step 3: package and dependency setup
mkdir aws-read-inventory
cd aws-read-inventory
git init -b main
python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install boto3 pytest
python -m pip freeze > requirements-lock.txt
mkdir -p src/aws_read_inventory tests evidence
touch src/aws_read_inventory/__init__.py
Create pyproject.toml:
[build-system]
requires = ["setuptools>=69"]
build-backend = "setuptools.build_meta"
[project]
name = "aws-read-inventory"
version = "0.1.0"
requires-python = ">=3.11"
dependencies = ["boto3"]
[project.scripts]
aws-read-inventory = "aws_read_inventory.cli:entrypoint"
[tool.pytest.ini_options]
pythonpath = ["src"]
addopts = "-q"
Record exact Python, pip, Boto3, botocore, and operating-system versions in release evidence. The broad dependency declaration supports packaging; the reviewed lock controls the tested build environment.
Step 4: implementation rules
Reuse and improve the models from AWS351. The collector must receive a session/client factory rather than constructing hidden globals. Apply finite connect/read timeouts and standard retry total attempts. Use the describe_vpcs paginator, enforce max_vpcs in the consumer, normalize sets to sorted tuples, and validate every required field.
Identity guard order:
- resolve credentials through the approved provider chain;
- call STS;
- classify caller ARN without printing it;
- compare account and allowed role pattern through a non-secret deployment configuration;
- stop before inventory if mismatch occurs.
Do not accept an expected account solely from attacker-controlled repository input in CI. The trusted workflow/environment owns that boundary.
All logs use allowlisted fields: level, event, operation, error code, HTTP status, request ID, retry count, page number, item count, and duration. Never serialize the full response or environment.
Step 5: required tests
Use Stubber or injected fakes. Tests must never discover local credentials. Supply inert values such as testing only inside client construction.
| Test | Required assertion |
|---|---|
| Help and invalid arguments | No API call; exit 0/2 correctly |
| Expected identity | Collection proceeds; output account redacted |
| Wrong account/role | Stops before EC2/S3; exit 3 |
| Zero VPCs | Complete empty list, not error |
| Two VPC pages | All items included once and sorted |
| Maximum VPCs | Stops at exact configured item count |
| Failure on page two | Output marked/withheld as incomplete; non-zero |
AccessDenied | Code/request ID logged; no full response |
| Endpoint/timeout | SDK category and non-zero exit |
| Malformed modeled field | Normalization error, not silent omission |
| Bucket/name redaction | Canary values absent from stdout/stderr |
| Determinism | Same modeled input produces byte-identical normalized result |
Also test that no method outside the four approved API operations is called. A unit test cannot prove IAM least privilege or live service pagination, so record those limits.
Step 6: local quality gate
python -m compileall -q src tests
pytest
git diff --check
git status --short
If approved tools are installed, add formatting, linting, type checking, dependency vulnerability review, license review, and secret scanning. Pin their versions in CI. A scanner finding must be triaged; suppressions require owner, rationale, scope, and expiry.
Create a fake secret canary in a test fixture, run the application and tests, then scan stdout, stderr, repository files, Git history, and built archive for it. The test should confirm default redaction while never using a real secret.
Step 7: CI design
The untrusted proposal workflow runs offline checks only and receives no AWS credential. A separate protected, manually authorized workflow may use OIDC federation for live read-only validation. Trust conditions bind organization/repository, branch or environment, audience, and workflow context as supported. Session duration and policy are minimal.
Required CI stages:
- clean checkout of exact commit;
- supported Python setup;
- dependency install from reviewed lock/cache policy;
- syntax/static checks;
- offline tests and coverage threshold;
- secret/dependency scan;
- package/source archive build;
- archive file-list inspection;
- SHA-256 digest and provenance manifest;
- artifact upload with explicit access and retention.
Pull requests from forks or untrusted authors must never execute with protected cloud credentials. Changing the workflow itself requires specialist ownership and independent review.
Step 8: optional live read-only validation
Only with account-owner approval and a temporary identity:
aws sts get-caller-identity --profile training --output json
aws-read-inventory --profile training --region ap-south-1 --max-vpcs 20 --format json
Privately compare identity with the approved target. Share only redacted output. Verify CloudTrail events if authorized and note that management-event recording/query availability depends on trail/event-history design. No cleanup of AWS resources is required because none were created.
Failure injection in live mode must remain read-only: wrong profile/account guard, disallowed Region input, deliberately insufficient read role in an approved sandbox, network denial in an isolated environment, and output-storage permission denial. Do not alter production policies merely to create a failure.
Step 9: review, release, and provenance
Create focused commits, open a review proposal, record exact base/head, require tests and sensitive-path ownership, and resolve findings. After merge, prove clean status and create an annotated tag:
git status --porcelain
pytest
git tag -a v0.1.0 -m "Read-only inventory training release v0.1.0"
git archive --format=tar.gz --output=../aws-read-inventory-v0.1.0.tar.gz v0.1.0
sha256sum ../aws-read-inventory-v0.1.0.tar.gz
git rev-parse v0.1.0^{}
Inspect the archive list before distribution. The release manifest binds tag object, peeled commit, archive digest, lock digest, test result, tool versions, reviewer decision, and optional live evidence ID. Do not move the published tag.
Operations and incident runbook
Document installation, identity prerequisites, Region behavior, API list, expected duration, output sensitivity, exit codes, log locations, quota/throttling behavior, troubleshooting, and escalation.
Incident cases include unexpected account, access denied, throttling, partial pagination, dependency compromise, leaked output, CI credential exposure, artifact mismatch, and source-host outage. For a suspected credential exposure: stop jobs, revoke/expire sessions or rotate affected material, preserve redacted evidence, assess CloudTrail and downstream use, correct trust/policy, and retest.
Final acceptance
Submit source, README/security/operations docs, policy design, threat model, dependency lock, 12-test evidence, canary scan, CI design, review record, release tag, archive and digest, provenance manifest, optional redacted live result, failure records, and residual risks.
Pass requires all offline tests, no real credentials, exactly the approved read operations, complete pagination, finite timeouts/retries, identity guard before inventory, non-zero partial failure, default redaction, deterministic output, no credentials in untrusted CI, reviewed immutable release, and production untouched.