AWS 352: AWS SDK credential providers, retries, pagination, waiters, and error handling
Why this lesson matters
A successful SDK demo can still be unsafe automation. Credentials may come from an unexpected provider, one API response may omit later pages, retries may amplify a non-idempotent change, a waiter may outlive the pipeline, and broad exception handling may report success after partial failure.
This lesson makes those boundaries explicit and testable. It uses Boto3/botocore, but the reasoning applies to all AWS SDKs.
Outcomes
You will be able to:
- explain credential and Region resolution without logging secret values;
- use roles and temporary credentials instead of embedded long-term keys;
- configure connection/read timeouts and standard retry behavior deliberately;
- understand why attempt counts differ by configuration surface;
- consume all paginator pages and control page/item limits;
- use waiters with bounded delay/attempts and post-wait verification;
- classify service, transport, throttling, credential, validation, and waiter failures;
- preserve request IDs and return non-zero outcomes without leaking responses.
Credential provider chain
Boto3 searches providers in precedence order and stops at the first complete credential set. Current documentation includes direct client/session parameters, environment variables, assume-role, web identity, IAM Identity Center, shared files, legacy Boto2 configuration, container credentials, and EC2 instance-role metadata.
Higher precedence is not “more secure.” Passing static credentials directly can override a safer workload role. Production code should normally create a session without secret arguments and let the approved workload provider supply short-lived credentials.
Inspect only non-secret metadata:
import boto3
session = boto3.Session(profile_name=None, region_name="ap-south-1")
credentials = session.get_credentials()
if credentials is None:
raise RuntimeError("no AWS credential provider resolved")
print({"provider_method": credentials.method, "region": session.region_name})
Do not call get_frozen_credentials() and print the result. Access key ID, secret key, and session token are credentials. Provider-method metadata helps diagnose precedence but does not prove authorization.
Region, endpoint, partition, and identity
Require a Region from an argument or controlled deployment configuration. Avoid silently relying on a developer's default profile. Some services are global or have special control Regions, but the SDK client still uses an endpoint model.
Before high-impact work, call STS and compare account/role against an allowlist. Account verification alone is insufficient because roles have different permissions. Partition-aware automation must not construct arn:aws: blindly for GovCloud or China partitions.
Use the SDK-generated ARN returned by services where possible, or derive partition through STS/get-partition model with tests.
Timeouts and retry configuration
from botocore.config import Config
config = Config(
connect_timeout=3,
read_timeout=15,
retries={"mode": "standard", "total_max_attempts": 4},
user_agent_extra="aws-academy-p20/1.0",
)
standard mode provides modeled retry behavior for transient failures and throttling. adaptive adds client-side rate behavior but is not automatically right for shared clients or every workload. Never stack multiple uncoordinated retry loops across SDK, application, queue, and orchestrator.
Attempt semantics are easy to misread. In AWS config or AWS_MAX_ATTEMPTS, max_attempts includes the initial request. In a botocore Config retry dictionary, documented max_attempts historically counts retries, while total_max_attempts expresses total requests and is clearer. Verify behavior against the installed botocore version.
Timeout and retry budgets must fit the caller deadline:
connect + read + backoff across attempts < task deadline < pipeline timeout
Retry only errors classified transient by the SDK/service design. Validation failures, access denial, missing resources caused by bad input, and unsafe partial mutations require a decision, not blind retry. Use idempotency/client tokens when an operation supports them and persist them across retries.
Pagination: one response is not a collection
List APIs often return a page plus token. Use a modeled paginator when available:
from collections.abc import Iterator
from typing import Any
def iter_bucket_names(s3_client: Any, *, max_items: int | None = None) -> Iterator[str]:
# ListBuckets is not paginated in all SDK models; this function is only an example shape.
response = s3_client.list_buckets()
count = 0
for item in response.get("Buckets", []):
name = item.get("Name")
if not isinstance(name, str):
continue
yield name
count += 1
if max_items is not None and count >= max_items:
return
For a truly paginated EC2 operation:
def iter_instance_ids(ec2_client):
paginator = ec2_client.get_paginator("describe_instances")
for page in paginator.paginate(
Filters=[{"Name": "instance-state-name", "Values": ["running", "stopped"]}],
PaginationConfig={"PageSize": 50},
):
for reservation in page.get("Reservations", []):
for instance in reservation.get("Instances", []):
instance_id = instance.get("InstanceId")
if isinstance(instance_id, str):
yield instance_id
PageSize is a service request hint, not necessarily final total. MaxItems can create SDK continuation tokens whose handling differs from service tokens. Stream results when possible; loading a million items can exhaust memory. Deduplicate only when semantics permit it and preserve stable ordering when output is compared.
Tests must include zero, one, and multiple pages, a token cycle/invalid token failure where relevant, malformed item, throttling on a later page, and a consumer that stops early.
Waiters: polling state, not orchestration
A waiter repeatedly calls a read API until an acceptor matches success/failure or attempts are exhausted:
from botocore.exceptions import WaiterError
def wait_for_instance(ec2_client, instance_id: str) -> None:
waiter = ec2_client.get_waiter("instance_status_ok")
try:
waiter.wait(
InstanceIds=[instance_id],
WaiterConfig={"Delay": 15, "MaxAttempts": 20},
)
except WaiterError as error:
last = error.last_response or {}
request_id = last.get("ResponseMetadata", {}).get("RequestId", "unknown")
raise RuntimeError(f"instance waiter failed request_id={request_id}") from error
Compute maximum polling duration, but remember each API call and retry also consumes time. A waiter success proves only its acceptor state. Follow with workload-specific validation such as health endpoint, data integrity, version, and alarm state. A waiter does not provide rollback.
Error taxonomy
| Class | Examples | Response |
|---|---|---|
| Input/model validation | Missing required parameter, invalid type | Fix caller; do not retry |
| Credential/configuration | No credentials, partial credentials, no Region | Fix runtime identity/config |
| Authorization | AccessDenied, KMS deny | Identify caller and policy path; least-privilege correction |
| Throttling/transient service | throttling, selected 5xx | SDK retry within budget; reduce concurrency |
| Network/endpoint | DNS, TLS, proxy, connect/read timeout | Diagnose path; bounded retry only if transient |
| Resource state/conflict | not found, conflict, incorrect state | Re-read state and decide; avoid blind retry |
| Waiter exhaustion/failure | timeout or terminal acceptor | Capture last safe metadata; recover/rollback |
| Partial result | later page fails after earlier output | Mark incomplete; never report full success |
Catch specific botocore exceptions before broad BotoCoreError. For ClientError, inspect Error.Code, HTTP status, RequestId, and retry attempts from ResponseMetadata while avoiding full sensitive responses.
def error_summary(error) -> dict[str, object]:
response = error.response
metadata = response.get("ResponseMetadata", {})
details = response.get("Error", {})
return {
"code": details.get("Code", "Unknown"),
"http_status": metadata.get("HTTPStatusCode"),
"request_id": metadata.get("RequestId", "unknown"),
"retry_attempts": metadata.get("RetryAttempts", 0),
}
Assume-role session boundaries
Role assumption requires permission by the caller and trust by the target role. Add ExternalId only when the trust design requires it; it is not a secret. Use a meaningful, non-sensitive session name and optional source identity/tags where governance requires them. Limit duration and scope. Cache temporary credentials in memory and refresh before expiration rather than writing them into repository, artifact, log, or command history.
Do not log STS response credentials. Test clock-skew/expiration behavior and ensure long jobs can refresh through the provider instead of capturing one frozen credential set at startup.
Failure-injection lab
Use Stubber or a fake adapter to test:
- first call succeeds;
- no credentials resolve;
- explicit
AccessDeniedwith request ID; - throttling followed by success at the SDK layer;
- multiple paginator pages;
- failure on page three after items were emitted;
- empty page with continuation behavior supported by the model;
- waiter success followed by failed application health;
- waiter terminal failure;
- waiter maximum attempts exceeded;
- wrong account or Region guard;
- token/credential expires during a long process.
Assert exit state, completion marker, redaction, request ID, and whether partial output is discarded or explicitly labelled incomplete.
Independent challenge and acceptance
Build a read-only iterator for one approved paginated operation. Add explicit Region, finite timeouts, standard retry with total attempts, account/role guard, stable output, a maximum-item control, and structured error summary. Test all 12 cases without live credentials. Optionally run one approved live read and redact identity.
Pass requires complete pagination, no hardcoded credentials, no secret logs, a bounded total deadline, no retry of permanent denial/validation, no success after partial-page failure, waiter postcondition verification, and version-recorded tests.