AWS 032: AWS API requests, SDK purpose, retries, and idempotency introduction
The problem
An application sends a create request. The network times out before the response arrives. The developer retries, creating a duplicate resource. Another developer retries an invalid request until it causes more load.
Reliable API use requires authentication, authorization, request scope, error classification, controlled retries, and idempotency.
Final outcome
You will map the lifecycle of an AWS API request, classify retryable and non-retryable outcomes, design an idempotency record for a create operation, and inspect a safe read-only request through the CLI and SDK.
Request lifecycle
client builds request
|
v
credential provider supplies temporary credentials
|
v
request is signed for service, time, and scope
|
v
AWS endpoint authenticates and authorizes
|
v
service validates and starts or completes operation
|
v
response, error, or ambiguous timeout
AWS APIs commonly use HTTPS and Signature Version 4. SDKs and the CLI handle request serialization, signing, endpoint selection, response parsing, and credential refresh. Writing SigV4 manually is rarely the correct application choice.
Control plane and data plane
A control-plane request creates or changes configuration, such as creating a queue or changing a route. A data-plane request works with the service's main workload function, such as putting an object or reading an item.
The boundary is service-specific. Both types require security, limits, logging, retries, and cost understanding.
Anatomy of a request
A useful request record includes:
- service and API operation;
- account and principal;
- Region or endpoint;
- resource identifier;
- parameters;
- client request token where supported;
- timestamp;
- timeout;
- attempt number;
- response status and AWS request ID;
- resulting resource identifier;
- audit event location.
Do not log secret keys, session tokens, authorization headers, or sensitive request bodies.
SDK purpose
An AWS SDK provides:
- language-native clients and data types;
- credential-provider integration;
- request signing;
- endpoint and Region handling;
- response and exception parsing;
- service pagination;
- retry behavior;
- waiters or polling utilities for some operations.
An SDK cannot choose the correct architecture, permission boundary, retry safety, or business transaction. Those remain application responsibilities.
Classify failures before retrying
| Outcome | Typical action |
|---|---|
| invalid parameter | correct request, do not blind retry |
| access denied | inspect identity and policy, do not blind retry |
| not found after expected creation | consider consistency and operation state |
| throttling | retry with backoff and jitter within limits |
| transient service error | bounded retry according to SDK/service guidance |
| network timeout | outcome may be unknown; use idempotency and verify |
| conflict | inspect current state and operation semantics |
HTTP status alone is not always enough. Read the service error code and operation documentation.
Backoff and jitter
Immediate synchronized retries can worsen a service problem. Exponential backoff increases delay after each failure. Jitter randomizes delays so many clients do not retry together.
Use the supported SDK retry mode for the language and version. AWS retry defaults and features can change. Record the chosen mode, maximum attempts, total time budget, and which errors qualify. Do not stack several independent retry layers without modelling the combined attempts.
For example, three retries in the SDK inside three job retries inside three workflow retries can produce far more calls than expected.
Idempotency
An idempotent operation can receive the same logical request more than once without applying the effect more than once.
Some AWS operations are idempotent by default. Some accept a client token. Others require the application to design deduplication around a business key or state machine.
For a supported client token:
first request + token T + parameters P -> resource R
retry request + token T + parameters P -> same logical result
request + token T + changed parameters -> mismatch error
Use a unique token per logical operation. Do not reuse one token for unrelated requests.
Assessment submission example
The training portal accepts one final submission per learner and assessment attempt.
Choose idempotency key:
submission:<learner-id>:<assessment-id>:<attempt-id>
Store:
- idempotency key;
- normalized request hash;
- processing status;
- result identifier;
- created and expiry times;
- final response needed for a retry.
On retry:
- same key and same request returns the existing result;
- same key with different request is rejected;
- in-progress state is handled without starting a second write;
- expired records follow a documented retention rule.
Do not place personal data directly in an externally visible token. Use opaque identifiers or a protected hash when appropriate.
Safe practical inspection
Use the course profile:
aws sts get-caller-identity \
--profile course \
--output json \
--no-cli-pager
printf 'cli_exit=%s\n' "$?"
Then use the SDK script from AWS 031. Record:
- operation:
GetCallerIdentity; - endpoint Region configuration;
- credential provider type;
- exit status;
- response fields;
- what was redacted;
- whether the operation mutated state.
Do not capture --debug output in shared evidence. Debug logging can include endpoints, headers, request details, and identifiers that require careful handling.
Retry decision exercise
Create api-safety.md and classify:
AccessDeniedwhen listing a resource.ValidationExceptionfor an invalid CIDR.- throttling on a read operation.
- timeout after a create request was sent.
- duplicate event delivery to a worker.
- client token reused with changed parameters.
For each write:
- retry or do not retry;
- verification source;
- idempotency control;
- maximum attempt/time boundary;
- final operator evidence.
Expected direction: do not retry access or validation errors unchanged; use bounded backoff for documented transient/throttling failures; verify ambiguous writes using token and resulting state; deduplicate repeated events; treat token/parameter mismatch as a caller defect.
Common failures
| Failure | Consequence | Correction |
|---|---|---|
| retry every exception | waste and duplicate effects | classify errors |
| no total timeout | worker hangs too long | bound attempts and elapsed time |
| random token on every retry | duplicates still occur | reuse token for same logical request |
| token reused for new request | mismatch or wrong association | unique operation key |
| success response before durable commit | accepted work can disappear | define acknowledgement boundary |
| logs contain auth data | credential exposure | structured redaction |
Certification decisions
Look for requirements such as duplicate delivery, at-least-once processing, timeout ambiguity, throttling, or safe automation. The answer often combines a durable state transition, an idempotent consumer, bounded SDK retries, and observable request identifiers.
Retries improve transient-failure handling. They do not replace high availability, queues, capacity planning, or correct authorization.
Completion gate
Pass when api-safety.md correctly classifies all six outcomes, the request lifecycle identifies identity and scope, the idempotency design prevents duplicate submission effects, and the learner explains why a timeout does not prove failure.
No AWS resources were created or changed.