AWS 109: S3 buckets, objects and keys
The real problem
A team recognizes the name S3 buckets, objects and keys but has not connected the feature to a real requirement, identity boundary, network or data path, failure mode, price dimension, and cleanup owner. A plausible configuration could still fail the workload.
Final outcome
The learner will produce a requirement-led artifact for S3 buckets, objects and keys, inspect the matching AWS control plane in the Management Console, run a matching CloudShell or AWS CLI query, interpret the output, diagnose one failure, defend one architecture choice, and prove cleanup or approved retained state.
The practical outcome is not a command transcript. It must show what was expected, what happened, what the result proves, what it does not prove, and which evidence would change the decision.
Learning objectives
By the end of this lesson, the learner can:
- explain bucket naming;
- explain flat namespace;
- explain object composition;
- explain size boundaries;
- explain metadata versus tags;
- connect control-plane state to the real data, network, identity, or application behavior;
- identify cost and cleanup ownership before any optional mutation;
- troubleshoot from evidence without opening broad access or adding broad permissions.
Relationship model
Requirement
|
v
Identity and policy -> AWS configuration -> network or data path -> workload behavior
| | | |
+--------------------+----------------------+--------------------+
|
v
monitoring, cost, recovery, cleanup
Use this model to separate an AWS object that exists from a result that actually works. Every arrow is a verification boundary.
Prerequisites, permissions, Region, and safety
- Learning baseline: This sequence assumes practical Linux knowledge but no prior cloud-computing or AWS knowledge. Cloud, networking, security, data, automation, and architecture concepts must come from completed earlier lessons. If a prerequisite checkpoint is incomplete, return to its linked lesson before continuing.
- Confirm a non-root caller with
aws sts get-caller-identityand keep the account number private. - Use
ap-south-1unless this lesson explicitly names a second Region. - Confirm the intended profile and Region with
aws configure listbefore interpreting an empty result. - Use read-only List, Get, and Describe permissions for the named services. Design exercises run locally and require no resource-creation permission.
- This is a no-create lesson. Console and CLI work is read-only, and every design artifact is created locally.
- Never publish account IDs, public addresses, ARNs containing private account data, session IDs, presigned URLs, object data, credentials, or KMS material.
- Do not use root, world-open SSH or RDP, disabled TLS verification, unowned resources, or irreversible retention controls in a training exercise.
Core model
| Concept | What the learner must understand |
|---|---|
| Bucket naming | A traditional general purpose bucket name is unique across accounts and Regions in its AWS partition. AWS also supports account Regional namespace names with a required account/Region suffix. Names must follow current rules, must not expose confidential information, and cannot be changed after creation. Deleting a shared-global-name bucket can let another account claim that name, so keep an empty protected bucket when name continuity matters. |
| Flat namespace | Object keys form a flat namespace. Slashes create prefix-based organization in tools and the Console, not real nested directories with filesystem semantics. |
| Object composition | An object includes data, a key, system metadata, optional user metadata, optional tags, access and encryption state, checksums, and a version ID when versioning is enabled. |
| Size boundaries | Current S3 documentation permits an object up to 48.8 TiB, derived from 10,000 multipart parts of at most 5 GiB each. A single PutObject request remains limited to 5 GiB; multipart upload is required beyond that and is normally considered around 100 MiB for resilience and parallelism. Verify SDK/tool support because older material commonly states the historical 5-TB object limit. |
| Metadata versus tags | User-defined metadata is set when an object is created or copied and is returned with object metadata. Object tags are a separate key-value control used by lifecycle, policy, replication, and cost allocation scenarios. |
| Ownership | With Bucket owner enforced Object Ownership, ACLs are disabled and the bucket owner owns uploaded objects. Authorization is then policy-centered. |
How it works
Design keys for retrieval, partitioning, events, lifecycle, and human operation without embedding secrets or mutable business meaning. Store searchable business indexes in an appropriate database rather than assuming S3 lists are a relational query engine.
Read the result in layers:
- Scope: account, Region, VPC, bucket, AZ, endpoint, principal, object version, or resource ARN.
- Control plane: the requested configuration exists and reached an expected state.
- Behavior: the request, connection, health check, replication, restore, or application result meets the requirement.
- Operations: monitoring, failure owner, cost, retention, rollback, and cleanup are known.
Control-plane success is necessary but not sufficient. A resource can be available while policy, routing, DNS, health, data, or application behavior remains wrong.
Architecture decision table
| Requirement | Preferred direction | Why |
|---|---|---|
| Need lifecycle by data category | Stable prefix or object-tag scheme | Rules can select by prefix and supported tag filters. |
| Need to change custom metadata | Copy the object with replacement metadata | User metadata is not edited in place. |
| Need exact recovery of one overwrite | Record the version ID | The key alone resolves the current version. |
| Multiple accounts upload to one controlled bucket | Bucket owner enforced plus bucket and identity policies | Avoid ACL-dependent ownership surprises. |
Professional questions normally contain several valid services. State the requirement that selects one option, why the nearest alternative fails it, and what changed requirement would reverse the choice.
Deep dive: bucket creation is an architecture decision
Before creation, record:
- bucket type and namespace;
- Region or Availability Zone scope;
- name owner and deletion/name-reuse policy;
- data classification and prohibited name/key content;
- Object Ownership and Block Public Access posture;
- encryption/key choice and KMS administrators/users when applicable;
- versioning, Object Lock, lifecycle, replication, backup, logging, event, and monitoring requirements;
- expected object count, sizes, request rates, transfer sources/destinations, retention, and cost owner;
- IaC source, change approval, recovery test, and final deletion authority.
For a shared-global-name bucket, choose a random/GUID-style suffix and never place account secrets, customer names, email addresses, or confidential project details in the DNS-visible bucket name. Account Regional namespace buckets use an AWS-defined account-and-Region suffix and cannot be recreated by another account, but their exact API/tool/integration support should still be validated in the target environment.
us-east-1 has historical API differences around CreateBucketConfiguration; other Regions normally require a matching location constraint. Always create through a Regional endpoint and verify returned location rather than assuming the CLI's configured Region changed where an existing bucket lives.
Deep dive: keys, prefixes, and object identity
The key is the entire UTF-8 object name. A design such as tenant=42/year=2026/month=09/report-uuid.json can support human operations, policy/lifecycle filters, and analytics partition discovery. It can also expose sensitive values in URLs, logs, CloudTrail, metrics, and support records. Use opaque identifiers where disclosure is a concern.
Key-design rules:
- Treat key spelling, case, spaces, reserved URL characters, Unicode normalization, trailing periods, and relative path-looking segments as interoperability concerns. Test the exact SDK, CLI, proxy, and consumer.
- Do not make a mutable display name the only permanent identifier. Prefer immutable IDs and keep searchable business attributes in a database/catalog.
- Decide overwrite semantics. Immutable unique keys avoid writer races; stable keys require versioning and/or conditional requests.
- Design prefixes and tags for lifecycle, replication, IAM, events, inventory, cost allocation, and deletion scope before data volume makes migration expensive.
- Remember that an empty console “folder” can be a zero-byte object whose key ends in
/; deleting that marker does not delete all similarly prefixed keys.
An exact versioned identity is (bucket or access-point identity, key, version ID). A URL without a version ID normally addresses the current version. A current delete marker can return not-found behavior while older versions still exist and are billed.
Deep dive: object data, metadata, tags, checksums, and ETags
- The object value is an opaque byte sequence to S3.
Content-TypeandContent-Encodingmetadata tell clients how to interpret it; incorrect values can create browser/application failures without changing the bytes. - User metadata is supplied as HTTP headers and is copied/replaced with the object. Changing it normally requires a copy operation, which can create a new version and incur request, storage, KMS, and transfer effects.
- Object tags are managed through separate operations. They can drive IAM conditions, lifecycle, replication, and Batch Operations. Plan who may change classification tags because a tag change can alter retention or access behavior.
- ETag is not a universal file hash. A single-part, unencrypted/simple case may resemble an MD5, but multipart and encryption behavior break that assumption.
- Use S3 checksum support and a selected algorithm for integrity requirements. Record whether the client calculated the whole-object checksum, how multipart composite/full checksums behave, and whether download validation was actually enabled.
Deep dive: object operations and their side effects
| Operation | Meaning | Common hidden effect |
|---|---|---|
HeadObject | Return metadata/status without object bytes | Still requires permission; KMS/checksum mode can add permission requirements. A 403 versus 404 can depend on list permission. |
GetObject | Retrieve the current or specified version, optionally by byte range | Archive state, KMS authorization, checksum mode, response overrides, transfer, and requester-pays behavior can matter. |
PutObject | Atomically create/replace one object value | Can overwrite a stable key, create a version, invoke KMS, trigger events/replication, and produce request/storage cost. |
| multipart upload | Initiate, upload numbered parts, then complete or abort | Parts are not the final object before completion; abandoned parts remain billable; completion order uses part numbers, not upload time. |
CopyObject/multipart copy | Server-side copy that can replace metadata, encryption, class, ownership, or key | It is a new write, not a free rename; large copies require multipart copy and may create versions/charges. |
DeleteObject | Delete current unversioned data or add a delete marker in an enabled bucket | It might hide rather than remove bytes. Permanent version deletion needs the version ID and permission. |
ListObjectsV2 | Paginated listing of current keys | It is not a database query or a list of every version. Handle continuation tokens and prefixes. |
ListObjectVersions | Paginated listing of versions and delete markers | Required for a truthful version-aware inventory and complete cleanup. |
Use conditional requests to make concurrency visible. If-None-Match: * can prevent creating a key that already has a current object; If-Match: <etag> can reject a write/delete if the expected object changed. A 412 Precondition Failed or concurrency-related 409 Conflict is useful control evidence, not a reason to disable the condition.
Endpoints, URLs, and temporary access
- Virtual-hosted-style requests place the bucket/access-point identity in the hostname. TLS and DNS rules make bucket naming operationally important.
- A Regional REST endpoint supports authenticated object API behavior. An S3 website endpoint is a different, HTTP-oriented website feature and does not provide the same private-origin/TLS model.
- Gateway and interface VPC endpoints have different routing, source, policy, DNS, availability, and cost properties. Private routing does not create authorization by itself.
- Transfer Acceleration uses an accelerate endpoint and edge network; it must be enabled, is not compatible with every naming/path scenario, and should be benchmarked.
- A presigned URL carries the signing principal's delegated permission for one operation until its effective expiry, subject to credential lifetime and policy controls. Treat it as a bearer secret: do not log or publish it, keep duration and operation narrow, and remember that generating a URL does not grant permissions the signer lacks.
Three required behavior tests
- Positive: upload a uniquely named object with explicit content type and checksum, then
HEADand download it; verify length, metadata, encryption response, version ID where enabled, and local checksum. - Negative: attempt an unauthorized prefix or anonymous request and preserve the denial. Do not “fix” it by making the bucket public.
- Dependency/concurrency: issue a conditional write with a deliberately stale ETag or existing-key condition and verify rejection. Explain how that prevents a lost update or unintended overwrite.
A passing test records expected status before execution, exact principal and Region with identifiers redacted, observed status/error code, what it proves, what it does not prove, and cleanup of all versions and incomplete multipart uploads.
AWS Management Console guided practice
Before opening a service page, write the expected account, Region, starting state, and evidence. Do not choose Create, Save, Purchase, Lock, or Delete unless the lesson explicitly authorizes the live track.
- Open an instructor-owned S3 bucket and inspect one object Name, prefix, Size, Type, Last modified, Storage class, Version ID, metadata, tags, and encryption.
- Enable Show versions in the supplied evidence view and separate current versions, noncurrent versions, and delete markers.
- Draft the two globally unique AWS 120 bucket names using the approved
nw-p06-source-{unique-suffix}andnw-p06-replica-{unique-suffix}pattern; check availability only during the lab.
For each step, capture the field name and value in text. A screenshot may support the record but does not replace the explanation. Console labels can evolve, so use the service search and current documentation if a navigation label differs.
CloudShell and AWS CLI practice
CloudShell is the default browser-based command environment taught in AWS 028. AWS 029 and AWS 030 cover local CLI installation and authentication. This lesson therefore does not assume that an unconfigured local shell is ready.
Start every session with:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account portion of the ARN before sharing. Then perform the topic query:
Inspect object key, size, timestamp, storage class, checksum or ETag context, and version ID on an approved bucket.
NW_BUCKET="replace-with-owned-bucket-name"
aws s3api list-object-versions --bucket "$NW_BUCKET" --prefix documents/ --query '{Versions:Versions[].{Key:Key,Version:VersionId,Latest:IsLatest,Size:Size,Class:StorageClass},DeleteMarkers:DeleteMarkers[].{Key:Key,Version:VersionId,Latest:IsLatest}}' --output json
Expected interpretation:
Version output distinguishes data versions from delete markers. An ETag must not be assumed to be a plain MD5 checksum for every encryption and multipart case.
Replace every replace-with-... sample value before running its command, and use only an explicitly owned resource. Explain each option first. These queries are read-only; a successful response does not authorize a later create or delete operation.
Practical work
Create p06-object-model.md. Define exact keys documents/policy.txt, documents/runbook.txt, and evidence/restore-proof.txt; allowed metadata; tags Project=NitWings-P06 and DataClass=Training; version-ID ledger; prefix ownership; case-sensitivity examples; naming privacy; checksum strategy; multipart threshold; and deletion behavior. No objects are uploaded.
The evidence package must contain:
- the problem and final requirement in the learner's own words;
- caller type and Region with private identifiers redacted;
- exact planned values, ownership, and cost class;
- one Console observation and matching CLI or API evidence;
- one behavior result or supplied data-plane record;
- one denied, failed, or counterexample result and evidence-led diagnosis;
- one architecture choice plus the rejected alternative;
- cleanup proof or explicit retained-state owner, expiry, and next lesson.
Verification standard
Use expected state before observed state. Record timestamps in UTC and preserve the original failure before changing anything. A passing submission answers all four questions:
- What exact requirement was tested?
- Which evidence proves the AWS configuration?
- Which evidence proves the workload behavior?
- What remains unproven or requires later monitoring?
If AWS returns no rows, verify account, Region, permission, filters, pagination, resource type, and deletion state before concluding that nothing exists.
Common failures and troubleshooting
| Symptom | Evidence first | Likely boundary | Smallest safe response |
|---|---|---|---|
| object appears missing | caller, Region, filters, pagination, tags | scope or read permission | align scope before creating a duplicate |
| state remains pending or unavailable | service state, events, dependencies, quotas | dependency or capacity | correct the named dependency and wait with a bound |
| AccessDenied | principal, action, resource, explicit-deny context | identity, resource, endpoint, organization, or KMS policy | change only the proven policy layer |
| configuration exists but behavior fails | route, DNS, security, listener, health, logs, object version | data path or application | test the next boundary and change one control |
| bill is higher than expected | hours, bytes, requests, AZs, addresses, retention | cost model or retained resource | stop optional work and reconcile the ledger |
| cleanup is blocked | dependency inventory and owning service | deletion order or immutable state | remove owned dependants in reviewed reverse order |
Do not troubleshoot by attaching administrator access, opening administration ports to the internet, disabling encryption, retrying uncontrolled creation, deleting unknown resources, or weakening retention.
Cost, cleanup, and retained state
No AWS resource is created. Close CloudShell and remove or redact downloaded evidence.
Cleanup evidence requires terminal state and an after-inventory. Search related ENIs, public IPv4 addresses, EBS volumes and snapshots, load balancers, target groups, Auto Scaling instances, endpoints, logs, S3 versions and delete markers, backup recovery points, and global IAM roles when they apply. Billing data can lag, so schedule a later review.
Architecture and certification decisions
- Certification coverage: SAA-C03; SOA-C03; SAP-C02; DOP-C02.
- Exam mapping: SAA D1-D4.
- Explain service scope, failure boundary, consistency, recovery, security, operations, and price rather than matching a keyword.
- Treat availability and durability, encryption and authorization, routing and filtering, health and lifecycle, backup and replication, and discount and capacity as separate concepts.
- Do not reproduce protected certification questions.
Knowledge check
- Are folders real S3 directories?
Expected direction: No. Tools render key prefixes as folders.
- What is required above the single PUT limit?
Expected direction: Multipart upload.
- Can user-defined metadata be edited in place?
Expected direction: No. Copy the object with replacement metadata.
- Why record version ID?
Expected direction: It identifies the exact object version for recovery and audit.
Completion gate and assessment
| Area | Points | Passing evidence |
|---|---|---|
| Requirement and model | 15 | Correct scope, terminology, and final outcome |
| Console evidence | 15 | Current path and interpreted fields |
| CLI or API evidence | 15 | Scoped command, expected result, and limitations |
| Behavior or decision exercise | 20 | Reproducible result or defensible architecture reasoning |
| Troubleshooting | 15 | Original symptom, hypothesis, one change, retest, rollback |
| Security and cost | 10 | Least privilege, data protection, current price dimensions |
| Cleanup and handoff | 10 | Terminal-state proof or approved retained-state record |
Pass at 80 out of 100 with no critical safety failure. A missing practical artifact, unexplained output, unsafe access, destructive action outside the owned scope, unplanned billed resource, or false cleanup claim requires remediation and a changed retest.