AWS 108: Amazon S3 concepts
Why object storage exists: the problem before the service
Before object storage, applications commonly saved data on a server's local disks or on a shared file server. That works at small scale, but it couples valuable data to machines, filesystems, mount points, capacity planning, RAID, backup jobs, and replacement procedures. Adding another application server creates a harder question: which server owns the file, and how do all servers see the same current copy?
Object storage changes the interface. An application sends an authenticated HTTPS request containing a bucket, an object key, data, and optional metadata. The storage service handles placement and durability behind that API. The caller does not select a disk, partition, inode, RAID group, or storage host. This separation makes object storage suitable for application assets, backups, logs, data lakes, software artifacts, media, static websites, and many other datasets that are retrieved as whole objects.
Amazon Simple Storage Service (Amazon S3) launched in March 2006 as one of AWS's earliest infrastructure services. Its architectural importance is larger than “a place to upload files”: S3 helped make durable storage an API-accessible service that independent compute fleets could share. That history explains several design choices today. S3 uses buckets and keys rather than disks and paths, HTTP operations rather than normal POSIX system calls, and policy-controlled service endpoints rather than a filesystem mount as its native interface.
The problem for an architect is therefore not merely creating a bucket. The architect must choose the right bucket type and storage class, design object identity and concurrency, authorize every request path, protect and recover versions, control lifecycle and replication, observe data access, and predict charges. A bucket can exist while the workload remains insecure, unrecoverable, inaccessible, unexpectedly expensive, or semantically incorrect.
Final outcome
The learner will build a complete mental model of S3, trace an object request from caller to stored version and downstream event, distinguish all current S3 bucket types, select the correct protection and access controls, and produce a requirement-led storage design. The learner will also inspect the matching AWS control plane, run read-only CLI queries, diagnose a denied or missing-object case, defend one architecture choice, and prove cleanup or approved retained state.
The practical outcome is not a command transcript. It must show what was expected, what happened, what the result proves, what it does not prove, and which evidence would change the decision.
Learning objectives
By the end of this lesson, the learner can:
- explain why object storage developed and how its API differs from block and file storage;
- define bucket, object, key, prefix, metadata, tag, version ID, delete marker, ETag, checksum, storage class, endpoint, access point, and replication;
- distinguish general purpose, directory, table, and vector buckets and avoid applying one bucket type's behavior to another;
- trace the identity, TLS endpoint, authorization, data-protection, storage, event, monitoring, and billing boundaries of an S3 request;
- explain regional placement, namespaces, strong consistency, object-level atomicity, concurrency, durability, and availability;
- separate encryption, authorization, public-access prevention, immutability, version recovery, backup, and replication;
- identify every major S3 price dimension and cleanup obligation before creating data;
- connect control-plane state to the real data, network, identity, or application behavior;
- identify cost and cleanup ownership before any optional mutation;
- troubleshoot from evidence without opening broad access or adding broad permissions.
Complete S3 request and data relationship model
Human, application, or AWS service
|
| credentials + signed HTTPS request (or presigned request)
v
DNS -> S3 endpoint / access point / Multi-Region Access Point
|
v
Authentication -> authorization policy evaluation -> Block Public Access guardrail
IAM identity bucket/access-point policy ACL if still enabled
session policy VPC endpoint policy Organizations controls
permissions boundary KMS key policy/grants when required
|
v
bucket type + Region/AZ scope -> key or table/index identity -> optional version ID
|
v
encryption + integrity check -> atomic object operation -> response/version/checksum
|
+--> lifecycle / storage-class transition / expiration
+--> replication / Batch Replication / backup
+--> event notification -> SQS, SNS, Lambda, or EventBridge
+--> CloudTrail data event / server access log / metrics
+--> Inventory / Storage Lens / analytics and cost records
Every arrow is a verification boundary. A successful CreateBucket call proves only control-plane creation. It does not prove that the intended principal can upload, that an application can resolve and reach the endpoint, that a KMS key permits decrypt, that a replica meets its recovery objective, or that lifecycle and retention preserve data for the required period.
Prerequisites, permissions, Region, and safety
- Learning baseline: This sequence assumes practical Linux knowledge but no prior cloud-computing or AWS knowledge. Cloud, networking, security, data, automation, and architecture concepts must come from completed earlier lessons. If a prerequisite checkpoint is incomplete, return to its linked lesson before continuing.
- Confirm a non-root caller with
aws sts get-caller-identityand keep the account number private. - Use
ap-south-1unless this lesson explicitly names a second Region. - Confirm the intended profile and Region with
aws configure listbefore interpreting an empty result. - Use read-only List, Get, and Describe permissions for the named services. Design exercises run locally and require no resource-creation permission.
- This is a no-create lesson. Console and CLI work is read-only, and every design artifact is created locally.
- Never publish account IDs, public addresses, ARNs containing private account data, session IDs, presigned URLs, object data, credentials, or KMS material.
- Do not use root, world-open SSH or RDP, disabled TLS verification, unowned resources, or irreversible retention controls in a training exercise.
Vocabulary: build the model before memorizing features
| Term | Precise meaning and design consequence |
|---|---|
| Bucket | The top-level S3 resource and policy/configuration boundary. A bucket has a type and placement scope. It is not a folder and does not contain a mounted filesystem. |
| Object | For a general purpose bucket, the stored value plus system and user metadata, identified by bucket and key and, when versioning is used, version ID. A normal object update replaces the whole object value atomically; applications should not assume in-place byte editing. |
| Key | The object's complete name. customers/42/photo.jpg is one key, not three real nested directories. The / characters support prefix grouping and console presentation. A key is case-sensitive. |
| Prefix and delimiter | A prefix filters keys that begin with the same characters; a delimiter groups results during listing. Prefix design affects organization, IAM conditions, lifecycle filters, event filters, and operations - not modern general-purpose-bucket request partitioning in the simplistic old “randomize every prefix” sense. |
| System metadata | S3-controlled facts such as creation time, size, storage class, encryption, and checksum information. Some metadata can be changed only by copying/replacing the object. |
| User-defined metadata | Key-value data supplied with an object. It is returned with object metadata but is not a general database query index. |
| Object tag | A separate key-value label set that policy, lifecycle, replication, and cost processes can use. Tags and metadata are not interchangeable. |
| Version ID | The identifier S3 assigns to one version when versioning is enabled. Bucket + key alone means “current version”; bucket + key + version ID identifies an exact stored version. |
| Delete marker | A versioned tombstone produced by a simple delete in a versioned bucket. It can make the current key appear deleted without removing older versions. Permanently deleting a specific version is different. |
| ETag | A response identifier often useful for cache or conditional-request logic, but not universally the MD5 hash of object data. Multipart upload and encryption can change its meaning. Use supported checksum fields when cryptographic integrity proof is required. |
| Checksum | An integrity value using a supported algorithm. It can detect corruption in transit and support end-to-end validation; understand whether the client, SDK, or service calculates and verifies it. |
| Storage class | The storage and retrieval/cost behavior assigned to an object. Storage class is an object-level economic and access decision, not a bucket type. |
| Endpoint | The DNS/API destination used by a request. Regional, zonal, access-point, website, dual-stack, FIPS, VPC endpoint, and acceleration paths have different capabilities and security properties. |
| Access point | A named access endpoint with its own policy that simplifies controlled access to shared datasets. It does not erase the bucket policy or other applicable guardrails. |
| Lifecycle rule | Declarative automation that transitions or expires eligible current versions, noncurrent versions, delete markers, or incomplete multipart uploads. It is not a backup. |
| Replication | Asynchronous copying under a configured IAM role to another bucket in the same or a different Region. Versioning is required for ordinary live replication. Replication is not the same as request routing or an application transaction. |
| Object Lock | Write-once-read-many protection for versioned objects using retention periods or legal holds. Governance and compliance modes have materially different bypass behavior. It is not equivalent to encryption. |
The four current S3 bucket types
Do not treat “S3” as one bucket implementation with identical APIs. Select the type before applying feature assumptions.
| Bucket type | Scope and data model | Choose it for | Important boundaries |
|---|---|---|---|
| General purpose | Regional object bucket; the original and broadest S3 bucket type. It supports the normal object API and all general-purpose storage classes except S3 Express One Zone. | Most object storage, data lakes, logs, backups, application assets, archives, and static origin data. | Names can use the shared partition-wide namespace or the newer account Regional namespace where supported. Verify current namespace, feature, and endpoint compatibility. |
| Directory | Zonal bucket built for S3 Express One Zone and hierarchical prefixes. | Performance-sensitive workloads needing single-digit millisecond access and very high request rates where one-AZ placement is acceptable. | Different naming, zonal endpoints, authorization sessions, API behavior, storage-class choice, and feature support. Do not promise multi-AZ resilience by habit. |
| Table | Regional purpose-built storage for tabular datasets represented as Apache Iceberg tables and namespaces. | Analytics tables queried through compatible engines such as Athena, Redshift, and Spark, with S3-managed maintenance. | Tables are subresources with S3 Tables APIs, policies, maintenance, quotas, and integration behavior - not arbitrary objects managed like a normal bucket. |
| Vector | Regional purpose-built storage for vector indexes, vectors, and similarity queries. | Cost-oriented vector storage for semantic search, recommendation, RAG, and AI workloads that fit its latency and query model. | Uses dedicated S3 Vectors APIs and vector-index configuration. Dimensions, distance metric, metadata filtering, quotas, encryption, and supported integrations must be designed explicitly. |
New services evolve. General purpose buckets remain the default teaching path in AWS 108–115 and the live lab in AWS 120. Directory, table, and vector buckets are introduced here so an architect recognizes that their scope and APIs differ; current feature support must be checked before production design.
Object storage is not a Linux filesystem
A Linux filesystem offers directories, byte-range updates, file descriptors, append behavior, rename operations, permissions attached to filesystem objects, and often POSIX locking semantics. Native S3 offers HTTP API operations against object keys. The console may draw folders and tools may present S3 through a filesystem-like interface, but presentation does not create full POSIX semantics.
Three consequences matter immediately:
- Renaming a general purpose bucket object is normally implemented as copy plus delete. It is not a metadata-only filesystem rename.
- Updating a normal object means putting or copying a replacement object. Readers see a complete old or complete new object, but a workflow that changes many keys is not one atomic transaction.
- An application that requires shared locking, frequent small random writes, or POSIX permissions should evaluate EFS, FSx, or block storage rather than assuming S3 is a drop-in mount.
Region, namespace, and data location
A general purpose bucket has one home Region. A bucket name's uniqueness scope is a naming rule; it does not mean the data is automatically global. AWS now supports both the traditional shared global namespace within an AWS partition and an account Regional namespace for eligible new general purpose buckets. The learner must inspect which namespace was selected rather than relying on an old universal statement that every bucket name is globally unique.
S3 does not copy data to another Region merely because clients access it globally. Cross-Region Replication, S3 Batch Replication, an application copy, backup copy, or another explicit mechanism creates another Regional copy. Directory buckets are zonal, so a single-AZ failure consideration is part of the type choice. Data-residency analysis must include replicas, backups, analytics exports, logs, KMS keys, and downstream event consumers - not only the source bucket.
Consistency, atomicity, and concurrent writers
General purpose bucket object PUT, COPY, and DELETE operations and related reads/listing have strong consistency: after a successful write, a subsequent read does not need an eventual-consistency delay to discover that write. Strong consistency does not create a multi-object database transaction.
Suppose two processes read version A and both upload a replacement to the same unversioned key. Each PUT can succeed, while the later completed write becomes current. The application can still lose an update. Controls include versioning, conditional requests such as If-Match or If-None-Match, checksums, application-level idempotency, object naming that avoids overwrites, and a transactional coordination service when several objects must change as one business operation.
One completed object operation is atomic: a reader does not receive half of an old object and half of its replacement. Multipart upload is a transfer protocol; the completed multipart object becomes visible after successful completion. Uncompleted parts can remain billable and should be aborted explicitly or by lifecycle automation.
Durability, availability, resilience, and recovery are different
- Durability asks whether stored data survives failures over time. S3 Standard is designed for eleven nines of durability, but that statement does not recover an authorized accidental deletion or malicious overwrite by itself.
- Availability asks whether a permitted request succeeds when needed. Storage class, Region/AZ scope, endpoint path, DNS, policies, KMS, quotas, and application retries affect it.
- Resilience asks how the workload continues or degrades through a component, AZ, or Regional event.
- Recovery asks which point in time can be restored and how long restoration takes. Versioning, Object Lock, AWS Backup, replication, archive restore, and tested runbooks address different threats.
Replication can also reproduce an unwanted change depending on configuration. Versioning preserves recoverable versions but still needs permission controls, lifecycle, monitoring, and tested recovery. Object Lock protects selected versions from deletion for a retention requirement but creates operational and legal consequences. No single checkbox means “backup, disaster recovery, ransomware protection, and compliance completed.”
Authorization: why an S3 request is allowed or denied
Evaluate the effective request, not only one policy document. Relevant layers can include:
- the caller's IAM identity policy, session policy, permissions boundary, and Organizations service control policy;
- the bucket policy or access-point policy;
- S3 Block Public Access at account, bucket, or access-point scope;
- Object Ownership and ACLs where ACLs remain enabled;
- a VPC endpoint policy and policy conditions such as source VPC endpoint, network origin, TLS, principal, prefix, object tags, or encryption headers;
- the AWS KMS key policy and grants for SSE-KMS or DSSE-KMS operations;
- Object Lock retention/legal hold for delete or overwrite attempts.
An explicit deny wins over an allow. s3:ListBucket targets a bucket ARN and can be constrained by prefix; object actions such as s3:GetObject target object ARNs. A principal may list a bucket but not read objects, read one prefix but not another, or read object metadata but fail decrypt because KMS authorization is missing. Block Public Access is a guardrail against public-policy/ACL exposure; it does not replace least-privilege private policies.
For modern general purpose buckets, prefer Bucket owner enforced Object Ownership so ACLs are disabled unless a documented integration requires them. Keep all four Block Public Access settings enabled unless a reviewed public distribution requirement proves otherwise. For public content, evaluate CloudFront with Origin Access Control so the bucket can remain private.
Encryption and integrity: separate choices
S3 encrypts new uploads at rest by default with SSE-S3, but an architect still chooses the required control and verifies existing data and policy behavior.
| Method | Key/control model | Main decision |
|---|---|---|
| SSE-S3 | S3 manages the encryption keys | Simple default at-rest encryption with no customer KMS policy or request charge. |
| SSE-KMS | AWS KMS key protects object data keys | Adds KMS policy control and auditability; requires KMS permissions, compatible Region/key design, quota and request-cost planning. S3 Bucket Keys can greatly reduce KMS request traffic/cost. |
| DSSE-KMS | Two independent layers of server-side encryption, with KMS control | Use when a requirement calls for dual-layer server-side encryption; S3 Bucket Keys are not supported. |
| SSE-C | Caller supplies the encryption key with each applicable request over TLS; S3 does not retain that key | Specialized compatibility case with substantial client key-management responsibility. Losing the key loses access. |
| Client-side encryption | Client encrypts before upload and controls envelope/key behavior | Protects plaintext before S3 receives it, but shifts implementation, rotation, metadata, recovery, and interoperability duties to the application. |
Encryption at rest does not authorize a caller. TLS protects transport but does not prove object integrity after application processing. Use checksum validation where required, require secure transport in policy, and test both S3 permission and KMS permission paths.
The S3 feature map for this module
This lesson establishes the system. The following lessons deepen and test each part without changing the 423-page sequence:
| Lesson | Detailed responsibility |
|---|---|
| AWS 109 | Bucket creation choices, object/key/prefix design, metadata, tags, endpoints, naming, upload/download/copy/delete semantics, conditional operations, and presigned access. |
| AWS 110 | Every general purpose storage class, access/availability trade-offs, archive restore, minimum-duration/minimum-size effects, monitoring, and cost decisions. |
| AWS 111 | Versioning states, delete markers, recovery, lifecycle evaluation, transitions, expiration, noncurrent versions, and incomplete multipart upload cleanup. |
| AWS 112 | IAM and bucket policies, Object Ownership/ACLs, Block Public Access, access points, VPC endpoint controls, TLS, encryption methods, KMS, CORS, and exposure analysis. |
| AWS 113 | Same-Region and Cross-Region Replication, Replication Time Control, Batch Replication, ownership/encryption, failure metrics, Multi-Region Access Points, and failover controls. |
| AWS 114 | Request scaling, key design, multipart upload, byte ranges, transfer acceleration, networking, retries, checksums, concurrency, and transfer tools. |
| AWS 115 | Object Lock governance/compliance retention, legal holds, archive retrieval, vault-style retention decisions, evidence, and recovery drills. |
| AWS 120 | A bounded live lab that secures, versions, lifecycles, replicates, deletes, restores, validates, and completely cleans up owned S3 data. |
Related capabilities that must appear in designs and later architecture lessons include event notifications and EventBridge, CloudTrail data events, server access logging, request metrics, Storage Lens, Inventory, Batch Operations, S3 Select alternatives/current support, static website hosting versus CloudFront, CORS, access logs, AWS Backup, Storage Gateway, DataSync, Snow Family transfer options, and cost allocation.
Service-completeness contract for an AWS architect
The S3 family is not complete until the learner can reason through each of these categories. “Know the service name” is never a pass.
- Origin and fit: why object storage exists; object versus block, file, database, queue, and cache behavior.
- Resource model: all current bucket types; buckets, objects, keys, prefixes, versions, metadata, tags, checksums, endpoints, tables/namespaces, and vector indexes.
- Interfaces and transfer: Console, CLI, SDK/REST, SigV4, presigned URLs, multipart upload/copy, range requests, Transfer Acceleration, DataSync, Storage Gateway, Snow, and private endpoint paths.
- Consistency and concurrency: strong consistency, atomic object replacement, conditional requests, idempotency, concurrent writers, multipart completion/abort, and absence of multi-key transactions.
- Storage economics: Standard, Intelligent-Tiering tiers and monitoring, Standard-IA, One Zone-IA, Glacier Instant Retrieval, Glacier Flexible Retrieval, Glacier Deep Archive, Express One Zone, archive restore, minimum storage duration, minimum billable object size, request/retrieval/transition/data-transfer charges, and lifecycle break-even reasoning.
- Identity and exposure: identity/resource policies, explicit deny, cross-account access, access points, MRAP, Access Grants where applicable, Object Ownership, ACL legacy cases, Block Public Access, policy validation, and Access Analyzer.
- Network and delivery: public service endpoints, gateway/interface endpoint differences, endpoint policies, DNS, TLS-only enforcement, CloudFront Origin Access Control, website endpoint limitations, CORS, acceleration, dual-stack/FIPS requirements, and on-premises paths.
- Data protection: default encryption, SSE-S3, SSE-KMS, DSSE-KMS, SSE-C, client-side encryption, KMS key policy/grants, Bucket Keys, checksum validation, versioning, MFA Delete limitations/operations, Object Lock, legal hold, backup, and restore.
- Lifecycle and governance: rule filters and precedence, transitions, current/noncurrent expiration, expired delete markers, incomplete multipart uploads, retention conflicts, legal discovery, data classification, ownership, and deletion approval.
- Replication and multi-Region design: SRR, CRR, live versus existing objects, Batch Replication, RTC, delete-marker behavior, replica modification/ownership, KMS permissions, metrics, failure reasons, MRAP routing and failover controls, RPO/RTO, and replication cost.
- Events and automation: notification destinations, EventBridge, duplicate/out-of-order event handling, recursion prevention, S3 Batch Operations, Inventory-driven remediation, IaC configuration, policy-as-code, and safe deployment/rollback.
- Observability and audit: CloudTrail management versus data events, server access logs, CloudWatch request/storage metrics, Storage Lens, Inventory, replication metrics, KMS logs, AWS Config/Security Hub controls, ownership of alarms, and evidence retention.
- Performance: request-rate scaling, parallelism, multipart thresholds, byte-range fetches, retry/backoff, connection reuse, checksums, small-object overhead, latency-sensitive directory buckets, and measurement before optimization.
- Failure and troubleshooting:
AccessDenied, wrong Region/endpoint, missing current version, delete marker, KMS deny, policy/BPA conflict, CORS versus authorization, signature/clock errors, replication failure, archive-not-restored response, throttling/retry behavior, incomplete multipart cost, and cleanup blocked by versions or retention. - Architecture and certification: select S3 from requirements; reject close alternatives with evidence; calculate failure, security, recovery, operational, and cost consequences; build and test with least privilege; then explain what remains unproven.
Completeness does not mean memorizing every REST header or obscure API parameter. It means no concept needed to understand, build, secure, automate, operate, troubleshoot, recover, price, compare, or select the service at the target architect level is silently omitted.
How a general purpose bucket request works
An S3 architecture starts with data owner, bucket type, object identity, Region, access pattern, retention, recovery point, recovery time, mutability, encryption, authorization, network path, replication, event consumers, audit evidence, and deletion obligations.
For a representative GetObject request:
- The application determines the exact endpoint, bucket/access point, key, and optional version ID. DNS resolves the endpoint and the client establishes TLS.
- The client signs the request with temporary credentials, uses a presigned signature, or invokes through an AWS service integration. Authentication identifies the principal and session context.
- AWS evaluates applicable organization, session, boundary, identity, endpoint, bucket/access-point, ACL, and Block Public Access controls. An explicit deny ends the request.
- S3 resolves the requested version. If the current version is a delete marker and no older version ID was requested, the result behaves as deleted.
- For SSE-KMS or DSSE-KMS, the KMS authorization/key state must also permit the cryptographic operation. S3 retrieves/decrypts the stored value and validates the service-side integrity path.
- S3 returns status, headers, metadata, checksum information where requested/supported, and bytes. The client must handle retries, ranges, checksums, and application validation correctly.
- Configured logs, CloudTrail data events, request metrics, and downstream application telemetry create different evidence. No one signal automatically provides a complete audit record.
- Storage, request, retrieval, transfer, KMS, acceleration, replication, analytics, and logging dimensions can generate charges depending on the chosen path.
Read the result in layers:
- Scope: account, Region, VPC, bucket, AZ, endpoint, principal, object version, or resource ARN.
- Control plane: the requested configuration exists and reached an expected state.
- Behavior: the request, connection, health check, replication, restore, or application result meets the requirement.
- Operations: monitoring, failure owner, cost, retention, rollback, and cleanup are known.
Control-plane success is necessary but not sufficient. A resource can be available while policy, routing, DNS, health, data, or application behavior remains wrong.
Architecture decision table
| Requirement | Preferred direction | Why |
|---|---|---|
| Durable application objects accessed by key | S3 general purpose bucket | The object API, lifecycle, versioning, policy, and storage classes fit. |
| Shared POSIX file access from Linux instances | Evaluate EFS | S3 object semantics do not provide normal filesystem locking and rename behavior. |
| Low-latency single-AZ object workload | Evaluate S3 Express One Zone directory bucket | Zonal placement and supported feature set must match the requirement. |
| Managed Apache Iceberg analytics tables | Evaluate S3 table bucket | It provides table/namespace resources and managed maintenance rather than requiring the team to treat a table as unrelated ordinary objects. |
| Cost-oriented similarity search over embeddings | Evaluate S3 vector bucket | Vector indexes and similarity-query APIs fit; validate latency, dimensional, filtering, Region, and integration requirements. |
| Need independent Regional recovery | Versioning plus tested replication and restore design | Durability inside one Region is not the same as a second-Region copy. |
Professional questions normally contain several valid services. State the requirement that selects one option, why the nearest alternative fails it, and what changed requirement would reverse the choice.
Worked examples: reason from requirements
Example 1: application images shared by an Auto Scaling fleet
Requirement: any EC2 instance may upload or retrieve a profile image; instances are disposable; images must survive instance replacement.
Choose a private general purpose bucket. Give the EC2 workload an IAM role scoped to the required object prefix, use HTTPS and default encryption, enable versioning if overwrite recovery is required, and define lifecycle/retention. Do not store the only copy on an instance-store disk or one instance's EBS volume. CloudFront may become the controlled read path when global caching and public delivery are required, while uploads remain authenticated to an API/S3 path.
What is not yet proven: image authorization between tenants, safe content validation, lifecycle economics, recovery behavior, and application handling of retries and concurrent writes. Those require tests.
Example 2: Linux application requiring shared append and file locking
Requirement: several EC2 instances open the same files, append small records, use filesystem locks, and expect POSIX rename behavior.
Do not select S3 merely because the data is called “files.” Evaluate EFS for a managed shared Linux filesystem or an appropriate FSx service when its protocol/features fit. Re-architecting the application to immutable objects may make S3 viable, but that is an application-design change, not a mount option.
The reversal condition: if the application writes complete immutable log segments and readers retrieve by object key, S3 can become the better durable destination.
Example 3: regulated records with seven-year non-deletion
Requirement: each final record must be retained for seven years, unauthorized deletion must be prevented, auditors need evidence, and a separate Region is required.
Evaluate a versioned general purpose bucket with Object Lock. Compliance mode may fit a true non-bypass requirement, whereas governance mode permits specifically authorized bypass. Add least-privilege write/read roles, encryption based on the key-control requirement, replication designed for locked objects, audit events, inventory/evidence, and a tested retrieval process. Confirm retention dates and legal-hold procedures with the data owner and legal/compliance authority before creating irreversible settings.
Versioning alone fails the non-deletion requirement; replication alone can copy bad writes/deletes depending on configuration; encryption alone does not prevent deletion.
Example 4: daily telemetry queried as changing SQL tables
Requirement: analysts use SQL; schemas and partitions evolve; table-file compaction and snapshot maintenance are an operational burden.
Evaluate an S3 table bucket and its Apache Iceberg integration with the intended query engines. A normal general purpose bucket can store Iceberg files, but the team would own more table maintenance. Validate Region availability, engine compatibility, IAM/resource policies, quotas, replication/recovery, and pricing before choosing the newer bucket type.
Cost model: calculate requests as well as bytes
Never copy a numeric price into a permanent lesson and assume it remains current. Use the current S3 pricing page for the selected Region, then calculate from measured or estimated units.
monthly estimate = storage GB-month by class
+ PUT/COPY/POST/LIST and transition requests
+ GET/SELECT and other request classes
+ retrieval GB and archive restore requests
+ internet, inter-Region, acceleration, or other transfer
+ replication destination storage and replication requests
+ Intelligent-Tiering monitoring where applicable
+ Inventory, Storage Lens advanced, Batch Operations, or analytics
+ KMS API usage and logging/monitoring destinations
+ early-deletion or minimum-object-size effects
Worked reasoning: one million 10-KiB objects total only about 9.54 GiB, but lifecycle transitions can create one million transition requests and some classes apply a minimum billable object size. “Very little data” therefore does not mean “negligible S3 bill.” Conversely, one 100-GiB object has low object count but can create meaningful retrieval and transfer charges. Always model object count, average size, request frequency, retention, retrieval probability, and transfer path together.
AWS Management Console guided practice
Before opening a service page, write the expected account, Region, starting state, and evidence. Do not choose Create, Save, Purchase, Lock, or Delete unless the lesson explicitly authorizes the live track.
- Open Amazon S3, General purpose buckets in ap-south-1, and distinguish account bucket inventory from the Region stored on each bucket.
- Select only an owned or instructor-supplied bucket and inspect Properties, Permissions, Management, Metrics, and object Version ID fields without changing them.
- Classify ten workload statements as object, block, or file storage and defend the S3 choices by API and consistency behavior.
For each step, capture the field name and value in text. A screenshot may support the record but does not replace the explanation. Console labels can evolve, so use the service search and current documentation if a navigation label differs.
CloudShell and AWS CLI practice
CloudShell is the default browser-based command environment taught in AWS 028. AWS 029 and AWS 030 cover local CLI installation and authentication. This lesson therefore does not assume that an unconfigured local shell is ready.
Start every session with:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account portion of the ARN before sharing. Then perform the topic query:
List owned general purpose buckets, then query the Region and security posture only for an instructor-approved bucket.
NW_BUCKET="replace-with-owned-bucket-name"
aws s3api list-buckets --query 'Buckets[].{Name:Name,Created:CreationDate}' --output table
aws s3api get-bucket-location --bucket "$NW_BUCKET"
aws s3api get-public-access-block --bucket "$NW_BUCKET"
Expected interpretation:
A bucket list is account-scoped but not proof of object access, encryption mode, versioning, public exposure, data residency, lifecycle, or recovery. Never query or modify an unowned bucket name.
Replace every replace-with-... sample value before running its command, and use only an explicitly owned resource. Explain each option first. These queries are read-only; a successful response does not authorize a later create or delete operation.
Practical work
Create p06-data-requirements.md for a protected document dataset. Define object key scheme, source Region ap-south-1, replica Region ap-southeast-1, data owner, expected object sizes, read/write frequency, retention, RPO, RTO, encryption, public-access prohibition, identity boundary, lifecycle, replication, restore test, evidence, and final deletion. No bucket is created.
The evidence package must contain:
- the problem and final requirement in the learner's own words;
- caller type and Region with private identifiers redacted;
- exact planned values, ownership, and cost class;
- one Console observation and matching CLI or API evidence;
- one behavior result or supplied data-plane record;
- one denied, failed, or counterexample result and evidence-led diagnosis;
- one architecture choice plus the rejected alternative;
- cleanup proof or explicit retained-state owner, expiry, and next lesson.
Verification standard
Use expected state before observed state. Record timestamps in UTC and preserve the original failure before changing anything. A passing submission answers all four questions:
- What exact requirement was tested?
- Which evidence proves the AWS configuration?
- Which evidence proves the workload behavior?
- What remains unproven or requires later monitoring?
If AWS returns no rows, verify account, Region, permission, filters, pagination, resource type, and deletion state before concluding that nothing exists.
Common failures and troubleshooting
| Symptom | Evidence first | Likely boundary | Smallest safe response |
|---|---|---|---|
| object appears missing | caller, Region, filters, pagination, tags | scope or read permission | align scope before creating a duplicate |
| state remains pending or unavailable | service state, events, dependencies, quotas | dependency or capacity | correct the named dependency and wait with a bound |
| AccessDenied | principal, action, resource, explicit-deny context | identity, resource, endpoint, organization, or KMS policy | change only the proven policy layer |
| configuration exists but behavior fails | route, DNS, security, listener, health, logs, object version | data path or application | test the next boundary and change one control |
| bill is higher than expected | hours, bytes, requests, AZs, addresses, retention | cost model or retained resource | stop optional work and reconcile the ledger |
| cleanup is blocked | dependency inventory and owning service | deletion order or immutable state | remove owned dependants in reviewed reverse order |
Do not troubleshoot by attaching administrator access, opening administration ports to the internet, disabling encryption, retrying uncontrolled creation, deleting unknown resources, or weakening retention.
Cost, cleanup, and retained state
No AWS resource is created. Close CloudShell and remove or redact downloaded evidence.
Cleanup evidence requires terminal state and an after-inventory. Search related ENIs, public IPv4 addresses, EBS volumes and snapshots, load balancers, target groups, Auto Scaling instances, endpoints, logs, S3 versions and delete markers, backup recovery points, and global IAM roles when they apply. Billing data can lag, so schedule a later review.
Architecture and certification decisions
- Certification coverage: SAA-C03; SOA-C03; SAP-C02; DOP-C02.
- Exam mapping: SAA D1-D4.
- Explain service scope, failure boundary, consistency, recovery, security, operations, and price rather than matching a keyword.
- Treat availability and durability, encryption and authorization, routing and filtering, health and lifecycle, backup and replication, and discount and capacity as separate concepts.
- Do not reproduce protected certification questions.
Knowledge check
- Is S3 a POSIX filesystem?
Expected direction: No. It is an object service.
- Does global bucket-name uniqueness make bucket data global?
Expected direction: No. A bucket is created in a Region.
- Does eleven nines durability equal constant availability?
Expected direction: No. Durability and availability are different requirements.
- What identifies a versioned object precisely?
Expected direction: Bucket name, object key, and version ID.
Completion gate and assessment
| Area | Points | Passing evidence |
|---|---|---|
| Requirement and model | 15 | Correct scope, terminology, and final outcome |
| Console evidence | 15 | Current path and interpreted fields |
| CLI or API evidence | 15 | Scoped command, expected result, and limitations |
| Behavior or decision exercise | 20 | Reproducible result or defensible architecture reasoning |
| Troubleshooting | 15 | Original symptom, hypothesis, one change, retest, rollback |
| Security and cost | 10 | Least privilege, data protection, current price dimensions |
| Cleanup and handoff | 10 | Terminal-state proof or approved retained-state record |
Pass at 80 out of 100 with no critical safety failure. A missing practical artifact, unexplained output, unsafe access, destructive action outside the owned scope, unplanned billed resource, or false cleanup claim requires remediation and a changed retest.
Official sources
- AWS origins and the problem infrastructure services were built to solve
- Amazon S3 overview, current bucket types, resources, and consistency model
- General purpose bucket namespaces
- Directory buckets and S3 Express One Zone
- S3 Tables and table buckets
- S3 Vectors and vector buckets
- S3 storage classes
- S3 security best practices
- S3 Block Public Access
- S3 Bucket Keys for SSE-KMS
- Amazon S3 pricing