AWS 195: Storage and data-transfer cost optimization
Why this lesson matters
Optimize from lifecycle and packet path: class, retention, requests, IOPS, throughput, snapshots, retrieval, AZ crossings, internet egress, NAT, and endpoints.
What you will be able to do
By the end, you can:
- explain storage and data-transfer cost optimization in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Optimize from lifecycle and packet path: class, retention, requests, IOPS, throughput, snapshots, retrieval, AZ crossings, internet egress, NAT, and endpoints. |
| Scope and boundary | The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Storage and data-transfer cost optimization. |
| Evidence of success | Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Storage and data-transfer cost optimization. |
| Cost model | Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design. |
| Safe rejection rule | Avoid deleting backups from age alone, moving hot data to archival class, or centralizing traffic without cross-AZ analysis. |
How the request flows
+-----------------------------+
| Data age and traffic path |
+-----------------------------+
|
v
+--------------------------------------+
| Storage or network price dimension |
+--------------------------------------+
|
v
+-----------------------------------------+
| Safe lifecycle or architecture change |
+-----------------------------------------+
|
v
+-----------------------------------------------+
| Performance, restore, and cost verification |
+-----------------------------------------------+
Optimize from the byte and request path
Storage price per GB is only one term. Total cost can include provisioned capacity, requests, IOPS/throughput, retrieval, early deletion, lifecycle transitions, replication, backup copies and data transfer. Draw where each byte is created, stored, read, copied and deleted, including AZ/Region/account/internet boundaries.
Storage decisions
| Workload characteristic | Direction and caveat |
|---|---|
| Unknown/changing object access | S3 Intelligent-Tiering where object size, monitoring/automation charge and access-tier behavior fit |
| Predictable colder objects | Lifecycle to Standard-IA/One Zone-IA/Glacier classes after modeling minimum duration, minimum billable size, retrieval and restore time |
| EBS general SSD | gp3 lets size, IOPS and throughput be adjusted independently within limits; rightsize all three |
| High sustained IOPS/latency requirement | io2 family after measurement; provisioned performance costs even when idle |
| Shared Linux files | EFS lifecycle/throughput/storage-class selection with mount/IA access behavior measured |
| Backup/retention | AWS Backup/native lifecycle, incremental behavior, copies, cold tier and immutable retention with restore tests |
Deleting a source volume does not delete retained snapshots; AMIs, orphan snapshots, unattached volumes, old RDS snapshots and backup vault recovery points need ownership. Snapshot size is changed blocks, not necessarily volume provisioned size. Never delete from age alone - map dependency, retention and restore evidence.
Lifecycle transitions can cost more for tiny/short-lived objects due to per-request charges, minimum storage duration and minimum billable object size. Model object count and churn, not just TB. Compression, columnar formats and partition pruning can reduce both storage and scan/query transfer, but require workload-compatible CPU and retrieval behavior.
Data-transfer topology
Common cost boundaries include internet egress, inter-Region transfer, cross-AZ paths, NAT Gateway hourly/data processing, interface endpoint hourly/data processing, Transit Gateway processing, replication and CDN viewer/origin transfer. Prices and exceptions are service/direction/Region specific; verify current pages.
Optimization patterns:
- Keep chatty application/database tiers in the same AZ only when that does not undermine required AZ resilience; otherwise budget cross-AZ traffic.
- Deploy zonal NAT gateways and route each private subnet to its local NAT to avoid cross-AZ hairpinning; compare NAT with gateway endpoints for S3/DynamoDB and interface endpoints for supported services.
- Use CloudFront caching/compression/origin shielding when it reduces origin requests/bytes and meets freshness/security needs.
- Avoid repeatedly moving bulk data across Regions; compute near data where latency, sovereignty and resilience allow.
- Inspect load balancer cross-zone behavior, Kubernetes topology, EFS mount targets and centralized inspection/TGW paths.
A VPC endpoint is not automatically cheaper: compare hourly endpoints per AZ, bytes and operational/security value against NAT path. A centralized inspection architecture can add TGW, firewall and cross-AZ processing by design; calculate both directions.
Measurement and safe change
Use Cost Explorer/CUR usage types to locate spend, then correlate with CloudWatch bytes/requests, S3 Storage Lens, EBS metrics, VPC Flow Logs and resource inventory. Test one change, measure performance/durability/RTO effects, and preserve rollback. Cost optimization that violates retention or removes recovery capacity is an architecture failure.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use lifecycle, right-sizing, compression, locality, and private service paths when evidence proves savings without breaking recovery or performance. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid deleting backups from age alone, moving hot data to archival class, or centralizing traffic without cross-AZ analysis. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Use the Console service search and open S3 Storage Lens or Metrics, EC2 volumes and snapshots, VPC NAT gateways and endpoints, Cost Explorer; confirm the account and Region before reading the page.
- Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
- Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
- Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws s3api list-buckets --query 'Buckets[].Name' --output table
aws ec2 describe-volumes --query 'Volumes[].{Type:VolumeType,GiB:Size,IOPS:Iops,Throughput:Throughput,State:State}' --output table
aws ec2 describe-snapshots --owner-ids self --query 'Snapshots[].{Id:SnapshotId,GiB:VolumeSize,State:State,Started:StartTime}' --output table
aws ec2 describe-nat-gateways --query 'NatGateways[].{Id:NatGatewayId,State:State,Subnet:SubnetId}' --output table
Expected interpretation
Inventory supplies resource quantities. Real optimization requires access and age evidence, restore need, request profile, network flow, performance, owner, and safe deletion or migration tests.
Practical work
Analyze ten cost findings: stale snapshot, oversized gp3, log retention, S3 lifecycle, retrieval spike, cross-AZ NAT, missing S3 endpoint, duplicate data, idle public IPv4, and global transfer. Rank savings against risk.
Diagnose this topic from its own evidence
- S3 bill rises after lifecycle: inspect transition/retrieval requests, minimum duration/size, object count and restore pattern.
- NAT dominates: identify source/destination bytes in flow logs, local-AZ routing and endpoint/CDN alternatives.
- Data transfer usage type unexplained: map service, direction, source/destination AZ/Region and both legs of intermediaries.
- Snapshot cost persists: inventory AMI/backup dependencies, incremental lineage and retention before deletion.
- Cheaper tier harms application: compare retrieval latency/fees, throughput and RTO against requirement and roll back.
Cost and cleanup
Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Optimize from lifecycle and packet path: class, retention, requests, IOPS, throughput, snapshots, retrieval, AZ crossings, internet egress, NAT, and endpoints.
- Which scope or ownership boundary must be proved first?
Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Storage and data-transfer cost optimization.
- What evidence is strong enough to accept the result?
Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Storage and data-transfer cost optimization.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid deleting backups from age alone, moving hot data to archival class, or centralizing traffic without cross-AZ analysis.
- Which cost dimensions and retained resources need an owner?
Expected direction: Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Lesson acceptance
- Produce a per-byte/per-request architecture with every charging boundary.
- Model lifecycle using object count, size, duration, transitions, retrieval and restore time.
- Right-size EBS/storage performance separately from capacity.
- Compare NAT, gateway/interface endpoints, CDN and centralized-network paths quantitatively.
- Implement a measured reversible optimization without weakening durability, security or recovery.