AWS 314: AWS Cost and Usage Report
Why this lesson matters
AWS billing dashboards answer common questions, but finance and engineering eventually need line-item evidence: which account, service, resource, usage type, pricing term, credit, tax, Savings Plan, or Reservation produced a number? AWS Cost and Usage data is the detailed source for allocation, anomaly investigation, commitment analysis, showback, and invoice reconciliation.
Cost and Usage Report 2.0 through AWS Data Exports is now the recommended detailed export. It provides a consistent schema, nested structures, additional identity/account fields, SQL-based column/row selection, S3 delivery, and integrations. Legacy CUR remains relevant to existing pipelines but has a dynamic schema and separate API. A learner must identify which generation exists before changing queries.
Billing data is not a real-time ledger. Current-month files are refreshed and can change as usage, refunds, credits, taxes, support, commitment allocation, and invoice finalization arrive. A query can be syntactically correct and financially wrong by double-counting files, summing the wrong cost column, mixing currencies, ignoring line-item types, or treating an estimated month as final.
What you will be able to do
By the end, you can:
- explain payer/management account, linked/member account, billing period, invoice, line item, usage, rate, cost, credit, and amortization;
- distinguish Data Exports CUR 2.0, legacy CUR, Cost Explorer, Billing Conductor, FOCUS, dashboards, and invoices;
- design a secure S3 delivery bucket and export resource;
- understand CUR 2.0 fixed/nested schema and legacy dynamic columns;
- select report granularity, file format, compression, overwrite/version behavior, and integration path;
- interpret line-item types and unblended, blended, net, amortized, and effective cost;
- query Parquet data with Glue/Athena without duplicate or partial-period errors;
- allocate by account, resource, tag, cost category, service, and workload;
- reconcile export totals to billing views and explain legitimate timing differences;
- detect missing files, stale partitions, schema drift, duplicate rows, tag gaps, and commitment allocation errors;
- control access to sensitive financial and usage identity data; and
- build a governed P19 cost-evidence pipeline.
Before you start
- Billing data is sensitive. It can expose account names/IDs, resource IDs, usage patterns, discounts, contracts, tags, principal identifiers, and business activity.
- The default T0 path is read-only inventory and local design. Do not create or change a management-account export without billing-owner approval.
- Data Exports is managed from the appropriate billing scope. A member account cannot assume it sees organization-wide data.
- Redact account IDs, invoice IDs, payer names, resource IDs, rates, credits, principals, bucket names, and commercial discounts.
- If creating a T1 export, use an approved dedicated bucket, KMS/access design, lifecycle plan, query workgroup, budget, and cleanup owner.
mkdir -p "$HOME/aws314-evidence"
cd "$HOME/aws314-evidence"
export AWS_DEFAULT_REGION="us-east-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Billing APIs are often operated from a management/billing context and can use us-east-1 endpoints, while exported resources reside in the configured S3 Region. Verify current endpoint behavior rather than copying the example blindly.
1. Build the billing-data mental model
AWS service usage and commercial adjustments
-> billing processing
-> CUR 2.0 table in AWS Data Exports
-> scheduled export
-> customer S3 prefix
-> Glue/Athena or warehouse/BI
-> governed cost views, reconciliation, allocation, decisions
The management or payer account receives the consolidated bill for organization member accounts under consolidated billing. A line item identifies billing-period, account, product, usage, pricing, and cost dimensions. Resource-level IDs appear only for services and configurations that support them and may require selecting resource identifiers.
The billing period is usually calendar-month based, but data can be updated after usage occurs. Invoice scope and billing period are related but not identical in every commercial situation. Never join or reconcile solely on month without account, currency, invoice/finalization, and line-item semantics.
2. Select the correct cost data product
| Need | Direction | Boundary |
|---|---|---|
| Detailed recurring line-item export | CUR 2.0 in Data Exports | Recommended current path; S3 data engineering remains yours. |
| Existing pipeline with legacy schema | Legacy CUR during governed migration | Dynamic schema and separate APIs; plan CUR 2.0 transition. |
| Interactive filtered cost trends | Cost Explorer | Aggregated service, not raw export replacement. |
| Standardized multi-cloud FinOps data | FOCUS export | Standard schema, but AWS-specific detail may still be needed. |
| Quick managed visualization | Cost and usage dashboard export/QuickSight | Dashboard does not replace source reconciliation. |
| Pro forma/custom billing view | Billing Conductor where applicable | Pro forma chargeback differs from AWS invoice economics. |
| Final amount owed | Invoice/billing documents | Reconcile to export; do not use an estimated CUR alone. |
Data Exports currently supports CUR 2.0, cost-optimization recommendations, FOCUS variants, carbon emissions, and a cost-and-usage dashboard path. Choose one because its schema and purpose fit, not because every export seems useful.
3. Understand CUR 2.0 versus legacy CUR
CUR 2.0 uses a fixed table schema and nests families such as product, resource tags, cost categories, and discounts into key-value structures. Data Exports can select, filter, and alias columns with its supported SQL subset. It adds fields such as payer/member account names and identity-related columns documented for CUR 2.0.
Legacy CUR columns can appear or disappear based on services, tags, cost categories, discounts, and usage. Queries using SELECT * or positional columns break when the schema evolves. Legacy CSV and Parquet tag-column normalization behavior also differs.
Migration choices:
- create CUR 2.0 using the legacy-compatible schema option for minimal pipeline change; or
- adopt the native CUR 2.0 schema and rewrite queries for nested fields and new names.
Run both for at least one closed billing period, compare totals by account/service/line-item type/currency, reconcile exceptions, then switch consumers. Dual running costs storage/query and must have an end date.
4. Design the export and S3 destination
An export definition includes name, table/query, selected columns or filters, refresh cadence, S3 destination, format, compression, and overwrite behavior according to current options.
Prefer Parquet for analytics because columnar compression reduces scan and preserves types. CSV is inspectable but larger and more prone to parsing/schema problems. ZIP/GZIP affects consumers and splittability. Record exact format and schema version in the data contract.
The S3 bucket policy must allow the billing export service to write only the required bucket/prefix under documented source-account/ARN conditions. Keep Block Public Access enabled. Use approved encryption, object ownership, versioning if required, access logging/data events according to risk, and lifecycle transitions/deletion aligned with finance retention.
Do not reuse a general application bucket. Separate:
s3://approved-cost-bucket/
raw/cur2/v1/ immutable service delivery
catalog/ crawler or schema metadata if used
query-results/ Athena workgroup output
curated/ versioned transformed tables
reconciliation/ signed/controlled close evidence
Restrict raw cost data to billing/FinOps pipelines. Analysts may receive curated views with account/resource/principal/commercial fields masked.
5. Interpret line-item types before summing
Common line-item concepts include usage charges, fees, discounts, credits, refunds, taxes, Savings Plan or Reservation-related lines, and support/other charges. Exact type values and applicable columns must come from the current dictionary.
Never write:
SELECT SUM(line_item_unblended_cost) FROM cur;
without time, currency, invoice/finalization, duplicate-file, and line-item interpretation. The number may be useful, but it has not been defined.
Cost views answer different questions:
- unblended cost: charge at the rate associated with each line;
- blended cost: consolidated-billing averaged rate behavior for applicable usage;
- net cost: includes applicable discounts after credits/discount constructs where fields are present;
- amortized cost: spreads upfront/recurring commitment fees across covered use for economic analysis;
- effective cost: commitment-covered usage calculated from allocated recurring/upfront components.
Use unblended/net for invoice-oriented questions according to contract and available fields. Use amortized/effective views for workload economics and commitment allocation. Do not mix columns across line-item types without a documented formula.
Savings Plan and Reservation analysis must distinguish commitment fees, covered usage, unused commitment, negation/discount lines, and allocation. A workload's on-demand public price is not its actual effective cost.
6. Understand important dimensions
| Dimension family | Questions |
|---|---|
| Bill | Billing period, payer, invoice, billing entity, currency? |
| Line item | Type, usage account, interval, usage type/operation/AZ, amount, normalization? |
| Product | Service/product code, Region/location, instance/storage/transfer attributes? |
| Pricing | Term, purchase option, rate code, public/on-demand rate? |
| Reservation/Savings Plan | ARN, term, covered usage, commitment, effective/amortized amount? |
| Resource | Resource ID present and meaningful for this service? |
| Tags/cost categories | Activated in time, populated, normalized, inherited? |
| Identity | Principal/user identifiers available, approved, and sensitive? |
Usage account tells where usage occurred, not necessarily which team benefits. Tags can be missing, late, mutable, unsupported, or technically correct but financially misleading. Cost categories apply rules centrally but require version/effective-date governance.
Data transfer often produces lines on multiple services or directions. Identify source/destination/operation and avoid allocating both sides as duplicate business traffic.
7. Catalog and query with Athena
Data Exports can integrate with Athena/CloudFormation according to selected path. Otherwise create a governed Glue table matching the exact format and prefix. Prefer partition projection or a reliable partition process over a crawler that repeatedly mutates financial schema without review.
Use an enforced Athena workgroup with encrypted results, scan limits, metrics, and controlled result prefix. Query only Parquet columns needed.
Illustrative native CUR 2.0 logic must be adapted to the current dictionary:
SELECT
line_item_usage_account_id,
product['product_name'] AS product_name,
line_item_line_item_type,
line_item_currency_code,
SUM(line_item_unblended_cost) AS unblended_cost
FROM finops.cur2
WHERE bill_billing_period_start_date = DATE '2026-08-01'
AND bill_billing_period_end_date = DATE '2026-09-01'
GROUP BY 1, 2, 3, 4
ORDER BY 5 DESC;
Do not copy names without checking your schema/query mode. Add:
- exact billing-period filter;
- one delivery prefix/table generation;
- currency grouping;
- exclusion/inclusion rule for credits/taxes/refunds;
- finalization status or close procedure;
- row/file count and source object manifest; and
- query version/checksum.
8. Prevent duplicate and incomplete results
Exports can refresh current-period files. Depending on configuration, delivery can overwrite or create new versions/manifest sets. Querying every historical object under a broad prefix can count multiple refresh generations.
Use the delivery manifest or documented overwrite layout, partition only approved current objects, and test uniqueness from stable line attributes where possible. Preserve S3 object version/ETag, export execution time, row count, byte count, minimum/maximum usage time, and query snapshot.
Current-month data is estimated and mutable. Define:
- intraday/daily operational view marked estimated;
- preliminary monthly close after period end;
- final close after invoice/export stabilization and reconciliation;
- reopen procedure for late credit/refund/tax adjustment.
Never silently rewrite a signed-off finance dashboard. Version the close and record adjustments.
9. Allocate cost responsibly
Allocation hierarchy can use:
- directly attributable resource/tag/account;
- usage metric tied to a consumer;
- shared-service allocation driver such as requests, users, storage, or revenue;
- fixed agreed percentage; and
- unallocated cost with owner and remediation.
Do not spread all shared cost by account spend merely because it balances. The driver should be causal, understandable, stable, and economical to collect.
Activate cost-allocation tags before expecting them in billing data. Define required keys, allowed values, case, ownership, inheritance, backfill limitations, and compliance. Untagged is a visible category, not a value to hide.
Containers and shared compute may need split cost allocation data or platform metrics. Separate infrastructure cost from allocation methodology and avoid claiming precision unsupported by measurements.
10. Reconcile to billing
Reconciliation sequence:
- lock export generation and billing period;
- confirm payer/account and currency scope;
- aggregate by line-item type and service;
- include/exclude taxes, credits, refunds, support, marketplace, and discounts according to stated objective;
- compare with Cost Explorer using matching cost metric/filter/time;
- compare invoice sections at final close;
- investigate timing, rounding, refunds, credits, exchange/currency, invoice entity, and commercial adjustments;
- document accepted differences and sign off.
Cost Explorer can have refresh timing and aggregation semantics different from export. Invoice values can include adjustments not visible when the earlier snapshot was taken. A mismatch is a diagnostic starting point, not proof one system is wrong.
11. Read-only inventory
aws bcm-data-exports list-exports --max-results 20 --output table
aws cur describe-report-definitions --output table
aws s3api get-bucket-location --bucket REDACTED
aws s3api get-bucket-policy-status --bucket REDACTED
aws glue get-databases --max-results 20 --output table
aws athena list-work-groups --output table
bcm-data-exports manages current exports; cur lists legacy report definitions. Permissions and endpoint/Region behavior differ. Read-only output proves configuration, not delivery completeness or financial accuracy.
12. Diagnose from evidence
| Symptom | Evidence | Response |
|---|---|---|
| Export has no recent files | export status, execution, bucket policy, prefix, source conditions | Correct exact service-write path and verify next delivery. |
| Athena table is empty | S3 objects, format, table location, partitions/projection, KMS | Trace raw object to table; do not recreate export blindly. |
| Query suddenly doubles | S3 keys/manifests/versions, partition scope, row keys | Remove duplicate refresh generations and add manifest control. |
| Column query breaks | CUR generation/schema, nested fields, legacy dynamic change | Version schema/query and migrate with parallel reconciliation. |
| Resource IDs missing | service support, export setting, line type, data dictionary | Allocate at supported level; do not invent IDs. |
| Tags absent | activation date, supported resource, key/value/case, usage time | Fix governance prospectively and expose unallocated historical cost. |
| Athena total differs from Cost Explorer | metric, filters, time, currency, line types, refresh | Align semantics and document timing. |
| Amortized total looks wrong | commitment fee/covered/unused/negation formulas | Rebuild using documented columns per line-item type. |
| Query cost spikes | CSV, broad columns/periods, duplicate files, no workgroup controls | Use Parquet, partition/prune, compact curated data, enforce limits. |
| Analyst sees sensitive discounts | raw access, view columns, result bucket, logs | Revoke, investigate, and publish least-privilege curated views. |
13. Build the P19 FinOps evidence pipeline
Scenario: one payer, 20 member accounts, two currencies/invoice entities, Savings Plans, Marketplace, shared networking, EKS, and 65 percent valid workload tags.
Produce:
- CUR 2.0 versus legacy migration ADR;
- payer/member/invoice/currency scope;
- Data Export query and schema contract;
- S3/KMS/bucket policy and lifecycle design;
- manifest/delivery completeness checks;
- Glue/Athena/workgroup design;
- raw-to-curated transformation with tests;
- line-item type and cost-metric formulas;
- commitment amortization/unused-allocation view;
- tag/account/cost-category/shared-cost hierarchy;
- current-month versus final-close process;
- Cost Explorer/invoice reconciliation;
- sensitive-data access and masked analyst views;
- ten failure diagnostics;
- pipeline and query cost model; and
- rollback, migration cutover, and cleanup evidence.
Acceptance tests include one known EC2 hour, one data-transfer path, one credit, one refund, one tax line, one Marketplace charge, one Savings Plan-covered use, one unused commitment, one untagged resource, and one late adjustment.
Cost and cleanup
Data Exports itself has current pricing/limits; the pipeline also incurs S3 storage/requests, KMS, Glue, Athena scanned bytes, ETL, warehouse, QuickSight, logs, transfer, backups, and engineering/finance operations. CUR data growth follows account/resource/service activity and selected columns/granularity.
For T0, retain only redacted designs. For an approved pilot, disable consumers, preserve required close evidence, delete the export only with billing approval, remove Glue/Athena/QuickSight resources, delete query results and raw objects according to retention, revoke bucket/KMS access, remove scheduled jobs, and verify no second legacy/current export remains unintentionally.
Knowledge check
- Recommended detailed export? CUR 2.0 through AWS Data Exports.
- Why is current month mutable? Usage and billing adjustments continue and invoice finalization occurs later.
- CUR 2.0 versus legacy? Fixed/nested customizable schema versus dynamic legacy schema and APIs.
- Why not sum one column blindly? Line-item types, metrics, currency, refresh generations, and scope alter meaning.
- Amortized versus unblended? Economic spreading of commitments versus charge on individual lines.
- Why can tags be absent? Activation timing, service/resource support, missing/mistyped values, and historical limits.
- How do duplicate totals arise? Multiple refreshed file generations or overlapping prefixes/partitions are queried.
- What proves monthly close? Locked source manifest, defined formula/scope, reconciliation, documented differences, and sign-off.
Lesson acceptance
Pass when a reviewer can reproduce one reported amount from export definition through exact S3 objects, schema, query, line-item formula, allocation, reconciliation, and close approval. Fail if the design uses broad raw access, confuses legacy and CUR 2.0 schemas, queries all refresh generations, ignores currency/line type, hides unallocated cost, calls an estimated month final, or cannot reconcile to billing.