AWS 208: AWS CloudTrail
Why this lesson matters
CloudTrail records AWS API activity so an operator can ask who used which identity, called which action, from where, against which resource, and whether AWS accepted the request. It is an audit source, not an application transaction log. A successful API event may precede asynchronous failure, and an absent event may reflect the wrong Region, time, event category or selector rather than proof that nothing happened.
What you will be able to do
By the end, you can:
- distinguish Event history, trails, event categories, Insights and CloudTrail Lake;
- decode
userIdentity, session issuer, event source/name, source IP, request and error fields; - explain Region and account scope for recent Event history;
- verify a multi-Region trail, S3 delivery, log-file validation and optional CloudWatch Logs integration;
- determine when management, data and network activity events require different selectors;
- correlate an API request with resource/service completion evidence;
- design durable organization audit delivery without exposing sensitive event fields.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Use Event History and standard trails to answer who called which API, when, from where, against what resource, and with what response. |
| Scope and boundary | For AWS CloudTrail, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup. |
| Evidence of success | Success requires matching Console, command-line, behavior, monitoring, and owner evidence for AWS CloudTrail. One green status is not enough. |
| Cost model | Requests, running capacity, stored telemetry, retained history, data transfer, and connected resources must be priced for the exact design. |
| Safe rejection rule | Avoid teaching CloudTrail Lake to a new customer, assuming every data-plane action is logged, or using event presence as proof of application success. |
How the request flows
+--------------------------+
| Signed AWS API request |
+--------------------------+
|
v
+---------------------------+
| CloudTrail event record |
+---------------------------+
|
v
+-----------------------------------+
| Event History or trail delivery |
+-----------------------------------+
|
v
+-----------------------------------------------------+
| Investigation, alert, and retained audit evidence |
+-----------------------------------------------------+
Event sources and retention, deeply
Event history exists automatically and provides the latest 90 days of management events for one account and one Region. It is separate from trails and event data stores, cannot aggregate an organization, searches only one attribute plus time, and does not show data, Insights or network activity events. Viewing Event history and using lookup-events have no CloudTrail charge.
A trail continuously delivers selected events to S3. Prefer an organization, multi-Region trail for governed audit coverage; enable log-file integrity validation, protect the destination bucket with least privilege and lifecycle/retention controls, and monitor delivery. CloudWatch Logs integration can support metric filters and alarms but is a second retained copy with separate permissions and cost. S3 Object Lock or isolated backup controls may be needed when evidence immutability is a requirement.
Event categories answer different questions:
| Category | Example | Selection boundary |
|---|---|---|
| management | change a security group or IAM role | Event history includes recent management events; trail read/write selectors control delivery copies |
| data | GetObject, Lambda invocation, DynamoDB item call | high-volume; explicitly select supported resources/events |
| Insights | unusual write/error/API-rate activity | requires eligible trail/event data store configuration and baseline |
| network activity | supported VPC endpoint API access outcomes | configure supported sources/selectors; not a packet-flow replacement |
Read eventTime, eventSource, eventName, awsRegion, sourceIPAddress, userAgent, requestParameters, responseElements, errorCode, errorMessage, requestID, eventID, readOnly and resources. For an assumed role, trace userIdentity.arn, principal ID and sessionContext.sessionIssuer; the displayed username alone can hide the real role/session chain. Event payloads can contain resource names and request fields, so redact deliberately without destroying the fields needed for investigation.
Current service change: CloudTrail Lake closed to new customers on May 31, 2026. Existing customers can continue using it under the documented transition, but new-customer architecture should use trails/S3 and evaluate CloudWatch-based analysis rather than promising a new Lake deployment.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use Event History for recent regional management events and standard multi-Region trails for durable governed audit delivery. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid teaching CloudTrail Lake to a new customer, assuming every data-plane action is logged, or using event presence as proof of application success. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Open CloudTrail > Event history in
ap-south-1; clear the default read-only filter when investigating both reads and writes. Set an absolute UTC range and one exact filter value. - Open the supplied security-group change. Expand the JSON and identify actor/session issuer, API, Region, source, request parameters, response/error, request ID and event ID.
- Switch Regions and explain why the Regional event list changes. Do not conclude absence until every candidate Region and the event's service behavior are checked.
- Open Trails and inspect the governed trail: home Region, multi-Region/organization flags, logging state, S3 destination/prefix, KMS key, validation, CloudWatch Logs and event selectors.
- Inspect the S3 bucket policy, public block, versioning, retention/lifecycle and most recent delivered digest/log object without downloading unrelated records.
- This lesson is read-only. Do not start/stop logging, alter selectors or delete audit data.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws cloudtrail describe-trails --include-shadow-trails false --query 'trailList[].{Name:Name,MultiRegion:IsMultiRegionTrail,Validation:LogFileValidationEnabled,S3:S3BucketName,CloudWatch:CloudWatchLogsLogGroupArn}' --output table
aws cloudtrail lookup-events --max-results 50 --query 'Events[].{Time:EventTime,Name:EventName,User:Username,Source:EventSource}' --output table
aws cloudtrail get-trail-status --name replace-with-approved-trail \
--query '{Logging:IsLogging,LatestDelivery:LatestDeliveryTime,DeliveryError:LatestDeliveryError,LatestDigest:LatestDigestDeliveryTime,DigestError:LatestDigestDeliveryError}' --output json
aws cloudtrail get-event-selectors --trail-name replace-with-approved-trail --output json
Expected interpretation
describe-trails proves configuration; get-trail-status proves recent delivery state; get-event-selectors proves intended coverage. lookup-events returns Event history metadata and embeds the full event as a JSON string, so use the full record for identity and request fields. None proves the downstream service completed asynchronous work - correlate the request ID and service events/state.
Practical work
Investigate a supplied security-group change. Find event time, principal, session issuer, source address, user agent, request parameters, response, related change, and rollback. State which data events need explicit configuration.
Diagnose this topic from its own evidence
| Symptom | First checks | Correction |
|---|---|---|
| expected event absent | Region, 90-day window, management vs data category, exact event name and selector | query correct scope/source; repair future selectors without fabricating history |
| actor appears unknown | assumed-role ARN, session issuer, principal ID and federation source | reconstruct identity chain and preserve session name/context |
| trail says configured but S3 is stale | IsLogging, latest delivery/error, bucket/KMS policies and destination ownership | repair the specific delivery permission/path |
| API event says success but resource failed | asynchronous service events/status and same request ID | diagnose service completion separately |
| log integrity questioned | digest delivery, validation chain and protected evidence copy | validate files and investigate missing/altered chain |
Positive test: locate the supplied revoke-security-group event and match every request field to the changed rule. Negative test: search the wrong Region, record the expected absence, then correct scope. Dependency-failure test: a trail cannot write to its S3 bucket; prove CloudTrail delivery failure separately from API event capture.
Cost and cleanup
Event history lookup is free, while additional trail copies, data/network activity events, Insights, S3 storage/requests, KMS, CloudWatch Logs ingestion/retention/querying, notifications and cross-Region copies can charge. High-volume data event selectors must be intentionally scoped. Audit cleanup follows retention/legal policy: never delete a shared trail, bucket, KMS key or evidence because a learning session ended.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Use Event History and standard trails to answer who called which API, when, from where, against what resource, and with what response.
- Which scope or ownership boundary must be proved first?
Expected direction: For AWS CloudTrail, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup.
- What evidence is strong enough to accept the result?
Expected direction: Success requires matching Console, command-line, behavior, monitoring, and owner evidence for AWS CloudTrail. One green status is not enough.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid teaching CloudTrail Lake to a new customer, assuming every data-plane action is logged, or using event presence as proof of application success.
- Which cost dimensions and retained resources need an owner?
Expected direction: Requests, running capacity, stored telemetry, retained history, data transfer, and connected resources must be priced for the exact design.
Lesson acceptance
- One management event is decoded from principal/session through request, response/error and resource.
- Wrong-Region absence is demonstrated and corrected without claiming no event occurred.
- Trail configuration, logging status, selector coverage, latest delivery and digest evidence agree.
- Management, data, Insights and network activity boundaries are distinguished.
- CloudTrail Lake's post-May-31-2026 new-customer restriction is stated, and durable trail/S3 ownership and cost are documented.