AWS 260: Centralized logging and security accounts
Why this matters to an architect
Logs are needed to detect attacks, reconstruct incidents, troubleshoot systems, prove compliance, and improve services. If workload administrators can alter the only copy, if delivery silently stops, or if responders cannot query within the required time, “central logging enabled” is false assurance.
A mature multi-account design separates two duties:
- Log Archive receives and durably protects original security and operational evidence.
- Security Tooling configures delegated security services, aggregates findings, detects, investigates, and responds.
They cooperate but should not share unlimited administration. An analyst who can investigate evidence should not automatically be able to erase its source; a storage custodian should not automatically suppress findings or remediate workloads.
Outcomes
By the end, you can:
- distinguish event generation, collection, transport, archive, analytics, findings, alerts, cases, and response;
- design separate Log Archive and Security Tooling accounts with least-privilege access;
- inventory CloudTrail, Config, network, identity, security-service, infrastructure, database, and application sources;
- design organization trails, selectors, S3/KMS policies, integrity validation, Object Lock, lifecycle, and replication;
- plan delegated administrators, member enrollment, Regions, Security Hub central configuration, and aggregation;
- classify and redact sensitive telemetry while preserving forensic value;
- prove source-to-archive-to-query-to-alert delivery and detect gaps;
- size ingestion, storage, query, transfer, KMS, findings, and retention cost;
- operate incident access, legal hold, restore, export, and decommissioning.
The end-to-end evidence chain
identity / API / network / OS / database / application / SaaS sources
-> recorder or agent -> transport/buffer -> central destination
-> immutable/protected archive -> normalization/catalog/index
-> query/detection -> finding/alert -> case -> response
-> retention/legal hold -> verified restore/export/disposal
Every arrow needs an owner, authentication, encryption, retry/dead-letter behavior, health signal, quota, cost, failure test, and recovery procedure. A green S3 bucket or Security Hub dashboard proves only one segment.
Account boundaries and separation of duties
Log Archive account
Use as a tightly controlled storage sink for original or authoritative logs, signed digests, security exports, and required backups. Limit workloads and interactive users. Protect S3 bucket policies, KMS keys, Object Lock configuration, lifecycle, replication, and access logging through reviewed automation, preventive controls, and independent alarms.
Archive administrators manage durability/retention, not ordinary security analysis. Producers receive narrowly conditioned write paths and no read/delete authority. Audit and incident roles get purpose-limited, logged access. Break-glass does not mean permanent administrator access.
Security Tooling account
Use for delegated security administration, findings aggregation, detection engineering, SIEM/SOAR, investigation, ticketing, and controlled remediation. Depending on current service support, this can administer CloudTrail, Security Hub CSPM, GuardDuty, Inspector, Macie, Detective, Access Analyzer, Firewall Manager, Security Lake subscribers, and observability integrations.
Delegation is service-specific and often Region-specific. It reduces daily management-account use but grants organization-wide power. Record what remains management-account-only, service-linked roles, enabled Regions, auto-enrollment, member status, and offboarding behavior.
Why not one security account?
Combining custody, analytics, and remediation can let one compromised role delete originals, suppress alerts, alter queries, and change workloads. Separate accounts do not solve this alone: use distinct identities, KMS/key policy, bucket policy, SCP/RCP, CI/CD roles, approvals, and external evidence. Small organizations can combine functions only after documenting the residual risk and compensating controls.
Build a source and requirement matrix
For every source record:
| Dimension | Questions |
|---|---|
| purpose | audit, detection, troubleshooting, performance, legal, business analytics |
| events | management/data/network activity, authentication, configuration, packets/flows, queries, app transactions |
| coverage | accounts, OUs, Regions, resources, environments, new-account/Region behavior |
| sensitivity | credentials/tokens, request parameters, object names, IPs, identities, customer/health/payment data |
| collection | native delivery, CloudWatch Logs, Firehose, OpenTelemetry/agent, subscription, export, Security Lake |
| evidence | source event time/ID, account/Region/resource, destination object/record, checksum/digest, query result |
| SLO | expected latency, acceptable loss/duplicates/order, retention, recovery/export time |
| cost | events/bytes, scans/query, transformations, transfer, copies, KMS calls, archive retrieval |
Minimum candidates include CloudTrail management/data/network activity, AWS Config, IAM Identity Center/IdP, VPC/TGW/Route 53 flow and query logs, ELB/CloudFront/WAF/Network Firewall, S3 access, operating systems, EKS/ECS/Lambda, databases, applications, CI/CD, KMS, backup, and security-service findings. Not every source belongs in one dataset or retention class.
CloudTrail organization design
An organization trail can capture enabled member accounts, including accounts later added. Only the management account or registered CloudTrail delegated administrator creates/manages it; member accounts can see but not modify the organization trail. Prefer a multi-Region trail unless residency or tightly defined scope requires additional design.
Decide separately:
- read/write management events and global service events;
- data events such as S3 object or Lambda invocation activity;
- network activity events where supported;
- advanced selectors and exclusions;
- S3 archive delivery, optional CloudWatch Logs integration, and CloudTrail Lake event data stores;
- KMS encryption, SNS notification, and log-file integrity validation.
Management-event coverage is not data-event coverage. Data/network events can be high volume and chargeable; select from threat/compliance use cases and measure before broad enablement. Avoid excluding security-relevant service events merely to reduce cost without an approved replacement.
Destination policy and encryption
The Log Archive bucket policy must permit the documented CloudTrail service actions and organization/account prefixes while enforcing bucket-owner control, expected ACL/header behavior, TLS, and a specific trail aws:SourceArn. The trail ARN uses the management-account ID even when managed by a delegate. Avoid a broad service-principal write grant.
The KMS key policy must allow CloudTrail encryption for the intended trail/context and tightly scoped administrator/audit/decrypt roles. KMS availability, key disable/deletion, policy changes, grants, quotas, rotation, and cross-account read paths are part of logging availability. Test them; “SSE-KMS selected” is not proof.
Integrity, delivery, and failure
CloudTrail integrity validation delivers hourly signed digest chains using SHA-256 and RSA. Enabling it does not perform validation. Schedule aws cloudtrail validate-logs or an approved independent process, preserve results, and understand that the CLI expects original S3 locations.
Monitor get-trail-status, including latest delivery/digest times and errors, bucket/KMS denials, logging stopped, selector changes, and object arrival. CloudTrail can retry an unreachable S3 destination for 30 days and retry attempts can incur charges. Alert and repair immediately; do not rely on eventual retry as the SLO.
Generate a benign canary API event in each account/Region and prove it reaches the object, validates, becomes queryable, triggers the expected test detection, and reaches the case owner within target time.
S3 archive durability and immutability
Use separate buckets/prefixes and retention classes where policy, access, residency, or lifecycle differs. Enable Block Public Access, bucket-owner controls, TLS-only policy, versioning, access telemetry, narrowly scoped writes, explicit denies for unsafe deletion/configuration, and replication where the recovery/regulatory requirement justifies it.
S3 Object Lock is version-level WORM protection:
- Governance mode: protected from normal deletion but principals with bypass authority can override it;
- Compliance mode: no user, including account root, can shorten retention or delete a protected version before expiry;
- legal hold: indefinite until a principal with legal-hold permission removes it; independent of retention.
Object Lock does not prevent new versions or delete markers and does not replace policy, versioning, backup, monitoring, or account protection. Compliance retention is a serious business/legal decision: test lifecycle and storage cost using non-production data before committing. Separate retention administration from routine analysts.
Replication requires destination bucket/key policies, ownership, Object Lock compatibility, replication health, failure alerts, and tested recovery/query. A second mutable copy controlled by the same compromised role is not independent resilience.
AWS Config and configuration evidence
Configuration recorders capture supported resource configuration and changes; delivery channels send snapshots/history, and rules/conformance packs evaluate compliance. An organization aggregator collects configuration/compliance views across accounts and Regions but does not itself enable recorders, create source data, remediate resources, or prove complete coverage.
Record recorder status, recording strategy/resource types, global-resource behavior, delivery channel, aggregation authorization/trusted access, account/Region sources, failed aggregation, rule scope, remediation, retention, and cost. Test a controlled configuration change through recorder, delivery, aggregator, evaluation, alert, and restoration.
Findings and delegated security services
Findings are derived signals, not substitutes for source evidence. A finding can be delayed, suppressed, archived, normalized, duplicated, or absent because coverage is disabled.
For each service record delegated administrator, trusted access/service-linked roles, Regions, all/current/new-account enrollment, optional protection plans, standards/controls, finding destination/aggregation, suppression workflow, response SLA, and cost.
Security Hub CSPM central configuration uses a home Region and optional linked Regions. Configuration policies can centrally manage accounts/OUs and the home Region also aggregates linked-Region findings. Opt-in Regions must be enabled in member accounts. Changing/removing the home Region can delete configuration policies/associations while members retain previous settings; treat it as migration, not a console preference.
GuardDuty delegation and organization auto-enable/protection plans are Region-aware; repeat and reconcile across intended Regions. “Auto-enable new” is not proof that existing accounts or every protection plan is enabled. Apply the same evidence discipline to Inspector, Macie, Detective, Security Lake, and other services because semantics differ.
Security Lake can normalize supported AWS and custom sources into OCSF and provide subscriber/query access. It adds an ingestion/data-lake architecture; it does not automatically replace every raw archive, operational log, or service-specific investigation path. CloudWatch cross-account observability/unified data capabilities can centralize operational query; keep long-term immutable copies and fine-grained analyst access separate.
Data classification, minimization, and residency
Logs can contain personal data, tokens, secrets accidentally logged by applications, request/response fields, SQL, object keys, source IPs, and architecture details. Centralizing them increases both detection value and breach blast radius.
Define prohibited fields, SDK/application redaction, sampling, masking/tokenization, schema controls, producer tests, quarantine, access by data class, residency/transfer, retention, subject/legal processes, and incident handling. Never redact the only forensic identifiers without a documented correlation method. Do not centralize regulated data across Regions merely because an organization trail makes it easy; obtain legal/data-owner decisions.
Query and incident-access architecture
Archive is not analysis. Choose query paths by latency and format: CloudWatch Logs Insights, Athena/Glue/Lake Formation over S3, CloudTrail Lake, OpenSearch, Security Lake subscribers, or a SIEM. Use partitioning by source/account/Region/date, schema/catalog ownership, workgroups/query limits, column/row permissions, masking, result-bucket protection, and query audit.
Responders assume short-lived, MFA-protected roles with case/purpose context and least privilege. Separate triage, broad forensic read, export, legal hold, and remediation. Emergency access needs independent credentials, approval where possible, alarms, session evidence, expiry, and post-incident review. Test access during IdP, network, key, or primary-Region failure.
Monitoring the monitoring system
Measure expected versus received accounts/Regions/sources, last event/object time, volume anomalies, schema/parse failure, queue age/DLQ, subscription/export failure, KMS/S3 denial, recorder/detector/member status, disabled controls, query/index lag, alert-to-case delivery, retention drift, replication, and canary success.
Use independent alarms for changes to trails, recorders, destinations, key/bucket policies, lifecycle/Object Lock, delegated administrators, suppressions, event selectors, and pipelines. Event-driven alarms need periodic reconciliation because delivery can be best effort.
Read-only evidence audit
With approved access, redact account IDs, bucket/key ARNs, log content, IPs, domains, and finding details:
aws organizations list-delegated-administrators
aws cloudtrail describe-trails --include-shadow-trails false --region ap-south-1
aws cloudtrail get-trail-status --name replace-with-trail-arn --region ap-south-1
aws cloudtrail get-event-selectors --trail-name replace-with-trail-arn --region ap-south-1
aws configservice describe-configuration-recorders --region ap-south-1
aws configservice describe-configuration-recorder-status --region ap-south-1
aws configservice describe-configuration-aggregators --region ap-south-1
aws securityhub get-finding-aggregator --finding-aggregator-arn replace-with-arn --region replace-with-home
aws guardduty list-detectors --region ap-south-1
aws s3api get-bucket-versioning --bucket replace-with-archive
aws s3api get-object-lock-configuration --bucket replace-with-archive
aws s3api get-bucket-replication --bucket replace-with-archive
Do not print bucket policies, KMS policies, findings, or events into shared terminals without authorization. Review them securely. Repeat every intended Region and account population; follow pagination.
Troubleshooting matrix
| Symptom | Evidence order | Likely gap |
|---|---|---|
| trail says logging, no new objects | trail status, S3/KMS policy, SourceArn, delivery errors | destination path failed |
| objects exist but cannot prove integrity | validation setting, digest arrival/chain, validation run | enabling confused with validating |
| new account has no findings | org membership, delegated admin, auto-enable mode, Region/protection plan | onboarding coverage gap |
| Security Hub home view misses a Region | Region enabled, linked, service enabled, aggregator | aggregation confused with enablement |
| Config aggregator looks empty | recorder/status/source authorization/Region | aggregator creates no source evidence |
| analysts can delete source logs | role, bucket/key policy, Object Lock/bypass | custody and analysis conflated |
| alert works but investigation times out | catalog, partitions, query rights/limits, KMS | storage without query SLO |
| logging cost spikes | source/selector, bytes/events, duplicate pipelines, query scans | no telemetry cost model |
| application leaks secrets into logs | schema/redaction, sampling, access and incident evidence | classification absent |
| archive survives but cannot be recovered | replica/key/catalog/role/query game day | storage durability confused with service recovery |
Cost and capacity model
Model each source as events or GB per account × Region × day, then apply collection, ingestion, transformation, delivery, S3 storage tiers/retention, replication/transfer, KMS requests, catalog/index, query scans, archive retrieval, findings/security coverage, SIEM licensing, alerts, support, and engineering labor. Include test/non-production, duplicate copies, malformed retries, and 30-day CloudTrail redelivery risk.
Lifecycle transitions reduce storage price but can add minimum-duration, retrieval, and incident-delay costs. Data-event, Config, GuardDuty protection-plan, Security Hub, Macie, CloudWatch, Firehose, OpenSearch, Security Lake, and Athena pricing dimensions differ. Assign budget/anomaly alerts and a service owner; optimize scope from risk and measured use, never by silently disabling evidence.
Failure, recovery, and decommissioning
Game-day S3 policy denial, KMS disable/deny, stopped trail/recorder, Region/account onboarding gap, broken subscription/DLQ, schema change, SIEM outage, compromised analyst, lost IdP, inaccessible archive tier, and failed replica. Measure loss, backlog, duplicate handling, detection/query recovery, and RTO/RPO.
Before changing/deleting a source, delegate, trail, bucket, key, pipeline, table, index, or account, inventory producers/consumers, retention/legal holds, investigations, replicas, queries, dashboards, detections, costs, and downstream automation. Preserve required originals and digest chain. Migrate and reconcile, freeze old writes, prove final delivery/query, revoke access, then dispose only under approved retention policy.
Practical lab
Download the AWS260 centralized evidence architecture pack. It contains a source/retention matrix, account-and-access workbook, eight failure cases, delivery/integrity game day, and answer directions.
Submit source-to-response diagrams; Log Archive/Security Tooling RACI; CloudTrail/S3/KMS/Object Lock design; Config and security-service coverage; Region/onboarding model; classification; query/incident roles; health SLOs; cost model; eight diagnoses; and reversible migration/decommission plan.
Knowledge check
- Why separate Log Archive from Security Tooling?
- Which stages exist between event generation and response?
- Why do organization trail, management events, and data events prove different coverage?
- What must S3 and KMS resource policies constrain?
- Why does enabling integrity validation not prove integrity?
- How do Object Lock governance, compliance, and legal hold differ?
- Why is a Config aggregator not a recorder?
- How do Security Hub home/linked Regions affect configuration and findings?
- Why must GuardDuty onboarding be reconciled per Region and protection plan?
- When can central logging create privacy/residency risk?
- Why is archive storage without tested query access incomplete?
- What monitors the logging system itself?
- Which dimensions dominate cost?
- What must be preserved before decommissioning a trail or account?
Lesson acceptance
Pass requires complete source inventory and coverage; separated accounts/duties; policy and encryption boundaries; immutable retention decision; integrity and canary proof; Config/finding/delegation semantics; Regions/new-account coverage; data classification; least-privilege query/incident access; health/reconciliation; capacity/cost; failure/recovery tests; eight case diagnoses; and retention-safe decommissioning.
Official sources
- AWS SRA Log Archive account
- AWS SRA Security Tooling account
- CloudTrail organization delegated administrator
- Prepare an organization trail
- CloudTrail integrity validation
- CloudTrail data events
- S3 Object Lock
- Config multi-account aggregation
- Security Hub central configuration
- GuardDuty organization administration
- Security Lake architecture