Lesson 246 · AWS Learning Path

AWS 246: Monitoring, logging, remediation, and performance

· Published · 9 min read

Labelled process diagram for AWS 246: Operational scenario to Telemetry and state evidence to Diagnosis or remediation decision to Score and targeted retest, with decision, proof and rejection evidence.

Why this checkpoint exists

This is a gate, not another guided lab. You must decide which signal is missing, read unfamiliar evidence, distinguish symptom from cause, select a bounded remediation, and explain performance and cost trade-offs without step-by-step prompting.

AWS renamed the operations certification from AWS Certified SysOps Administrator – Associate to AWS Certified CloudOps Engineer – Associate for SOA-C03. The prior SOA-C02 exam ended September 29, 2025, and SOA-C03 began September 30, 2025. In the current guide, Domain 1 - Monitoring, Logging, Analysis, Remediation, and Performance Optimization - is 22% of scored content. This checkpoint follows the current three-task blueprint rather than the retired six-domain outline.

This course gate is intentionally stricter than the certification's compensatory scoring model. AWS does not require a pass in every domain; this course requires balanced operational competence so a strong score cannot hide a dangerous weakness.

What is assessed

SOA-C03 taskCompetence tested hereQuestions
1.1 Metrics, alarms, filters, loggingtelemetry choice, agent, alarm semantics, dashboards, SNS1–10
1.2 Analysis and remediationevidence correlation, EventBridge, Systems Manager Automation, guardrails11–20
1.3 Performance optimizationEC2/network, EBS, S3, EFS/FSx, RDS/RDS Proxy, containers21–30

The 30 questions are original course material, not copied certification items or an exam dump. Service lists and console labels can change; reasoning from requirements matters more than memorizing a screen.

Required pack and assessment rules

Download the AWS246 Domain 1 checkpoint pack. It contains:

  • DOMAIN1_CHECKPOINT_WORKBOOK.md - answer sheet, evidence tasks, scoring, and retest plan;
  • DOMAIN1_SCENARIOS.md - 30 single- and multiple-response scenarios;
  • SUPPLIED_EVIDENCE_CASES.md - CLI/log/performance evidence for practical interpretation;
  • INSTRUCTOR_ANSWER_DIRECTIONS.md - answers and reasoning; keep closed until the timed attempt ends.

Rules:

  1. Use 75 minutes for questions 1–30 and 45 minutes for the three supplied evidence cases.
  2. No answer key, search engine, AI assistant, or course notes during the timed attempt.
  3. Select the stated number of responses. A multiple-response item earns one point only when the complete set is correct.
  4. Record confidence as high, medium, or low before grading.
  5. After grading, explain every incorrect or low-confidence answer from first principles and cite an official source.
  6. Retest with changed symptoms; do not memorize option letters.

The operating model you must apply

user impact or objective
        |
        v
resource and application signals
        |
        +--> metrics: numeric behavior over time
        +--> logs: timestamped events and context
        +--> traces: request path and latency contribution
        +--> changes: who changed what and when
        +--> configuration: current/recorded resource state
        v
falsifiable diagnosis
        v
bounded decision: observe / contain / remediate / escalate
        v
customer verification + telemetry recovery + retained evidence

Metrics, logs, traces, and changes are not interchangeable

  • A metric is efficient for trend, threshold, rate, percentile, saturation, and alarm evaluation. It usually lacks per-request detail.
  • A log records discrete events and dimensions chosen by the producer. Logs Insights queries do not invent fields that were never emitted.
  • A trace links work across components and reveals latency/error contribution, but sampling means it is not a complete ledger.
  • CloudTrail records supported AWS control-plane activity and selected data events; it is not an application transaction log.
  • AWS Config records supported resource configuration and compliance over time; it does not inspect Linux process state or secret plaintext.

An architect starts with the question and chooses the signal. “Enable every log forever” is not an observability strategy; it is an unowned cost and privacy risk.

CloudWatch concepts that must be automatic

A metric is identified by namespace, metric name, and exact dimension set. Average, Sum, Minimum, Maximum, sample count, and percentiles answer different questions. Period is aggregation width; evaluation periods define how many recent periods are considered; datapoints-to-alarm defines M-of-N behavior. Missing data can be treated as missing, breaching, not breaching, or ignored - each can be correct or dangerous depending on whether silence means healthy, idle, or broken telemetry.

A metric alarm watches one metric or metric-math expression. A composite alarm combines other alarm states and reduces notification noise; it cannot directly perform EC2 or Auto Scaling actions. Alarm actions generally occur on state transitions, not repeatedly merely because a state remains ALARM. Verify action enablement, SNS policy/subscription, Region, and state history before blaming the metric.

High-resolution custom metrics allow sub-minute periods but cost more. Standard EC2 metrics do not include guest memory or filesystem utilization; collect those with the CloudWatch agent or another approved telemetry path. Agent success requires configuration, host/container permissions, network/DNS reachability, credentials/instance role, correct Region, and a running process.

Metric filters transform matching log events into metrics; subscription filters stream matching log events to a supported destination; Logs Insights queries stored log data interactively. Retention limits future storage duration - it does not replace an archive, legal hold, or protected audit design.

Analysis and safe remediation

Form a hypothesis before changing state. Align timestamps in UTC, account, Region, resource ID, deployment/change ID, and request/trace ID. Correlation is not causation: CPU and errors rising together do not prove CPU is the cause.

EventBridge rules match events; targets receive matched events. Diagnose source emission, event bus/account policy, pattern structure, target role/resource policy, retry policy, dead-letter queue, throttling, and target behavior separately. An event pattern is an exact structural matcher, not a text search. Archive/replay can redeliver events, so consumers and remediations must be idempotent.

Automation must read before writing, validate ownership/current state, use least privilege, bound retries/timeouts/concurrency, verify customer behavior, and explicitly compensate or escalate. MaxErrors stops dispatching new work after the threshold but does not cancel work already running. Approval is useful only before mutation and with evidence. AWS245 is the minimum safety baseline.

Performance reasoning by bottleneck

Performance optimization means identifying the limiting resource and requirement before resizing.

EC2 and containers

Correlate CPU, memory, disk, network, load average, run queue, throttling, application latency, and downstream dependencies. CPU credit metrics matter for burstable instances. Enhanced networking, instance bandwidth, ENA queue behavior, placement groups, and packet-per-second limits can matter even when CPU is low. In ECS/EKS, separate task/pod limits, node saturation, scheduler placement, load balancer health, and application concurrency.

EBS

Volume type defines independent or coupled IOPS/throughput behavior. Inspect queue length, latency, operations, bytes, burst balance where applicable, instance EBS bandwidth, filesystem, and application I/O size/pattern. Increasing gp3 IOPS cannot solve an instance bandwidth ceiling; increasing throughput does not solve a small random-I/O IOPS limit. Snapshot initialization can cause first-read latency unless blocks are initialized or Fast Snapshot Restore is planned.

S3 and transfer

Use multipart upload for large objects and parallelism/resume; lifecycle for storage transitions/expiry; DataSync for managed movement with verification/scheduling; Transfer Acceleration when internet distance and measured benefit justify its cost. S3 request distribution no longer requires random key prefixes. Diagnose client concurrency, object size, network path, retries, KMS/request limits, and transfer pricing.

Shared file storage

Choose EFS for elastic NFS semantics, FSx families for specific Windows/Lustre/NetApp/OpenZFS capabilities, and S3 when object semantics fit. EFS throughput/performance mode, lifecycle, access pattern, client mount/network path, and small-file metadata operations affect results. A file-system requirement cannot be solved merely by renaming S3 objects as files.

Databases

For RDS/Aurora, correlate database load, top waits/SQL, CPU, free memory, connections, storage latency/queue, replica lag, locks, and application pool behavior. A read replica scales eligible reads, not writes or strongly read-after-write paths. Multi-AZ primarily improves availability. RDS Proxy manages connection pooling and failover behavior; it does not optimize inefficient SQL. Cache only when consistency, invalidation, and access patterns permit it.

Practical evidence cases

The pack supplies three unfamiliar cases:

  1. Silent alarm: a metric is present but notification did not arrive. Determine whether evaluation, state transition, disabled action, SNS, or Region is responsible.
  2. Remediation loop: EventBridge and Automation repeatedly touch an already-correct resource. Find the missing idempotency and loop-prevention controls.
  3. Slow database application: CPU is low while connection and wait evidence is high. Reject blind instance resizing and choose the next discriminating evidence/action.

For each, submit:

  • impact and scope;
  • three ranked hypotheses;
  • evidence supporting and contradicting each;
  • next read-only discriminator;
  • bounded remediation with rollback/compensation;
  • customer, telemetry, security, and cost verification.

Read-only AWS inventory option

The supplied track is complete. If the account owner permits read-only inspection, prove caller and Region first and redact identifiers:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws cloudwatch describe-alarms --query 'MetricAlarms[].{Name:AlarmName,State:StateValue,Actions:ActionsEnabled,TreatMissing:TreatMissingData}'
aws logs describe-log-groups --query 'logGroups[].{Name:logGroupName,Retention:retentionInDays,Bytes:storedBytes}'
aws events list-rules --query 'Rules[].{Name:Name,State:State,Bus:EventBusName}'
aws ssm describe-automation-executions --max-results 10
aws rds describe-db-instances --query 'DBInstances[].{Id:DBInstanceIdentifier,Class:DBInstanceClass,MultiAZ:MultiAZ,Storage:StorageType}'

Inventory output does not prove collection quality, delivery, least privilege, customer behavior, or performance. Do not execute remediation in this checkpoint.

Scoring and gate

Question score: 30 points. Practical evidence: 30 points, ten per case. Total: 60.

Course pass requires all of:

  • at least 24/30 scenario points (80%);
  • at least 24/30 practical points (80%);
  • at least 8/10 in each task group;
  • no critical safety miss: destructive guessing, broad privilege as a fix, secret exposure, unbounded automation, or success claimed without customer verification.

This threshold is a course rule, not an AWS-reported exam passing score. A failed task group produces a targeted study plan tied to AWS204–AWS245 and a changed retest after at least one deliberate practice exercise.

Error taxonomy and retest

Label every miss:

  • concept gap - did not understand the service behavior;
  • scope gap - missed account, Region, dimension, resource, identity, or time window;
  • evidence gap - acted without a discriminating signal;
  • wording gap - missed “most operationally efficient,” “least cost,” or response count;
  • safety gap - selected broad, destructive, irreversible, or unbounded action;
  • confidence gap - correct guess without an explainable model.

Do not merely reread. For each gap, write why the distractor was tempting, change one condition that would make it correct, perform a small read-only interpretation, and answer a changed case.

Cost and privacy

This checkpoint creates no AWS resources. In real designs, estimate metric/custom/high-resolution ingestion, API calls, log ingestion/storage/query/scanning/delivery, dashboards, alarms, traces, notifications, Automation steps/script duration, data transfer, and the resources retained to produce telemetry. Define retention by operational, security, legal, and cost needs. Redact account IDs, full ARNs, IPs where sensitive, customer data, secrets, tokens, and presigned URLs.

Acceptance evidence

Submit the completed workbook with timed start/end, original answers and confidence, graded task scores, practical evidence analyses, error taxonomy, official-source corrections, targeted practice, changed retest, and final no-create statement. Passing requires reasoning; an answer sheet copied after opening the key is not evidence.

Official sources

Advertisement