Lesson 211 · AWS Learning Path

AWS 211: AWS X-Ray, CloudWatch Application Signals, and OpenTelemetry tracing

· Published · 8 min read

Labelled process diagram for AWS 211: Instrumented request to OTel SDK and collector or agent to X-Ray trace and Application Signals to Correlated service, SLO, log, and fault evidence, with decision, proof and...

Why this lesson matters

Distributed tracing follows one request across process and service boundaries so latency and errors can be assigned to the responsible dependency. It works only when trace context is propagated, spans use stable service/resource attributes, sampling preserves representative evidence and logs can correlate by trace/request ID. For new instrumentation, AWS recommends OpenTelemetry; the older X-Ray SDKs and daemon are now in maintenance mode.

What you will be able to do

By the end, you can:

  • distinguish trace, span, parent/child relationship, link, event, attribute and resource;
  • propagate W3C trace context across synchronous calls and reason about asynchronous links;
  • choose OpenTelemetry/ADOT instrumentation and a CloudWatch agent or OTel collector path;
  • explain head sampling, sampling bias and why one trace is not an availability statistic;
  • use X-Ray trace/map evidence and Application Signals service/dependency/SLO views together;
  • correlate trace ID with structured logs and metrics without recording secrets or personal data;
  • diagnose missing spans, broken parentage, service-name explosion and collector export failure.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeInstrument new tracing with OpenTelemetry or ADOT, send traces to X-Ray and CloudWatch, and correlate services, spans, metrics, logs, sampling, and service-level objectives.
Scope and boundaryFor X-Ray, Application Signals, and OpenTelemetry, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup.
Evidence of successSuccess requires matching Console, command-line, behavior, monitoring, and owner evidence for X-Ray, Application Signals, and OpenTelemetry. One green status is not enough.
Cost modelTrace recording, Application Signals, custom metrics, logs, collector capacity, and retained telemetry can charge.
Safe rejection ruleAvoid starting new code on maintenance-mode X-Ray SDKs or recording secrets and high-cardinality personal data in span attributes.

How the request flows

+------------------------+
|  Instrumented request  |
+------------------------+
            |
            v
+-----------------------------------+
|  OTel SDK and collector or agent  |
+-----------------------------------+
                 |
                 v
+---------------------------------------+
|  X-Ray trace and Application Signals  |
+---------------------------------------+
                   |
                   v
+----------------------------------------------------+
|  Correlated service, SLO, log, and fault evidence  |
+----------------------------------------------------+

Trace and signal model, deeply

A trace represents one end-to-end transaction. A span represents one timed operation and carries a trace ID, span ID, parent (for a tree), start/end, status, name and attributes. Resource attributes identify the telemetry-producing service/runtime; span attributes describe the operation. Events add timestamped details. Links connect causally related work when a strict parent/child relationship is wrong, such as producer and later queue consumer.

The caller injects trace context into supported HTTP/message headers; the receiver extracts it and starts a child span. Broken propagation creates separate traces even when every service emits spans. Do not put tokens, passwords, full request bodies, sensitive SQL parameters or unapproved user identity in attributes. Unbounded values also make search and service naming expensive/confusing.

For new code, use upstream OpenTelemetry or AWS Distro for OpenTelemetry instrumentation, then export through the CloudWatch agent, ADOT/OTel collector or supported endpoint to X-Ray/Application Signals. Since February 25, 2026, X-Ray SDKs and daemon receive maintenance/security fixes only; AWS documentation recommends migration to OpenTelemetry. Plan migration before the documented February 25, 2027 end of support for those SDKs/daemon. X-Ray as a trace analysis destination is not the same thing as the maintenance-mode instrumentation libraries.

Application Signals uses collected metrics/traces to present services, operations, dependencies and service-level objectives. An SLO needs an SLI, target, evaluation period and owner - for example, successful requests under a latency threshold - not simply CPU. Trace maps show sampled relationships and faults; they do not prove every request was captured.

Sampling controls volume. Head sampling decides near request start and can miss rare failures; parent-based sampling should keep a trace decision coherent. Record the effective rule, reservoir/rate or probability, upstream decision and collector behavior. Increase sampling only with cost/privacy approval and restore it after the investigation.

Architecture decision table

SituationDirectionReason
Requirement matchesUse OpenTelemetry or ADOT for new instrumentation and X-Ray or Application Signals as supported analysis destinations.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid starting new code on maintenance-mode X-Ray SDKs or recording secrets and high-cardinality personal data in span attributes.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open CloudWatch > Application Signals > Services in the correct Region. Select a service and compare call volume, fault/error rate and latency for one absolute UTC range.
  2. Open one operation, one dependency and its SLO. Record service/operation names, SLI query, target, evaluation interval and burn/error-budget evidence.
  3. Follow a trace from the service map. Draw each span with service, operation, parent/link, duration, status and downstream call; identify the critical path rather than adding every span duration.
  4. Copy the redacted trace ID into Logs Insights and prove matching application logs. Compare event timestamps with span boundaries.
  5. Inspect instrumentation library/version, propagator, sampling rule and collector/agent configuration from approved evidence. Record exporter destination and IAM/network path.
  6. This lesson is read-only. Do not change sampling, enable Application Signals or deploy collectors without cost and workload-owner approval.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

start="$(date -u -d '1 hour ago' +%Y-%m-%dT%H:%M:%SZ)"
end="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
aws xray get-service-graph --start-time "$start" --end-time "$end" --output json
aws xray get-trace-summaries --start-time "$start" --end-time "$end" --output json
aws cloudwatch list-metrics --namespace ApplicationSignals --output table

Expected interpretation

The graph is reconstructed from sampled traces in the selected window. A missing node can mean no traffic, no sampling, broken propagation, missing instrumentation or exporter failure. A trace summary locates candidate traces; retrieve the batch trace and inspect segment/subsegment documents for the real timeline. Application Signals metrics are aggregate evidence and may exist even when one expected trace was not sampled.

Practical work

Draw a request across API, Lambda, database, and external dependency. Define OTel trace and span IDs, propagation, sampling, collector or CloudWatch agent, redaction, service name, SLO, logs correlation, and missing-span test.

Diagnose this topic from its own evidence

SymptomProve firstSmallest safe correction
no traces anywhereSDK auto/manual instrumentation, sampling, collector health, IAM and endpointrepair first failed export hop
one service missingincoming/outgoing propagation and instrumented framework/clientenable supported instrumentation and preserve W3C context
traces split at queueproducer context injection and consumer extraction/link modelpropagate message attributes and model async causality
service map explodesunstable service.name/resource attributes or URL values in namesnormalize low-cardinality service/operation naming
trace says success but user failedspan status/exception mapping and business assertionrecord meaningful error/status and correlate logs/SLI

Positive test: one supplied request produces a connected trace and matching log correlation ID. Negative test: a controlled downstream error marks the responsible span and service signal. Dependency-failure test: collector export is denied while the application still responds; distinguish telemetry loss from workload failure.

Cost and cleanup

Trace recording, Application Signals, custom metrics, logs, collector capacity, and retained telemetry can charge.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Instrument new tracing with OpenTelemetry or ADOT, send traces to X-Ray and CloudWatch, and correlate services, spans, metrics, logs, sampling, and service-level objectives.

  1. Which scope or ownership boundary must be proved first?

Expected direction: For X-Ray, Application Signals, and OpenTelemetry, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup.

  1. What evidence is strong enough to accept the result?

Expected direction: Success requires matching Console, command-line, behavior, monitoring, and owner evidence for X-Ray, Application Signals, and OpenTelemetry. One green status is not enough.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid starting new code on maintenance-mode X-Ray SDKs or recording secrets and high-cardinality personal data in span attributes.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Trace recording, Application Signals, custom metrics, logs, collector capacity, and retained telemetry can charge.

Lesson acceptance

  • A trace diagram identifies trace/span IDs, parent/link relationships, critical path, statuses and resource attributes.
  • W3C propagation and asynchronous messaging boundaries are explained.
  • OpenTelemetry is selected for new instrumentation, with the X-Ray SDK/daemon maintenance and end-of-support dates recorded correctly.
  • Connected positive, downstream-error and collector-failure cases are diagnosed from traces, logs, metrics and collector evidence.
  • Sampling, SLO, sensitive-data, cardinality, retention and telemetry-cost owners are named.

Official sources

Advertisement