Lesson 210 · AWS Learning Path

AWS 210: AWS Systems Manager

· Published · 8 min read

Labelled process diagram for AWS 210: Authorized operator to Systems Manager API and document to Managed node agent to Command, session, compliance, and audit evidence, with decision, proof and rejection evidence.

Why this lesson matters

Systems Manager provides several distinct fleet-operation capabilities behind one console. The shared foundation is an authorized operator, a service API/document, a correctly registered managed node, SSM Agent, IAM credentials and outbound service connectivity. Installing the agent satisfies only one dependency. This lesson teaches the foundation first so later Run Command, Patch Manager, State Manager and Automation work is diagnosable and auditable.

What you will be able to do

By the end, you can:

  • distinguish control-plane operator permission from managed-node permission;
  • prove SSM Agent version, registration, ping state, platform and last association state;
  • compare an instance profile, Default Host Management Configuration and hybrid activation;
  • map required outbound DNS/TLS/endpoints without opening inbound SSH;
  • choose Session Manager, Run Command, State Manager, Patch Manager, Automation, Inventory, Parameter Store or OpsCenter by purpose;
  • explain document permissions, parameters, target controls, output and CloudTrail evidence;
  • diagnose ConnectionLost, missing-node, access-denied and delivery failures layer by layer.

Before you start

  • Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
  • CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
  • The course example Region is ap-south-1. Global services and services with a required control Region are called out in their commands.
  • Run aws sts get-caller-identity privately. Redact the account number before sharing evidence.
  • Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
  • This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
  • Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.

The core model

QuestionWhat it means in this lesson
PurposeOperate managed nodes and workloads through Session Manager, Run Command, Patch Manager, State Manager, Automation, Inventory, Parameter Store, OpsCenter, and Explorer with scoped roles.
Scope and boundaryFor AWS Systems Manager, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup.
Evidence of successSuccess requires matching Console, command-line, behavior, monitoring, and owner evidence for AWS Systems Manager. One green status is not enough.
Cost modelRequests, running capacity, stored telemetry, retained history, data transfer, and connected resources must be priced for the exact design.
Safe rejection ruleAvoid broad document permissions, public SSH as the default, secrets in command parameters, or assuming an installed agent is an online managed node.

How the request flows

+-----------------------+
|  Authorized operator  |
+-----------------------+
           |
           v
+------------------------------------+
|  Systems Manager API and document  |
+------------------------------------+
                  |
                  v
+----------------------+
|  Managed node agent  |
+----------------------+
           |
           v
+----------------------------------------------------+
|  Command, session, compliance, and audit evidence  |
+----------------------------------------------------+

Managed-node foundation and capability map

A managed node needs four independent conditions:

  1. A supported, running SSM Agent with valid local registration state.
  2. Node credentials: an EC2 instance profile, Default Host Management Configuration (DHMC), or hybrid/multicloud activation role appropriate to the node type.
  3. Outbound HTTPS and DNS to the Regional Systems Manager messaging/service endpoints, through internet/NAT or approved VPC endpoints and their policies.
  4. Service and operator permissions for the selected document/action, plus KMS/S3/CloudWatch Logs permissions when those destinations are used.

DHMC can automatically provide management permissions to EC2 instances in an account/Region. Current prerequisites include IMDSv2 and SSM Agent 3.2.582.0 or later. SSM Agent tries instance-profile credentials before DHMC credentials, so an attached but inadequate instance profile can produce surprising authorization behavior. Organization-wide Quick Setup can apply DHMC and agent updates across selected accounts/Regions; treat that as broad administrative scope requiring change control.

CapabilityPrimary purposeDo not confuse with
Fleet Managerinventory/view/manage individual nodesconfiguration compliance policy
Session Managerinteractive shell/tunnel without inbound SSHproof that every command was logged unless logging is configured and supported
Run Commandfan-out non-interactive document executionlong-lived desired state
State Managerperiodically apply/verify an associationpackage vulnerability assessment
Patch Managerbaseline and orchestrated scan/install/complianceapplication dependency testing
Inventorycollected node metadatareal-time process monitoring
Automationmulti-step runbooks, branching/approvals/AWS APIsarbitrary unbounded scripts
Parameter Storehierarchical configuration and SecureString valuesautomatic rotation provided by Secrets Manager
OpsCenter/Exploreroperational work and aggregated viewsthe original metric/log/audit source

Documents are versioned execution definitions. Control who may call the API, which document/version they may use, which tag/resource targets are allowed, maximum concurrency/errors, and where output is retained. Secure parameters can still leak if a command echoes them; IAM/KMS access and application logging must prevent disclosure.

Architecture decision table

SituationDirectionReason
Requirement matchesUse Systems Manager to reduce inbound administration paths and standardize auditable fleet operations.Select only after scope, behavior, security, recovery, operations, and price evidence agree.
Requirement does not matchAvoid broad document permissions, public SSH as the default, secrets in command parameters, or assuming an installed agent is an online managed node.Rejecting an attractive service is a valid architecture result.
No create permission or cost approvalUse supplied evidence and local design workLearning does not depend on creating an hourly resource.
Existing resource is unknown or unownedInspect only, then stopNever change or delete a resource merely because it resembles a course example.

AWS Management Console, step by step

Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.

  1. Open Systems Manager > Fleet Manager > Managed nodes in ap-south-1. Record one approved node's ID, ping status/time, platform, SSM Agent version, IP/computer name and association status.
  2. Open node details and inspect IAM role/registration source, tags, inventory, associations, patch compliance and command history. Do not start a session or command on this read-only track.
  3. Open Settings and inspect DHMC status/role and update settings. Determine whether the selected node uses its instance profile first.
  4. In EC2/VPC read-only views, prove IMDSv2 setting, SG egress, subnet route/DNS and any ssm, ssmmessages or other documented VPC endpoints/policies used by the design.
  5. Open one approved command/association and identify document owner/version, parameters, targets, concurrency/errors, output destinations and per-node result.
  6. Correlate the operator API event in CloudTrail with Systems Manager execution ID and node/agent output. Redact commands that contain sensitive values.

CloudShell and AWS CLI, step by step

Start with a known caller and Region:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account part of the ARN in shared evidence. Now run the topic queries:

aws ssm describe-instance-information --query 'InstanceInformationList[].{Id:InstanceId,Ping:PingStatus,Platform:PlatformName,Version:AgentVersion,Association:AssociationStatus}' --output table
aws ssm list-associations --query 'Associations[].{Name:Name,Association:AssociationId,Status:Overview.Status}' --output table
aws ssm describe-instance-associations-status --instance-id replace-with-approved-managed-node-id --output json
aws ssm list-command-invocations --details --max-results 20 \
  --query 'CommandInvocations[].{Command:CommandId,Node:InstanceId,Status:Status,Document:DocumentName,Requested:RequestedDateTime}' --output table
aws ssm get-service-setting \
  --setting-id arn:aws:ssm:ap-south-1:replace-with-account-id:servicesetting/ssm/managed-instance/default-ec2-instance-management-role \
  --output json

Expected interpretation

PingStatus=Online proves recent agent/service communication, not that a specific document is authorized or succeeds. Association overview can summarize multiple targets; inspect per-node execution. Command status Success means the document/plugin reported success, not that the application-level postcondition passed. DHMC service setting is Regional and account-scoped; redact account IDs in shared evidence.

Practical work

Create a responsibility map showing the operator identity, Systems Manager service, instance profile or hybrid activation, agent, endpoints, logs, KMS, documents, commands, and audit trail.

Diagnose this topic from its own evidence

SymptomCheck in orderSmallest safe correction
node absentaccount/Region, supported node, agent install/run, registration and credential sourcerepair first failed prerequisite; avoid reinstalling blindly
ConnectionLostlast ping, agent log, DNS, route/SG/NACL/proxy, endpoint policy/TLS timerestore outbound service path or agent health
AccessDeniedoperator policy, node role/DHMC precedence, document/resource/KMS/S3 policyadd only the missing action/resource/condition
command pending/undeliverablenode online state, target tag snapshot, delivery timeout and concurrencyrestore node path or retarget only approved nodes
command says success but service is brokenplugin output/exit code and application postconditionfix script exit handling and add independent health test

Positive test: supplied evidence shows one online node and a harmless read-only document completing with expected stdout/exit status. Negative test: a nonmatching tag targets zero nodes. Dependency-failure test: the agent is running but endpoint access is blocked; diagnose from agent log and network path without opening SSH to the internet.

Cost and cleanup

Core Systems Manager capabilities have feature-specific pricing and limits. Price advanced/paid node tiers where applicable, Parameter Store advanced parameters/API interactions, Automation steps, OpsCenter/incident features, CloudWatch Logs and S3 output, KMS, VPC interface endpoint hours/data and the managed compute itself. Cleanup only course-owned associations, commands' retained output, activations, parameters, OpsItems and endpoints according to retention; never deregister or alter a shared production node.

Knowledge check

  1. What operational purpose is this lesson solving?

Expected direction: Operate managed nodes and workloads through Session Manager, Run Command, Patch Manager, State Manager, Automation, Inventory, Parameter Store, OpsCenter, and Explorer with scoped roles.

  1. Which scope or ownership boundary must be proved first?

Expected direction: For AWS Systems Manager, separate account and Region scope, identity, configuration, data or network behavior, failure ownership, evidence retention, and cleanup.

  1. What evidence is strong enough to accept the result?

Expected direction: Success requires matching Console, command-line, behavior, monitoring, and owner evidence for AWS Systems Manager. One green status is not enough.

  1. Which tempting design or shortcut must be rejected?

Expected direction: Avoid broad document permissions, public SSH as the default, secrets in command parameters, or assuming an installed agent is an online managed node.

  1. Which cost dimensions and retained resources need an owner?

Expected direction: Requests, running capacity, stored telemetry, retained history, data transfer, and connected resources must be priced for the exact design.

Lesson acceptance

  • The responsibility map identifies operator, API/document, node credentials, agent, network endpoints, outputs, KMS and audit sources.
  • One managed node's identity, version, ping time, platform and credential source are proved.
  • Instance profile, DHMC and hybrid activation boundaries are explained, including DHMC prerequisites and credential precedence.
  • Positive, zero-target and blocked-endpoint cases are diagnosed from separate service and agent evidence.
  • The learner selects each major Systems Manager capability by purpose and states permission, cost, retention and cleanup owners.

Official sources

Advertisement