Lesson 306 · AWS Learning Path

AWS 306: AWS IoT services overview

· Published · 23 min read

Labelled process diagram for AWS 306: Device identity and protocol to IoT Core messaging and fleet control to Edge or cloud processing to Action, telemetry, audit, and revocation evidence, with decision, proof and...

Why this lesson matters

An Internet of Things system connects physical equipment to software. That simple sentence hides unusual architecture risks. A web request can usually be retried from a managed computer on a stable network. An IoT device may have little memory, an inaccurate clock, a cellular connection that disappears for hours, a certificate installed in a factory, and an actuator capable of moving real machinery. Ten thousand devices can turn a minor reconnect bug into ten thousand simultaneous TLS handshakes and message bursts.

AWS IoT is a portfolio, not one product. AWS IoT Core provides secure device connectivity, MQTT messaging, device shadows, rules, and control features. Device Management adds fleet provisioning, indexing, Jobs, software-package visibility, and fleet operations. Device Defender audits configuration and detects unusual device behavior. AWS IoT Greengrass V2 runs and deploys software at the edge. AWS IoT SiteWise models and monitors industrial data. AWS IoT TwinMaker builds operational digital twins over existing data sources. Specialized products have different availability boundaries.

The portfolio has also changed. AWS IoT Events ended support on May 20, 2026 and must not be selected for new architecture. AWS IoT FleetWise is no longer open to new customers, although existing customers can continue using it. AWS IoT Greengrass V1 is approaching end of support and its V1 resources become unavailable after October 7, 2026 according to the migration guide current at this review. New designs use Greengrass V2. An architect must check current service status instead of copying an old reference diagram.

This workshop begins with device and message concepts, then follows identity from manufacturing to retirement. It ends with a reviewable design for 10,000 devices, staged software deployment, disconnected operation, detection, recovery, cost, and evidence.

What you will be able to do

By the end, you can:

  • explain a device, thing, certificate, MQTT topic, message, shadow, rule, Job, fleet index, component, asset, and digital twin in plain language;
  • separate device data-plane access from human and service control-plane access;
  • select MQTT, MQTT over WebSocket Secure, HTTPS, or local edge messaging from device constraints;
  • design topic names and IoT policies that prevent one device from impersonating another;
  • explain MQTT QoS 0 and QoS 1, duplicate delivery, sessions, retained messages, and offline limits;
  • use desired, reported, delta, and version fields in classic or named device shadows without creating update loops;
  • route telemetry with IoT rules and design idempotent downstream processing and error handling;
  • provision unique device identities and manage activation, rotation, revocation, ownership transfer, and retirement;
  • use fleet indexing, dynamic thing groups, Jobs, and package information for controlled fleet operations;
  • distinguish Device Defender audit findings from behavior detection and plan safe mitigation;
  • explain Greengrass V2 cores, nucleus, components, recipes, artifacts, dependencies, deployments, local communication, offline behavior, and rollback;
  • select SiteWise for industrial asset modeling and TwinMaker for contextual operational twins;
  • handle IoT Events, Greengrass V1, and FleetWise according to current lifecycle status;
  • estimate message, connectivity, registry, indexing, edge, storage, logging, and transfer cost drivers;
  • diagnose certificate, policy, MQTT, shadow, rule, Job, Greengrass, Defender, and SiteWise failures from evidence; and
  • produce and defend a secure 10,000-device lifecycle architecture.

Before you start

  • Complete AWS049 through AWS051 for IAM and AWS175 for PKI foundations.
  • Use synthetic device IDs and telemetry. Do not connect production machinery, vehicles, medical equipment, home controls, or safety systems.
  • This lesson's default T0 path creates nothing. Run only read operations against resources you own or use the supplied design evidence.
  • The optional T1 path needs account-owner approval, a budget alarm, a dedicated test thing, and complete cleanup. Never reuse a production device certificate.
  • Use one unique certificate and private key per device. Never submit private keys, claim credentials, full certificate IDs, account IDs, endpoint hostnames, or customer telemetry.
  • Confirm the account and Region before every inventory. IoT resources are Regional, while devices may physically move across network and national boundaries.
  • CloudShell cannot act like an intermittently connected embedded device. A local MQTT client or approved simulator is required for behavior testing.
  • A cloud command must never be the only physical safety control. Emergency stops, interlocks, safe local defaults, and authorized human procedures remain outside the message broker.

Create a local evidence directory:

mkdir -p "$HOME/aws306-evidence"
cd "$HOME/aws306-evidence"

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
date -u +%FT%TZ

Redact the account portion of the caller ARN before sharing it.

1. Build the IoT mental model

A device is the physical or virtual client: a sensor, gateway, appliance, controller, camera, or test process. An AWS IoT thing is a registry record that represents a device or logical asset. Creating a thing does not connect hardware, and deleting a thing does not erase a private key from hardware. The representation and the physical object have related but separate lifecycles.

Telemetry reports observations from device to cloud, such as temperature. A command asks a device to do something. Configuration states how it should operate. Firmware or software changes the code it runs. These flows deserve separate topics, authorization, retention, approval, and failure behavior. Treating every payload as generic data makes safety and ownership unclear.

physical device or edge gateway
  |  unique identity, TLS, MQTT or HTTPS
  v
AWS IoT Core data endpoint
  +--> message broker --> authorized subscribers
  +--> device shadow --> durable desired/reported state document
  +--> rules engine --> Lambda, streams, queues, storage, analytics
  +--> Jobs control --> staged remote operation or software update

fleet control and evidence
  +--> registry, groups, indexing, package inventory
  +--> Device Defender audit and behavior detection
  +--> CloudWatch logs/metrics and CloudTrail control-plane history

edge and industrial context
  +--> Greengrass V2 local applications and disconnected processing
  +--> SiteWise industrial assets, properties, metrics, and dashboards
  +--> TwinMaker entities, connectors, knowledge graph, and 3D scenes

The data plane handles device connections and messages. Devices normally authenticate to a Regional AWS IoT data endpoint with X.509 certificates. The control plane creates things, policies, certificates, Jobs, deployments, and security profiles. Administrators and AWS services authenticate to that plane through IAM. Giving a device an IAM user's long-lived access key confuses the planes and is unsafe.

2. Select the service before drawing the implementation

RequirementPrimary directionImportant boundary
Secure device connectivity, publish/subscribe, shadows, rulesAWS IoT CoreIt is not a time-series warehouse or fleet software runtime.
Provision, search, group, monitor, and update a device fleetAWS IoT Device Management capabilitiesA remote Job still needs safe device-side implementation.
Audit IoT configuration or detect unusual device metricsAWS IoT Device DefenderDetection does not prove compromise and automated mitigation can disconnect healthy devices.
Run modular workloads locally and tolerate cloud lossAWS IoT Greengrass V2Edge hardware, OS patching, disk, local users, and physical security remain yours.
Model industrial equipment and derived operational metricsAWS IoT SiteWiseModel quality, source protocol collection, and operational semantics remain engineering responsibilities.
Combine models and existing sources into an operational twinAWS IoT TwinMakerIt references and visualizes data; it is not the device broker or authoritative safety control.
Collect selected standardized vehicle signalsFleetWise only for an existing customerIt is closed to new customers; new designs assess Connected Mobility guidance and composable services.
Event state machines formerly implemented with IoT EventsCurrent rules, streams, Lambda, Step Functions, or another assessed patternIoT Events support ended May 20, 2026. Do not design a new dependency.
New edge applicationGreengrass V2Do not begin on V1; migrate existing V1 estates before end of support.

Service selection starts with protocol, connection pattern, offline requirement, volume, latency, safety consequence, ownership, and lifecycle. A digital twin presentation layer does not solve identity. A broker does not solve industrial semantics. An edge runtime does not remove the need for cloud authorization.

3. Understand IoT Core connectivity and MQTT

Retrieve the account's device data endpoint without exposing it in shared evidence:

aws iot describe-endpoint \
  --endpoint-type iot:Data-ATS \
  --query endpointAddress \
  --output text

Use the Amazon Trust Services endpoint. A typical embedded device uses MQTT with X.509 mutual TLS on port 8883, or port 443 with the required application-layer protocol negotiation. MQTT over WebSocket Secure commonly uses Signature Version 4 on port 443 for browser or IAM-authenticated applications. HTTPS supports device publication but does not provide MQTT subscriptions. Do not select a protocol only because a firewall permits its port.

MQTT uses topic strings such as:

telemetry/v1/tenant-17/device-0042/temperature
state/v1/tenant-17/device-0042/reported
command/v1/tenant-17/device-0042/request
command/v1/tenant-17/device-0042/result

Topic design is an authorization design. Put stable dimensions in a documented order, version the contract, limit topic length and cardinality, and separate command from telemetry. MQTT wildcards + and # belong in subscription filters. Broad filters and broad IoT policy resources can expose another tenant's data.

MQTT QoS 0 sends at most once. It has lower overhead but a message can be lost. QoS 1 sends at least once. It improves delivery assurance but duplicates can occur, so a consumer needs a message ID, device ID, event time, schema version, and idempotent write rule. Neither level means exactly-once business processing. A broker acknowledgement also does not prove that a motor acted safely.

Persistent sessions can preserve eligible subscriptions and queued QoS 1 messages for a disconnected client within service limits. Retained messages store the latest message on a topic for future subscribers. These are different features and neither replaces a device shadow, durable event store, or replay design. Verify current limits, expiry settings, and costs rather than assuming indefinite offline delivery.

4. Design device identity and least-privilege IoT policies

The common device identity is an X.509 certificate and private key. AWS IoT authenticates the certificate, checks whether it is active and attached appropriately, then authorizes the requested operation with an AWS IoT policy. An IAM policy for the engineer and an IoT policy for the device are separate policy systems.

Each manufactured unit should receive a unique key pair. Prefer generating the private key inside protected hardware when the platform supports it. Record non-secret manufacturing identity, certificate association, product model, firmware baseline, owner, and lifecycle status. Protect certificate-authority and fleet-provisioning credentials with much stronger controls than ordinary device credentials because their compromise can enroll a large fleet.

A conceptual policy binds the MQTT client and topic path to the connecting thing name:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "iot:Connect",
      "Resource": "arn:aws:iot:REGION:ACCOUNT:client/${iot:Connection.Thing.ThingName}"
    },
    {
      "Effect": "Allow",
      "Action": "iot:Publish",
      "Resource": "arn:aws:iot:REGION:ACCOUNT:topic/telemetry/v1/${iot:Connection.Thing.ThingName}/*"
    }
  ]
}

This is a teaching fragment, not a complete deployment policy. Test policy variables, attachment mode, thing-name/client-ID relationship, shadow topics, Jobs topics, and deny behavior in an isolated account. iot:Publish authorizes topic resources, while iot:Subscribe authorizes topic-filter resources and iot:Receive authorizes delivered topic resources. Confusing them produces either denial or excessive access.

Provisioning choices include individual factory registration, just-in-time registration/provisioning with a trusted CA, and fleet provisioning by claim or trusted user. A claim certificate is a bootstrap credential, not a permanent shared identity. Restrict what it can provision, protect and rotate it, monitor enrollment, and replace it with a unique device certificate as soon as provisioning succeeds.

Lifecycle controls must include:

  1. manufacture and key creation;
  2. enrollment and ownership assignment;
  3. certificate activation and least-privilege policy attachment;
  4. normal rotation before expiry or algorithm retirement;
  5. emergency revocation and quarantine;
  6. repair or ownership transfer with old-owner removal; and
  7. decommissioning that disables cloud identity and securely erases local secrets.

5. Use shadows for state, not as an event log

A device shadow is a cloud JSON document available even while the device is offline. Applications write desired state, devices write reported state, and the service calculates delta when they differ. Shadow metadata and version numbers help diagnose staleness and concurrent updates.

{
  "state": {
    "desired": {"sampleSeconds": 60},
    "reported": {"sampleSeconds": 300, "firmware": "2.4.1"}
  },
  "version": 87
}

The device should validate desired input, apply it safely, and report the result. It must not blindly copy arbitrary desired content into an actuator. Use optimistic version checks where concurrent writers matter. Reject stale commands, define who owns each field, and prevent application and device writers from continuously overwriting one another.

A classic shadow provides one default document. Named shadows separate concerns, for example connectivity, configuration, and maintenance. Separation can improve ownership but increases policy, lifecycle, and cost complexity. Shadows retain latest state; telemetry stores the history of observations elsewhere. Do not send every sensor reading through a shadow merely to create history.

6. Route data with the rules engine

An IoT rule evaluates MQTT messages with an SQL-like expression and invokes one or more actions. Common destinations include Lambda, S3, DynamoDB, Kinesis, SQS, SNS, CloudWatch, OpenSearch, and republished MQTT topics. Use a service role scoped to the exact destination. Configure an error action or equivalent failure route, and monitor invocation and failure metrics.

SELECT
  topic(4) AS deviceId,
  temperature,
  timestamp() AS receivedAt
FROM 'telemetry/v1/+/+/temperature'
WHERE temperature IS NOT NULL

Treat payloads as untrusted. Validate schema version, type, range, unit, event timestamp, device identity, and maximum size. Rules and downstream actions can retry or duplicate work. Use deterministic object keys or deduplication keys when business effects must be idempotent. Do not republish into a topic that triggers the same rule without a loop guard.

For high-volume architecture, decouple ingestion from slow consumers. A stream or queue absorbs bursts and gives replay or back-pressure behavior appropriate to the service. Partition choices must avoid concentrating all traffic on one tenant or device. The rules engine routes data; it does not make a weak data contract reliable.

7. Operate a fleet with groups, indexing, Jobs, and packages

Static thing groups express assigned membership such as product, plant, or release ring. Dynamic thing groups use fleet-indexing queries to identify devices matching current indexed attributes. Fleet indexing can include registry data, shadows, connectivity, Device Defender violations, and software-package information depending on configuration. Enabling more sources improves search but adds indexing cost and data-governance impact.

Use groups to create deployment rings:

engineering devices -> 1 percent canary -> 10 percent wave -> regional wave -> fleet

AWS IoT Jobs distributes a remote-operation document and tracks each execution. Jobs can support firmware installation, certificate rotation, diagnostics, or configuration, but the device agent performs the work. A safe agent validates signature and compatibility, checks power and disk, downloads over authenticated TLS, stages the artifact, verifies a cryptographic hash, installs atomically where possible, health-checks, rolls back, and reports a terminal state.

Configure rollout rate, timeout, retry, abort criteria, and target selection. Stop when a small wave exceeds a failure threshold. Never target the entire fleet first. Presigned download URLs expire; devices that remain offline need a renewal strategy. Code signing and secure boot strengthen different links: signing proves the approved artifact, while secure boot helps prevent untrusted code at startup.

Software package and package-version records can describe expected inventory. Device-reported or indexed inventory is evidence, not proof by itself. Compare it with deployment records, observed health, and vulnerability ownership.

Read-only fleet inventory examples:

aws iot list-things --max-results 20 --output table
aws iot list-thing-groups --max-results 20 --output table
aws iot get-indexing-configuration --output json
aws iot list-jobs --max-results 20 --output table
aws iot list-fleet-metrics --max-results 20 --output table

An empty result can mean an empty Region, denied visibility, pagination, or a different account. Record caller, Region, command, timestamp, and errors.

8. Detect security problems with Device Defender

Device Defender has two different jobs. Audit evaluates configuration controls, such as overly permissive policies or certificate practices. Detect compares reported device-side or cloud-side behavior metrics with a security profile and raises violations when behavior crosses defined thresholds or learned expectations.

Useful behavior signals include messages sent or received, authorization failures, connection attempts, listening ports, bytes transferred, and TCP connections where supported. A sudden rise may indicate compromise, a firmware bug, an expired certificate retry loop, or an operations event. A violation is a lead for investigation, not a verdict.

aws iot list-scheduled-audits --max-results 20 --output table
aws iot list-security-profiles --max-results 20 --output table
aws iot list-active-violations --max-results 20 --output table
aws iot list-mitigation-actions --max-results 20 --output table

Mitigation actions can update groups, policies, certificates, or Jobs depending on design. Automatic quarantine can stop an attack but can also disable an entire healthy fleet after a bad threshold. Begin with alert and human approval, test on canaries, preserve forensic data, make reversal explicit, and retain an out-of-band recovery route.

9. Run local workloads with Greengrass V2

A Greengrass core device runs the Greengrass nucleus or, for constrained supported scenarios, nucleus lite. A component is a versioned deployable software unit. Its recipe declares metadata, platform selection, dependencies, artifacts, configuration, lifecycle commands, and access requirements. Components can be private, AWS-provided, or community-provided; trust and support boundaries differ.

Greengrass supports local publish/subscribe and interprocess communication, local processing, connectors, stream buffering, secrets integration, and deployment from the cloud. It helps a factory gateway continue collecting and deciding during WAN loss. It does not guarantee unlimited offline storage. Define buffer size, eviction rule, event priority, local disk encryption, clock behavior, and replay order after reconnection.

A deployment targets one core or a thing group and resolves component versions and dependencies. Deployments are continuous desired state: changing a deployment can upgrade, downgrade, configure, or remove components. Pin tested component and nucleus versions where repeatability matters. An unbounded dependency request can select a newer patch during a later deployment change.

Roll out in rings and configure failure handling. A bad component can exhaust disk, crash repeatedly, monopolize CPU, or block a dependency. Run it as a dedicated least-privilege OS user where possible, limit local file and device access, declare IPC permissions narrowly, sign and scan artifacts, and collect component logs. Test restart, power loss, disk full, cloud loss, expired credentials, and rollback before fleet deployment.

aws greengrassv2 list-core-devices --max-results 20 --output table
aws greengrassv2 list-deployments --max-results 20 --output table
aws greengrassv2 list-components --scope PRIVATE --max-results 20 --output table

Cloud deployment and over-the-air nucleus update require functioning network, device certificate/key, credentials-provider access, and artifact access such as S3. Private VPC paths may therefore need IoT data, credential-provider, and S3 connectivity. A local application should enter a safe state when those dependencies fail.

Do not start new work on Greengrass V1. As of this lesson's review, AWS directs migration to V2 and states that V1 resources become unavailable after October 7, 2026. Verify the current notice again before planning any remaining migration.

10. Model industrial data with SiteWise

SiteWise collects, stores, organizes, and monitors industrial equipment data. A SiteWise gateway can collect from on-premises sources and send data to AWS. An asset model defines reusable equipment structure. Measurements receive raw values, transforms calculate values from properties, metrics aggregate over intervals, and hierarchies connect assets such as lines and machines.

plant
  +--> production line
         +--> mixer asset
                +--> motor temperature measurement
                +--> vibration measurement
                +--> running-state transform
                +--> availability or OEE-related metric

Data quality and unit are part of the contract. A value of 40 is meaningless without Celsius versus Fahrenheit, timestamp source, quality, sampling frequency, and asset identity. Gateway buffering helps with connection loss but needs disk sizing, retention, duplicate handling, and recovery tests. SiteWise Monitor can present operational portals, but dashboard visibility does not authorize a person to control machinery.

aws iotsitewise list-asset-models --max-results 20 --output table
aws iotsitewise list-assets --max-results 20 --output table
aws iotsitewise list-gateways --max-results 20 --output table

After IoT Events end of support, do not assume an old SiteWise alarm design remains the correct current path. Review AWS's migration guidance and choose a supported detection and notification pattern with equivalent state, deduplication, acknowledgement, and escalation behavior.

11. Add context with TwinMaker and handle specialized services

TwinMaker builds operational digital twins of physical and digital systems. A workspace contains entities, component types, components, relationships, scenes, and resources. Connectors can reference time-series and contextual data in SiteWise, Kinesis Video Streams, or custom sources. A knowledge graph connects equipment, spaces, processes, documents, and observations. Scene Composer overlays that context on web-optimized 3D assets, and Grafana can expose an application to operators.

TwinMaker normally references data where it already lives. It is not a duplicate universal database, device identity service, MQTT broker, or substitute for human safety monitoring. Define source-of-truth ownership, connector role permissions, source latency, stale-data display, model versioning, workspace isolation, and behavior when a source is unavailable.

FleetWise models vehicle signals, decoders, vehicles, fleets, and campaigns, with an edge agent selecting valuable data for cloud transfer. However, it is no longer open to new customers. Existing customers can continue, but a new customer must evaluate AWS Connected Mobility guidance and build the needed capability from currently available components. Vehicle designs also require privacy, driver consent, safety, data residency, bandwidth, and over-the-air risk review.

Other IoT-adjacent choices include IoT Core for LoRaWAN for compatible long-range gateway/device networks, Amazon Sidewalk integration for eligible products and Regions, and Kinesis Video Streams for video or WebRTC. Availability, hardware ecosystem, radio regulation, and support vary. Do not add a specialized service merely because the device is called IoT.

12. Design the 10,000-device reference architecture

Scenario: a manufacturer operates 10,000 refrigerated-display controllers across stores. Each sends temperature and compressor state every minute, needs a configuration shadow, can be offline for eight hours, and receives quarterly signed software. A dangerous temperature must alert operators, but cloud loss must not disable the local cutoff.

Create aws306-architecture.md with these decisions:

AreaRequired decision and evidence
Device identityUnique hardware-protected key, certificate-to-thing association, bootstrap process, rotation, revocation, transfer, retirement owner
Topic contractVersioned telemetry, status, command, and result topics; tenant/device authorization; schema and maximum payload
Delivery semanticsQoS choice, deduplication key, event time, late event rule, session/expiry behavior, eight-hour local queue
StateShadow fields and owners, desired validation, reported confirmation, version-conflict handling
IngestionIoT rule, validation, durable stream or queue, raw archive, error route, replay and idempotency
Edge safetyLocal cutoff and alarm behavior independent of cloud; Greengrass only if gateway processing is justified
Fleet operationsStatic release rings, dynamic unhealthy group, Job rollout rate, abort threshold, signed artifact, rollback image
Security detectionDefender audits, behavior profiles, alert triage, approved quarantine and reversal
ObservabilityConnect/auth failures, rule errors, queue age, offline fleet percentage, Job success, component health, stale shadows
ResilienceRegional endpoint dependency, DNS/TLS/clock failure, WAN outage, reconnect jitter, cloud recovery, support escalation
Cost and quotaPer-minute message count, payload-size billing units, rules actions, indexing, logs, storage, transfer, edge hardware, quota headroom
RetirementDisable certificate, detach policy, erase key, preserve required records, remove indexes and data under retention policy

At one telemetry message per minute, 10,000 devices produce 14.4 million messages per day before command, shadow, Jobs, retries, and duplicate traffic. Calculate the actual billable message units from payload size and current pricing. Model normal, outage-reconnect, and firmware-rollout peaks. Request quota increases before deployment and add randomized exponential reconnect backoff so all devices do not reconnect simultaneously.

Failure exercise: assume an issuer certificate expires, 30 percent of gateways have incorrect clocks, the WAN is unavailable for eight hours, and release 2.5.0 crashes on one hardware revision. Show which local controls continue, what evidence appears, how rollout aborts, how devices return to 2.4.1, how queued data drains without overwhelming ingestion, and who authorizes certificate recovery.

13. Diagnose from evidence

SymptomEvidence to inspectLikely boundarySafe response
TLS connection failsDevice clock, CA chain, endpoint/SNI, certificate status, Region, device logsNetwork, time, trust, or identityCorrect clock/trust/endpoint; never disable certificate validation.
Connect succeeds but publish is deniedIoT logs, client ID, topic, policy variable expansion, certificate attachmentsIoT authorizationNarrowly correct action/resource matching; do not grant iot:* on *.
Device disconnects repeatedlyDuplicate client ID, keepalive, network loss, throttling, reconnect intervalMQTT session or client behaviorGive unique IDs, add jitter/backoff, inspect broker reason and metrics.
Messages are missing or duplicatedQoS, session expiry, offline duration, retained setting, consumer dedupe keyDelivery semanticsAccept the selected semantics and make downstream effects idempotent.
Shadow oscillatesUpdate documents, versions, desired/reported ownership, multiple writersState contractStop competing writer, use version checks, assign field ownership.
Rule does not deliverRule SQL match, role trust/permissions, action errors, error route, destination healthRules engine or destinationTest synthetic message, repair least privilege, replay from durable source.
Job remains queued or in progressTarget membership, device subscription, presigned URL expiry, execution timeout, agent logsJob agent, network, artifactPause rollout, renew safely, diagnose a canary, preserve rollback.
Greengrass component failsNucleus/component logs, recipe platform, dependency versions, disk, OS user, artifact accessEdge runtime or componentStop wave, restore pinned version/configuration, verify local safe state.
Defender reports a fleet spikeMetric baseline, deployment/change record, affected group, raw logsAttack, bug, or expected eventInvestigate before mass quarantine; preserve reversible mitigation.
SiteWise values look wrongSource timestamp, unit, quality, property alias, gateway queue, transformSource or industrial modelQuarantine bad series, correct mapping/model, backfill only with provenance.
Twin scene is staleSource connector, component query, role, source latency, UI cacheData source or presentationMark data stale, repair connector, do not infer safe physical state.

Enable IoT logging deliberately. DEBUG can be useful for a small named group but expensive and sensitive at fleet scale. Prefer scoped temporary diagnostic logging, redact payloads, define retention, and return to the approved steady-state level. CloudTrail records control-plane changes; it is not the complete device message history.

14. Cost, quotas, and cleanup

Cost can arise from connectivity duration, messages and payload units, rules actions, registry and fleet operations, indexing, Jobs, Device Defender audits/detection, Greengrass active devices, SiteWise ingestion/storage/queries/monitors, TwinMaker queries and resources, CloudWatch logs, downstream compute/storage, data transfer, and cellular service. Pricing dimensions and free offers change, so record Region, date, pricing link, traffic assumptions, and uncertainty.

For the T0 path, prove no creation with before/after inventories and preserve only redacted local evidence. For an approved T1 simulator, delete only resources carrying the lesson ownership tag and exact recorded IDs. Stop subscriptions, Jobs, and deployments before removing their dependencies. Detach policies and certificates, deactivate the test certificate, delete the test thing, remove rules and destinations, remove Greengrass test deployments/components where owned, and verify that no log group, bucket object, stream, SiteWise asset, or alarm remains unexpectedly billed.

Never run broad wildcard deletion. A thing may represent a physical device still using its credential. Keep the certificate private-key deletion and physical-device reset as separate verified actions.

Practical submission

Submit one redacted evidence bundle containing:

  1. a one-page plain-language glossary;
  2. a portfolio selection table including the current status of Events, FleetWise, and Greengrass V1;
  3. the 10,000-device architecture and data-flow diagram;
  4. an identity lifecycle from manufacture through retirement;
  5. a topic tree and least-privilege policy explanation;
  6. a QoS, offline queue, duplicate, replay, and reconnect plan;
  7. a shadow schema with writer ownership and conflict handling;
  8. a rules path with schema validation, error route, and idempotency;
  9. fleet groups and staged Job rollout/abort/rollback design;
  10. Device Defender audit, detect, triage, and reversible mitigation design;
  11. Greengrass component/deployment and disconnected-operation design;
  12. a SiteWise asset model and a reasoned TwinMaker use or rejection;
  13. read-only CLI evidence with caller, Region, timestamps, and errors;
  14. completed diagnosis for the four-part failure exercise;
  15. normal and peak volume, quota, and cost worksheet; and
  16. cleanup proof or T0 no-create proof plus an architecture decision record.

Knowledge check

  1. Why is a thing not the same as a physical device?

Expected direction: the thing is a cloud registry representation; physical keys, firmware, ownership, and erasure have separate lifecycle evidence.

  1. Why can QoS 1 still produce duplicate business effects?

Expected direction: it is at-least-once message delivery, so consumers and effects require idempotency.

  1. When is a shadow better than a telemetry stream?

Expected direction: for durable latest desired/reported state while a device may be offline, not historical observations.

  1. Why must iot:Subscribe and iot:Receive be considered separately?

Expected direction: one authorizes topic filters and the other delivery on topic resources.

  1. What does fleet indexing prove?

Expected direction: searchable indexed state from enabled sources at that time, not physical health or trustworthy self-report by itself.

  1. Why can an automatic Defender mitigation be dangerous?

Expected direction: a false positive or shared bug could revoke or quarantine healthy devices at fleet scale.

  1. What makes a Job safe enough for software rollout?

Expected direction: canaries, signatures, prechecks, bounded rate, abort criteria, health confirmation, and rollback.

  1. What does Greengrass add during WAN loss?

Expected direction: local components, messaging, processing, and bounded buffering, provided local dependencies and safe-state design work.

  1. When do SiteWise and TwinMaker differ?

Expected direction: SiteWise organizes industrial measurements and assets; TwinMaker connects contextual models and sources into an operational twin and visualization.

  1. Why should a new architecture reject IoT Events and FleetWise?

Expected direction: Events ended support; FleetWise is closed to new customers, so currently supported alternatives must be assessed.

Lesson acceptance

You pass when a reviewer can trace one device from manufacture, authentication, connection, topic authorization, message delivery, state synchronization, fleet operation, detection, edge behavior, industrial context, failure recovery, cost, and retirement without an unexplained trust or ownership gap. The design must use unique identities, distinguish control and data planes, tolerate duplicate and delayed data, preserve local safety during cloud loss, stage remote updates, document current service lifecycle status, and include reversible incident actions.

Fail the lesson if it uses shared production certificates, grants wildcard fleet access without a justified containment boundary, assumes QoS 1 is exactly once, treats a shadow as unlimited history, deploys to all devices first, relies on cloud messaging as the physical emergency stop, selects a retired/unavailable service for new work, omits rollback, or presents an inventory screenshot as proof that the system works.

Official sources

Advertisement