AWS 149: Amazon Kinesis, MSK and Amazon MQ
Why this lesson matters
Choose stream or broker technology from ordering, replay, throughput, protocol compatibility, operations, and consumer requirements.
What you will be able to do
By the end, you can:
- explain amazon kinesis, msk and amazon mq in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Choose stream or broker technology from ordering, replay, throughput, protocol compatibility, operations, and consumer requirements. |
| Scope and boundary | The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Kinesis Data Streams, Amazon MSK, and Amazon MQ. |
| Evidence of success | Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Kinesis Data Streams, Amazon MSK, and Amazon MQ. |
| Cost model | Provisioned capacity, broker hours, storage, throughput, cross-AZ transfer, enhanced fan-out, connectors, and monitoring can charge; these are T0 design exercises. |
| Safe rejection rule | Avoid choosing Kafka for every queue, using one partition key that creates a hot shard, or ignoring broker client compatibility. |
How the request flows
+----------------------+
| Producer protocol |
+----------------------+
|
v
+-------------------------------------+
| Partition, topic, or broker queue |
+-------------------------------------+
|
v
+----------------------+
| Consumer group |
+----------------------+
|
v
+-------------------------------------+
| Lag, replay, and failure evidence |
+-------------------------------------+
For Kinesis Data Streams, Amazon MSK, and Amazon MQ, the important boundary is this: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Kinesis Data Streams, Amazon MSK, and Amazon MQ. Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Kinesis Data Streams, Amazon MSK, and Amazon MQ. That is why the lesson pairs the Console with CLI output and a practical artifact. One interface may hide a field, use a cached view, or be scoped differently. Matching evidence is stronger than a screenshot alone.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use the service that matches required protocol, ordering scope, replay, scale, and team operating model. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid choosing Kafka for every queue, using one partition key that creates a hot shard, or ignoring broker client compatibility. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
Three different compatibility contracts
Kinesis Data Streams is an AWS-native retained stream. Producers choose a partition key; its hash maps records to shards/internal capacity and ordering is per shard/partition-key path, not global. Consumers maintain checkpoints and replay from sequence/time positions before retention expires. Provisioned mode sizes shards; current on-demand modes scale managed capacity with different performance/cost characteristics. Shared-throughput polling consumers contend for shard read capacity; enhanced fan-out registers consumers with dedicated push throughput and added cost. PutRecords returns per-record success/failure, so retry only failed entries with the same logical identity and tolerate duplicates.
Amazon MSK runs Apache Kafka-compatible clusters. Topics have partitions, replicas and retention; ordering is per partition. Producer acknowledgments/idempotence/transactions, replication factor and minimum in-sync replicas shape durability. Consumer groups assign a partition to one active consumer in a group, commit offsets and rebalance. MSK Serverless, Provisioned Standard and Provisioned Express brokers have different capacity, storage, configuration, scaling, feature and cost models; Express is not simply “faster Standard” and currently restricts some Kafka features. AWS manages brokers, but teams still own topic design, partitions/skew, client versions, schema compatibility, quotas, ACL/IAM/SASL/TLS, upgrades, lag and connector behavior.
Amazon MQ is for compatible message-broker protocols and semantics when migrating/integrating ActiveMQ or RabbitMQ applications. Queues, topics/exchanges, routing, acknowledgements, prefetch, persistence, redelivery and dead-letter behavior follow the selected engine/protocol and version. Single-instance deployment is for development or accepted downtime; active/standby or cluster deployment choices and storage/network design determine availability. Managed broker patching does not remove client reconnect, failover, flow control or poison-message responsibilities.
Use SQS/SNS/EventBridge when AWS-native queue/fanout/routing meets the requirement with less broker operation. Choose Kinesis for ordered retained AWS stream processing, Kafka/MSK for Kafka ecosystem/API and independent consumer replay, and MQ for existing JMS/AMQP/MQTT/STOMP/OpenWire or RabbitMQ exchange/queue semantics as supported. Protocol compatibility must be proved with actual clients/plugins/features.
Partitioning, retention and failure
Partition keys distribute throughput and define order. A tenant/date constant can hot-spot one shard/partition while fleet averages look healthy. Calculate records/s and bytes/s both before and after encoding/overhead, identify the largest key, and preserve enough key cardinality. More partitions improve parallelism but increase open files, metadata, rebalance/checkpoint and cost; they do not accelerate one ordered key.
Retention is a replay window, not an indefinite system of record. Consumer outage tolerance must be shorter than retention with alarm margin; archive to S3 when longer history is required. Monitor Kinesis write/read throttles and iterator age, Kafka producer errors/under-replicated partitions/ISR/controller/broker disk-CPU-network and consumer lag, and MQ enqueue/dequeue/unacked/dead-letter/storage/connection plus broker health. Encrypt in transit/at rest, use private subnets, multi-AZ placement, scoped IAM/SASL/SCRAM/mTLS or broker users, Secrets Manager/KMS and audited admin access.
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Use the Console service search and open Kinesis, MSK, and Amazon MQ resource lists; confirm the account and Region before reading the page.
- Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
- Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
- Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws kinesis list-streams --output table
aws kafka list-clusters-v2 --query 'ClusterInfoList[].{Name:ClusterName,Type:ClusterType,State:State}' --output table
aws mq list-brokers --query 'BrokerSummaries[].{Name:BrokerName,Engine:EngineType,Mode:DeploymentMode,State:BrokerState}' --output table
Expected interpretation
The inventory separates a managed shard-based stream, managed Kafka, and ActiveMQ or RabbitMQ-compatible brokers. It does not prove protocol fit or end-to-end processing.
Practical work
Compare telemetry replay, existing Kafka analytics, and a legacy JMS application. Choose one service for each, then specify partition key, ordering, retention, consumer state, HA, security, and cost.
Add a numeric model for average/peak records, byte size, largest key, consumers, replay window and outage catch-up. Test hot partition, partial producer batch failure, duplicate record, consumer crash before/after checkpoint, retention expiry, Kafka rebalance, broker-AZ loss, invalid TLS/auth and DLQ recovery. Compare Kinesis provisioned/on-demand and shared/enhanced fan-out; MSK Serverless/Standard/Express; ActiveMQ/RabbitMQ deployment. Include schema registry/evolution, cross-Region DR, migration rollback and all idle/throughput/storage/transfer/monitoring costs.
Diagnose this topic from its own evidence
High Kinesis iterator age with write success means consumers are behind/blocked; write throttles isolated to keys suggest partition skew. Inspect failed entries in every producer batch. Kafka lag can come from slow consumers, too few partitions, rebalance loops, broker saturation or downstream failure; correlate partition-level lag and broker/ISR metrics. MQ queue growth with stable enqueue and low acknowledgement means consumer/prefetch/processing failure; connection churn suggests DNS/TLS/auth/failover. Always identify stream/topic/partition/shard, sequence/offset/message ID and consumer group before scaling.
Cost and cleanup
Provisioned capacity, broker hours, storage, throughput, cross-AZ transfer, enhanced fan-out, connectors, and monitoring can charge; these are T0 design exercises.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Choose stream or broker technology from ordering, replay, throughput, protocol compatibility, operations, and consumer requirements.
- Which scope or ownership boundary must be proved first?
Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Kinesis Data Streams, Amazon MSK, and Amazon MQ.
- What evidence is strong enough to accept the result?
Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Kinesis Data Streams, Amazon MSK, and Amazon MQ.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid choosing Kafka for every queue, using one partition key that creates a hot shard, or ignoring broker client compatibility.
- Which cost dimensions and retained resources need an owner?
Expected direction: Provisioned capacity, broker hours, storage, throughput, cross-AZ transfer, enhanced fan-out, connectors, and monitoring can charge; these are T0 design exercises.
Lesson acceptance
Pass when the learner selects Kinesis, MSK or MQ from protocol/ecosystem/order/replay evidence; models partition capacity/skew and consumer state; and designs retention, HA/DR, security, observability, poison/duplicate behavior and cost. Fail if Kafka is selected as a generic queue, order is claimed across partitions, producer batch HTTP success hides failed entries, or managed brokers are treated as operation-free.