AWS 400: SNS, ticket, chat, and stakeholder notification design
Why this lesson matters
A machine event, human page, ticket, collaboration message and stakeholder update have different audiences and guarantees. Sending one alarm to every channel causes noise, leaks sensitive data and still may not produce acknowledgement. Design communication around urgency, action, ownership, delivery evidence and accessibility.
Channel contract
| Channel | Purpose | Required behavior |
|---|---|---|
| Event bus/queue | Machine integration | Structured schema, retry, DLQ, idempotency |
| Page/on-call | Immediate human action | Ack/escalation, concise impact and runbook |
| Ticket/OpsItem | Durable ownership/backlog | Dedup/update, severity/SLA, closure evidence |
| Chat room | Responder coordination | Incident channel, timestamps, decisions, access |
| Email/SMS/push | Reach/fallback | Tested address, privacy and delivery limitations |
| Stakeholder/status | Business/customer understanding | Impact, scope, workaround, next update, no speculation |
Severity is based on customer/business/security impact and urgency, not the emitting service’s adjective. Build a routing policy for severity, environment, service owner, time zone, maintenance, and existing incident correlation. One incident key updates the same page/ticket/chat rather than creating a storm.
SNS delivery and filtering
Use separate encrypted topics when audience, data classification, Region, account or delivery policy differs. Topic policy controls publishers; subscription endpoint policy and KMS key policy must also allow delivery. Subscription filters can match message attributes or body properties and are eventually consistent after changes, so deploy/test before relying on a new route.
SNS delivery is at least once and retry behavior differs by endpoint type. Attach an encrypted SQS DLQ to each eligible subscription, not the topic; monitor DLQ messages and delivery-failure metrics. A published message or healthy subscription does not prove a human saw or acknowledged it. Do not expose unsubscribe URLs, task tokens, secrets, raw findings or customer data.
For chat integration, use current Amazon Q Developer in chat applications guidance and supported message format; avoid unsupported raw-delivery settings. Chat is not the system of record. Restrict workspace/channel, guard interactive command roles, log actions and provide a fallback when chat or identity is unavailable.
Message design and incident lifecycle
A page contains incident key, severity, detected UTC time, service/environment/Region, user impact and evidence, current state, safe first action, dashboard/runbook and acknowledgement link. It avoids huge JSON. A stakeholder update states confirmed impact, affected users/regions, start/time confidence, mitigation, workaround, next update and owner, clearly separating facts from investigation.
Lifecycle: detected, acknowledged, investigating, identified, mitigating, monitoring, resolved, retrospective. Every transition updates ticket/timeline and appropriate audiences. Silence/dedup windows must not hide a worsening incident. Define handoff, missed acknowledgement escalation, comms lead, legal/security review and accessible alternatives.
Keep an audience matrix with contact owner, approved data classification, language/time-zone/accessibility needs, primary/fallback channel, acknowledgement expectation, escalation delay and quarterly test date. Automate roster synchronization where possible but require review for departed staff and external recipients. During identity-provider or chat outage, responders need a preapproved secondary path that does not rely on the same failed control plane.
Measure detection-to-page, page-to-acknowledgement, escalation success, duplicate/suppressed count, delivery/DLQ failure, ticket correlation, update cadence and stakeholder correction rate. Review notifications after every serious incident. Remove messages that did not change a decision, but preserve signals required for audit or safety. Test templates with screen readers and mobile display rather than assuming rich formatting survives every endpoint.
Read-only inspection and workshop
aws sns get-topic-attributes --topic-arn TOPIC_ARN --region ap-south-1
aws sns list-subscriptions-by-topic --topic-arn TOPIC_ARN --region ap-south-1
aws sns get-subscription-attributes --subscription-arn SUBSCRIPTION_ARN --region ap-south-1
aws sqs get-queue-attributes --queue-url DLQ_URL --attribute-names All --region ap-south-1
Design notifications for availability outage, security finding, failed backup, cost anomaly and nonurgent patch drift. Define event envelope, routing/dedup, SNS topics/policies/KMS, filter fixtures, subscription retry/DLQ, pager/ticket/chat integrations, acknowledgement/escalation and stakeholder templates.
Test 20 failures: severity wrong, owner missing, duplicate pages, stale incident update, SNS publisher deny, KMS deny, filter mismatch, filter propagation delay, endpoint deleted, HTTP 4xx, endpoint throttled, DLQ permission deny, DLQ unmonitored, email delivered to old list, SMS truncates, chat raw format fails, interactive role broad, secret in message, acknowledgement absent, and resolved notice sent while SLI remains bad.
Cost and acceptance
Price SNS requests/delivery by protocol, SMS/carrier, SQS DLQ, Lambda/API/chat/ticket integrations, KMS, logs, paging vendor and human interruption. This lesson creates nothing.
Submit channel matrix, severity/routing policy, secure SNS design, filter/retry/DLQ tests, message templates, incident state machine, escalation/fallback, 20 failures, cost and retention/privacy. Pass requires action-oriented paging, human acknowledgement, no sensitive leakage, durable record and tested delivery failure.