Lesson 388 · AWS Learning Path

AWS 388: Resilience Hub policies, assessments, and recommendations

· Published · 4 min read

Labelled process diagram for AWS 388: Versioned intent to Automated validation to Controlled AWS change to Observed result and retained evidence, with decision, proof and rejection evidence.

Why this lesson matters

AWS Resilience Hub models an application, compares estimated workload RTO/RPO against a resilience policy, and recommends alarms, SOPs, and tests. An assessment is design evidence, not proof that people, automation, data, capacity, and dependencies actually recover within objectives.

Application and policy model

An application is a versioned collection of resources organized into application components. Resources may be imported from supported sources and grouped to reflect workload behavior. Publish a new application version before assessing it. Missing dependencies, shared services, external systems, or stale resource collections produce misleading conclusions.

Disruption scopePolicy questionRequired live evidence beyond assessment
Application/softwareCan bad code/config recover?Deployment rollback/forward game day
Infrastructure/hardwareCan failed resources replace?Health replacement and capacity test
Availability ZoneCan surviving zones carry load/data?AZ impairment test and load evidence
RegionCan another Region recover and receive traffic?Artifact/data/secret/capacity failover exercise

A resilience policy defines targeted RTO/RPO by disruption type. Objectives come from business impact and dependencies, not from current architecture estimates. Zero targets can be configured but estimated values are near zero and can breach a literal zero policy. Document measurement start/end and data-loss boundary.

Assessment interpretation

The assessment analyzes discovered configuration, estimates component/workload recovery, compares it to policy, and produces compliance plus recommendations. Review every component, unsupported/missing resource, estimate assumption, previous-assessment drift, and recommendation. A high score or compliant status does not guarantee recoverability.

Validate resource discovery against an independently maintained architecture and data-flow diagram. Count expected versus imported resources by account, Region, stack, tag, and component. Explicitly model shared identity, DNS, network, observability, artifact, secret/key, data replication, external provider, and operator dependencies. Assign a confidence level when a dependency cannot be represented and prevent that gap from being hidden by the overall compliance label.

Alarm recommendations improve detection; SOP recommendations offer recovery automation direction; test recommendations can integrate FIS or other validation. Generated/recommended artifacts need owner review, IAM/security checks, parameterization, stop conditions, version control, test results, and lifecycle. Do not deploy broad automation simply to increase a score.

Changes require republishing/reassessment. Scheduled daily assessment can detect posture drift, but result freshness, application-version drift, policy change, and unresolved recommendation age need alerts. Distinguish Resilience Hub application drift from CloudFormation resource drift and runtime data drift.

Govern recommendations through normal engineering change: triage benefit and assumptions, map owner and due date, test in a representative environment, measure actual recovery, update runbooks, and reassess. Reject recommendations that conflict with business consistency, security, cost, or service constraints, but record the rationale and compensating control. Track breached objectives and overdue high-risk gaps as operational risk, not merely backlog.

Read-only inspection and workshop

aws resiliencehub list-apps
aws resiliencehub describe-app --app-arn APP_ARN
aws resiliencehub list-app-versions --app-arn APP_ARN
aws resiliencehub list-app-assessments --app-arn APP_ARN
aws resiliencehub describe-app-assessment --assessment-arn ASSESSMENT_ARN
aws resiliencehub list-alarm-recommendations --assessment-arn ASSESSMENT_ARN
aws resiliencehub list-sop-recommendations --assessment-arn ASSESSMENT_ARN
aws resiliencehub list-test-recommendations --assessment-arn ASSESSMENT_ARN

Given a supplied two-Region API assessment, verify inventory against architecture, map shared/external dependencies, compare estimates to business targets, classify recommendations by risk/value/owner, and create a remediation backlog. For each target define real evidence: backup restore, Auto Scaling/AZ test, dependency impairment, control-plane unavailability, regional data/secret/artifact recovery, DNS traffic, failback, and measured user outcome.

Create a traceability table from business capability to component, resource, failure scope, target RTO/RPO, estimated RTO/RPO, alarm, SOP, test, last execution, actual result, gap, owner, and due date. Reassess only after changes are deployed and discovered correctly.

Failure game day

Analyze 16 cases: resource import omitted, shared dependency missing, wrong component grouping, stale published version, policy targets unapproved, Region target omitted, unsupported resource assumed safe, assessment permission failure, estimated RTO mistaken for observed, recommendation IAM broad, alarm absent, SOP untested, test has no stop condition, daily assessment disabled, compliance drifts, and score improves while user recovery fails.

For each identify misleading signal, additional evidence, customer risk, remediation, owner, and acceptance test. Keep assessment artifacts and game-day evidence with timestamps because architecture and recommendations evolve.

Cost and acceptance

Price current Resilience Hub assessment/application usage, recommended alarms/SOP automation/FIS tests, logs, duplicate recovery resources, and engineering exercises. Confirm current pricing. This lesson creates nothing.

Submit application/component inventory, policy rationale, report critique, recommendation backlog, business-to-test traceability, 16 failure analyses, measured-test plan, cost, and governance cadence. Pass requires complete dependency scope, business-approved objectives, no assessment-as-proof claim, tested operational artifacts, and observed recovery evidence.

Official sources

Advertisement