AWS 235: Resilience Hub assessment and operational recommendations
Why this lesson matters
A backup, Multi-AZ label or architecture diagram does not prove that a customer journey can survive a failure. Resilience starts with a business objective, an accurate dependency model and named owners. It becomes credible only when alarms detect the failure, runbooks recover the service, controlled tests measure the result, and a new assessment confirms that the intended change is represented.
AWS Resilience Hub brings those activities into one service. It can assess a modeled workload, identify likely failure modes and recommend operational controls. Its result is decision support: an estimated RTO/RPO, recommendation, finding or high score is never the same as a successful recovery exercise.
Outcomes
You will be able to:
- explain resilience, availability, disaster recovery, RTO and RPO from zero;
- distinguish the existing application experience from next-generation systems,
user journeys and services;
- choose a model and policy that match the business rather than the resource list;
- explain resource import, discovery, AppComponents and dependency blind spots;
- publish/version a model and interpret assessments, findings and drift;
- evaluate alarm, Systems Manager SOP and Fault Injection Service recommendations;
- explain what the classic resilience score measures and what it cannot prove;
- inspect both
resiliencehubandresiliencehubv2safely with the AWS CLI; - design IAM, cross-account, data-protection and cost ownership boundaries;
- diagnose incomplete topology, policy breaches, stale assessments and failures;
- turn recommendations into controlled changes, tests and reassessment evidence.
Safety boundary and workbook
- This lesson is read-only. Do not create or publish an application, start an
assessment, deploy recommendations, run an SOP or start a fault experiment.
- Use only an explicitly owned workload. Resource discovery and dependency data
can reveal sensitive names, topology, tags, accounts and third-party services.
- Never use a resilience assessment as authorization for a production change.
Each recommendation needs an owner, security/cost review, change approval, rollback plan and independently defined success and stop conditions.
- Redact account IDs, internal DNS names, resource identifiers, customer data and
generated reasoning from submitted evidence.
Download the Resilience Hub assessment workbook or complete archive.
Foundations: resilience is a business property
Availability asks whether a service is usable now. Resilience is the ability to resist or recover from disruption while meeting an agreed business outcome. Disaster recovery (DR) focuses on restoring service after a major event. They overlap but are not synonyms.
| Term | Plain meaning | Example evidence |
|---|---|---|
| RTO | maximum acceptable interruption time | checkout restored in 30 minutes |
| RPO | maximum acceptable data loss measured in time | at most 5 minutes of orders lost |
| SLO | measurable reliability objective over a period | 99.9% successful checkout/month |
| failure mode | a specific way a user journey can fail | writer unavailable, queue backlog, expired certificate |
| control | prevention, detection or recovery mechanism | Multi-AZ, alarm, tested runbook |
| residual risk | risk remaining after controls | Region loss requires manual DNS approval |
RTO is not merely the infrastructure restore duration. Measure from disruption through detection, diagnosis, decision, recovery, dependency readiness, data validation, routing/cutover and business acceptance. RPO is not backup frequency: it is the age of the newest usable and validated state after reconciliation.
Product history and the two experiences
AWS Resilience Hub originally organized work around applications, AppComponents, resiliency policies and assessments. On May 28, 2026, AWS announced the generally available next-generation experience. Existing customers can continue with the original experience and adopt the new one at their own pace. Therefore screenshots, APIs and terminology can legitimately differ.
| Existing experience | Next-generation experience |
|---|---|
| application | service (closest conceptual mapping) |
| application resources grouped into AppComponents | systems contain user journeys and services |
| static resource checks and estimated disruption RTO/RPO | GenAI-powered failure-mode assessments |
| one resiliency policy across disruption classes | modular availability, DR and data-recovery policies |
CLI namespace aws resiliencehub | CLI namespace aws resiliencehubv2 |
Do not mix identifiers or commands between experiences. First record which model the account uses, its Region availability, and which team owns migration. A next-generation system expresses a business boundary; a user journey is a customer or operator flow; a service is a deployable capability and maps most closely to an old application.
End-to-end control loop
business journey + RTO/RPO/SLO + owner
|
v
accurate resources, data-flow edges and dependencies
|
v
published application version OR current service topology
|
v
assessment -> breach/failure-mode findings -> recommendations
|
v
reviewed IaC + alarms + SOPs + controlled fault tests
|
v
measured recovery/data loss -> reassess -> accept residual risk
Every arrow needs evidence. Resilience Hub does not operate the application for you, approve generated changes, guarantee business continuity or replace game days, backup restores, security reviews and incident management.
Model the workload before assessing it
Start with a user journey such as “place an order.” Trace its entry point, authentication, synchronous calls, queues, databases, object stores, DNS, secrets, certificates, observability and third-party dependencies. Separate:
- hard dependencies whose failure stops the journey;
- soft dependencies with degradation, cache or manual fallback;
- sources of truth from disposable caches and derived indexes;
- zonal, Regional, cross-Region and external failure boundaries;
- shared resources whose owner or blast radius crosses services;
- control-plane dependencies needed only during recovery.
Existing applications and AppComponents
Resources can be imported from supported sources such as CloudFormation stacks, resource groups/tags, Terraform state and EKS. Resilience Hub groups resources that fail and recover together into AppComponents. Automatic grouping is a starting hypothesis. Incorrect components produce misleading estimates, so the application owner must review missing, extra, unsupported and shared resources.
An application draft is not assessable evidence. Publish the application to create a version, then assess that published version. If resources change after assessment, record drift and publish/reassess rather than presenting old results.
Next-generation systems, journeys and services
Next-generation resource discovery can use CloudFormation, Terraform state in S3, tags and EKS. A tag source containing several tags matches them with AND; separate tag sources are combined with OR. Broad tags can silently import unrelated resources, while inconsistent tags omit dependencies.
A failure-mode assessment needs a valid topology, including at least two resources connected by a data-flow relationship, and an invoker role. Service creation can consider up to five Regions. Record exact source versions because a mutable stack, state file or tag query can produce a different model later.
Dependency discovery uses observed DNS/query behavior and a documented 35-day lookback. It can identify internal AWS and third-party dependencies, but it is not omniscient. Direct IP calls, inactive paths and shared-VPC attribution may be missed. A discovered edge is evidence to review, not proof of business criticality.
Define policy from business impact
Existing resiliency policy
The existing experience compares estimated workload recovery against target RTO and RPO for four disruption classes:
- application/software;
- infrastructure/hardware;
- Availability Zone;
- Region.
A zero target is accepted, but an estimated RTO or RPO cannot be exactly zero, so that objective will be breached. “Not applicable” needs a business reason; it must not hide an unsupported or unmodeled dependency. Policy selection should come from business-impact analysis, not the architecture's current capability.
Next-generation modular policies
The new model separates availability SLO, disaster recovery and data recovery policy components. Apply the components relevant to each service and journey. This allows different services to carry different objectives without pretending that every resource has the same criticality.
For either experience, record the objective's owner, measurement window, assumptions, scope and exception expiry. A policy is a requirement, not a control.
What an assessment actually does
Existing assessment
An assessment evaluates the imported application version against its attached policy. It estimates RTO/RPO from supported resource configurations and generates recommendations for resilience, CloudWatch alarms, Systems Manager standard operating procedures (SOPs) and AWS Fault Injection Service (FIS) tests.
Compliance/status can include Assessed, Not assessed, Policy breached and Changes detected. Read the component-level reason and recommendation; do not reduce the result to one color. Excluding a recommendation changes score only after reassessment, and exclusions must have an owner and rationale.
Classic resilience score
The maximum classic score weights recommendation coverage approximately as:
| Category | Maximum contribution |
|---|---|
| policy compliance | 40 points |
| implemented alarms | 20 points |
| implemented SOPs | 20 points |
| implemented tests | 20 points |
A score of 100 means the service recognized the recommended coverage. It does not prove an alarm fired, a runbook succeeded, a fault remained contained, recovery met RTO/RPO, data was correct or operators were available. Preserve raw details, assessment version/time and measured test evidence beside the score.
Next-generation failure-mode assessment
The next-generation assessment analyzes topology and configuration to identify failure modes and findings. Its lifecycle commonly moves from PENDING to IN_PROGRESS, then SUCCESS or FAILED; AWS documents a typical duration of about 5–15 minutes. GenAI-produced reasoning must be validated against the actual architecture, service quotas, supported features, data classification and threat model. Only assessed resources influence results; inspect which discovered resources were omitted or unsupported before accepting coverage.
Operational recommendations are proposals
CloudWatch alarms
An alarm recommendation is useful only after checking metric namespace, dimensions, statistic, period, evaluation periods, missing-data treatment and threshold against a known steady state. Test both positive and negative paths: it must detect the condition and avoid persistent false alarms. Verify routing, deduplication, escalation, dashboard, retention and owner response time.
Systems Manager SOPs
An SOP recommendation can map to an SSM Automation runbook. Review every step, input, branch, timeout, retry, rollback and required IAM permission. Separate the operator who starts automation, the automation service role and roles assumed in other accounts. Test success, partial failure, idempotent retry and manual escape. A document that exists but has never run is not recovery evidence.
Fault Injection Service experiments
FIS recommendations can turn a failure hypothesis into a controlled experiment. Define exact targets, selection mode, blast radius, steady-state metrics, stop conditions, duration, rollback and emergency owner before execution. Use tags and account/Region checks to prevent targeting the wrong resources. Start in a safe environment and never treat generated templates as production authorization. AWS236 covers experiment design and safety in depth.
Resiliency recommendations
Configuration recommendations can improve redundancy, backup or recovery, but may increase cost, complexity or correlated failure. Validate supported engine/ Region behavior, KMS and network paths, quotas, deployment order and rollback. Implement through reviewed infrastructure as code where practical, then publish the new model and reassess.
Console walkthrough: read-only evidence
Console labels differ by experience. Confirm account and Region first.
Existing experience
- Open AWS Resilience Hub → Applications and select an owned application.
- Record application ARN/status, assessment schedule, policy and current
published version. Do not select Publish or Start assessment.
- Review AppComponents and resources. Compare them with the workload diagram
and record missing, extra, unsupported and incorrectly grouped resources.
- Open the newest assessment. Record version, time, compliance, estimated
disruption RTO/RPO, score and each recommendation category.
- Inspect implemented/excluded recommendations and drift. Record the evidence
that would be needed to validate each implementation independently.
Next-generation experience
- Open the next-generation Resilience Hub experience and inspect an owned
system, its user journeys and services.
- Record source types/versions, Regions, topology, dependency-discovery status,
policy components and invoker-role ownership.
- Inspect assessed versus discovered resources and data-flow relationships.
- Open the latest failure-mode assessment and record ID, lifecycle, scope,
failure modes, reasoning, findings and recommendations.
- Mark each recommendation accept, change, reject or needs evidence;
do not deploy or run anything in this lesson.
Return to the inventory and prove that no resource or assessment was created.
AWS CLI: read-only inventory
Start with private caller and Region checks:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --output json
aws configure list
aws --version
Never assume ap-south-1 supports every feature shown in another Region. Confirm the current regional table before planning deployment.
Existing resiliencehub API
aws resiliencehub list-apps --output json
aws resiliencehub list-resiliency-policies --output json
export APP_ARN="replace-with-owned-application-arn"
aws resiliencehub describe-app --app-arn "$APP_ARN" --output json
aws resiliencehub list-app-versions --app-arn "$APP_ARN" --output json
aws resiliencehub list-app-assessments --app-arn "$APP_ARN" --output json
aws resiliencehub list-alarm-recommendations \
--assessment-arn "replace-with-assessment-arn" --output json
aws resiliencehub list-sop-recommendations \
--assessment-arn "replace-with-assessment-arn" --output json
aws resiliencehub list-test-recommendations \
--assessment-arn "replace-with-assessment-arn" --output json
Use describe-app-assessment for the selected assessment and inspect resource, component and recommendation APIs exposed by the installed CLI version. Follow pagination; a first page is not a complete inventory. Do not share full ARNs.
Next-generation resiliencehubv2 API
The exact operations available depend on the installed AWS CLI model. Discover them before copying commands from an older lesson:
aws resiliencehubv2 help
aws resiliencehubv2 list-systems --output json
aws resiliencehubv2 list-services --output json
aws resiliencehubv2 list-services --generate-cli-skeleton input
Then use the documented read-only get/list operations for the selected system, service, resources, relationships, policies and assessments. If the namespace is unknown, update the AWS CLI through an approved process; do not silently fall back to classic APIs and claim next-generation coverage. Never run a create, update, publish, start, import or delete operation in this lesson.
IAM, cross-account and data boundaries
Resilience Hub needs permissions to inspect modeled resources and, depending on the experience, discover topology or invoke assessment capabilities. Apply least privilege and separate these actors:
- learner/read-only inventory identity;
- service-linked or service role used by Resilience Hub;
- next-generation invoker/discovery role;
- cross-account roles in member/workload accounts;
- deployment role for approved recommendations;
- SSM Automation and FIS experiment roles.
Trust policy answers who may assume the role; permissions policy answers what the session may do. Restrict account, organization, external ID/source conditions and resources where supported. Monitor CloudTrail for role assumption and mutating calls. AWS Organizations integration can centralize next-generation visibility, but it does not transfer workload ownership or approval authority.
Protect exported assessments and generated templates as architecture-sensitive data. Set an owner, encryption, retention and deletion policy for S3, logs and local evidence. Do not embed secrets in tags, diagrams, SOP parameters or IaC.
Cost model and cleanup
Pricing changes, so calculate from the current regional pricing page before approval. As reviewed on September 15, 2026:
- the existing experience prices assessed applications per month after its
documented introductory allowance;
- next-generation pricing starts per assessed service after its first failure-
mode assessment, includes a limited assessment/resource allowance, and can add charges for larger or additional assessments;
- optional dependency discovery has a per-service monthly charge;
- FIS experiments and the workload resources, telemetry, storage, data transfer,
backups and recovery capacity are billed separately.
The cost owner must count services/applications, Regions/accounts, assessment frequency, assessed resource count, dependency discovery, FIS action duration, CloudWatch metrics/logs/alarms and resources added by recommendations. Budgets and tags do not stop charges automatically.
This practical creates no AWS resources. Cleanup proof is the before/after read-only inventory plus local artifact inventory. In a real implementation also remove obsolete generated templates, temporary test resources and stale evidence according to retention approval; never delete an unknown application model.
Troubleshooting from evidence
| Symptom | Likely causes | Evidence and safe correction |
|---|---|---|
| no applications/services | wrong account/Region, permissions or experience | caller, Region, both namespaces, CloudTrail denial |
| assessment cannot start | unpublished classic app, invalid topology, missing invoker role/policy | version, data-flow edge, role trust and validation message |
| assessment fails | permission, unsupported resource, quota or transient service issue | status reason, CloudTrail, service events, support matrix |
| result misses a dependency | wrong source/version/tag logic, inactive DNS path, direct IP/shared VPC | source diff, 35-day observation window, flow/DNS/application evidence |
| unexpected resources included | broad tags, shared stack/state or stale import | source query, owner and resource-level diff |
| policy breached | target stricter than estimated configuration | component estimate and business objective; improve design or accept risk, never weaken target silently |
| score did not change | change not recognized, exclusion not reassessed, stale app version | implementation status, publish/version and new assessment ID |
| high score but game day fails | score measured recommendation coverage, not runtime outcome | alarm history, SOP execution, FIS timeline, actual RTO/RPO |
| recommendation is unsafe | incomplete context, unsupported design, excessive privilege/blast radius | architecture/security/cost review; change or reject with reason |
| next-gen CLI unknown | old AWS CLI or unavailable regional endpoint | aws --version, approved update, current regional documentation |
Diagnose from the first divergent layer: caller/Region → model experience → input source/version → topology → policy → IAM → assessment status → recommendation → implementation → measured test. Randomly editing resources destroys evidence.
Practical assessment: P05 supplied evidence
Do not use a live workload. Complete every section of the workbook using the P05 architecture and supplied course evidence.
- Choose classic or next-generation modeling and explain why. Also map its terms
to the other experience so an operator cannot confuse APIs.
- Trace one user journey and list hard/soft dependencies, source of truth,
failure boundaries and owner. Include one dependency-discovery blind spot.
- Define justified RTO/RPO and, where applicable, availability/DR/data-recovery
policy components. Do not copy objectives from current architecture.
- Record exact resource source/version, missing/unsupported/shared resources and
grouping or relationship corrections.
- Review at least three findings: accept one, modify one and reject one. For each,
document assumptions, blast radius, IAM, cost, owner and evidence required.
- Design an alarm, SSM SOP and FIS hypothesis for one failure mode. Include
negative test, idempotent retry, stop condition, rollback and approval boundary.
- Define measured RTO/RPO proof and a publish/reassessment sequence. State what a
high score or successful assessment still cannot prove.
- Finish with residual risk, exception expiry and exact no-create evidence.
Knowledge check
- Why is an RTO not the same as instance restart time?
Because it includes detection through business validation and service return.
- Why must a classic application be published before assessment?
The assessment evaluates a specific published application version, not an unversioned draft.
- What do AppComponents represent?
Resources expected to fail and recover together; owners must verify grouping.
- Why can a score of 100 coexist with a failed recovery exercise?
The score measures recognized recommendation coverage, not measured outcomes.
- What is required for a next-generation failure-mode assessment topology?
At least two resources connected by a data-flow relationship, plus valid scope and invoker permissions.
- Why may dependency discovery miss a critical service?
It relies on observed behavior; direct IP, inactive and shared-VPC paths can be blind spots.
- Are multiple tags in one next-generation tag source OR conditions?
No. Tags inside one source are AND; separate tag sources are OR.
- What is wrong with changing the policy target merely to clear a breach?
It hides business risk unless impact analysis and the policy owner approve it.
- When is an SOP implemented?
Not merely when a document exists; permissions, success, failure, retry and outcome must be tested and evidenced.
- Why reassess after changes?
To bind findings to the new model/configuration and reveal remaining drift.
Lesson acceptance
The lesson is complete only when the learner submits:
- a filled workbook naming experience, account/Region scope and owners;
- a journey diagram with resources, relationships, dependencies and blind spots;
- business-owned RTO/RPO/SLO or modular policy rationale;
- exact source/version and included/missing/unsupported-resource evidence;
- assessment/finding register with reasoned accept/change/reject decisions;
- alarm, SOP and FIS designs with roles, tests, stop conditions and rollback;
- measured-recovery and reassessment plan that does not misuse the score;
- current cost calculation inputs and explicit no-create/cleanup proof;
- no secrets, full account IDs, internal endpoints or customer data.
Official sources
- What is AWS Resilience Hub?
- Resiliency policies
- Resilience score
- Run an application assessment with APIs
- Applications and AppComponents
- Operational recommendations
- Standard operating procedures
- Review an assessment
- Next-generation Resilience Hub
- What changes in the next-generation experience
- How next-generation assessments work
- Create and model a next-generation service
- View discovered dependencies
- Minimum topology requirements
- AWS Resilience Hub pricing