Lesson 235 · AWS Learning Path

AWS 235: Resilience Hub assessment and operational recommendations

· Published · 15 min read

Labelled process diagram for AWS 235: Application resources and policy to Resilience assessment to Recommendation, alarm, SOP, and test to Verified RTO/RPO and reassessment, with decision, proof and rejection evidence.

Why this lesson matters

A backup, Multi-AZ label or architecture diagram does not prove that a customer journey can survive a failure. Resilience starts with a business objective, an accurate dependency model and named owners. It becomes credible only when alarms detect the failure, runbooks recover the service, controlled tests measure the result, and a new assessment confirms that the intended change is represented.

AWS Resilience Hub brings those activities into one service. It can assess a modeled workload, identify likely failure modes and recommend operational controls. Its result is decision support: an estimated RTO/RPO, recommendation, finding or high score is never the same as a successful recovery exercise.

Outcomes

You will be able to:

  • explain resilience, availability, disaster recovery, RTO and RPO from zero;
  • distinguish the existing application experience from next-generation systems,

user journeys and services;

  • choose a model and policy that match the business rather than the resource list;
  • explain resource import, discovery, AppComponents and dependency blind spots;
  • publish/version a model and interpret assessments, findings and drift;
  • evaluate alarm, Systems Manager SOP and Fault Injection Service recommendations;
  • explain what the classic resilience score measures and what it cannot prove;
  • inspect both resiliencehub and resiliencehubv2 safely with the AWS CLI;
  • design IAM, cross-account, data-protection and cost ownership boundaries;
  • diagnose incomplete topology, policy breaches, stale assessments and failures;
  • turn recommendations into controlled changes, tests and reassessment evidence.

Safety boundary and workbook

  • This lesson is read-only. Do not create or publish an application, start an

assessment, deploy recommendations, run an SOP or start a fault experiment.

  • Use only an explicitly owned workload. Resource discovery and dependency data

can reveal sensitive names, topology, tags, accounts and third-party services.

  • Never use a resilience assessment as authorization for a production change.

Each recommendation needs an owner, security/cost review, change approval, rollback plan and independently defined success and stop conditions.

  • Redact account IDs, internal DNS names, resource identifiers, customer data and

generated reasoning from submitted evidence.

Download the Resilience Hub assessment workbook or complete archive.

Foundations: resilience is a business property

Availability asks whether a service is usable now. Resilience is the ability to resist or recover from disruption while meeting an agreed business outcome. Disaster recovery (DR) focuses on restoring service after a major event. They overlap but are not synonyms.

TermPlain meaningExample evidence
RTOmaximum acceptable interruption timecheckout restored in 30 minutes
RPOmaximum acceptable data loss measured in timeat most 5 minutes of orders lost
SLOmeasurable reliability objective over a period99.9% successful checkout/month
failure modea specific way a user journey can failwriter unavailable, queue backlog, expired certificate
controlprevention, detection or recovery mechanismMulti-AZ, alarm, tested runbook
residual riskrisk remaining after controlsRegion loss requires manual DNS approval

RTO is not merely the infrastructure restore duration. Measure from disruption through detection, diagnosis, decision, recovery, dependency readiness, data validation, routing/cutover and business acceptance. RPO is not backup frequency: it is the age of the newest usable and validated state after reconciliation.

Product history and the two experiences

AWS Resilience Hub originally organized work around applications, AppComponents, resiliency policies and assessments. On May 28, 2026, AWS announced the generally available next-generation experience. Existing customers can continue with the original experience and adopt the new one at their own pace. Therefore screenshots, APIs and terminology can legitimately differ.

Existing experienceNext-generation experience
applicationservice (closest conceptual mapping)
application resources grouped into AppComponentssystems contain user journeys and services
static resource checks and estimated disruption RTO/RPOGenAI-powered failure-mode assessments
one resiliency policy across disruption classesmodular availability, DR and data-recovery policies
CLI namespace aws resiliencehubCLI namespace aws resiliencehubv2

Do not mix identifiers or commands between experiences. First record which model the account uses, its Region availability, and which team owns migration. A next-generation system expresses a business boundary; a user journey is a customer or operator flow; a service is a deployable capability and maps most closely to an old application.

End-to-end control loop

business journey + RTO/RPO/SLO + owner
                |
                v
accurate resources, data-flow edges and dependencies
                |
                v
published application version OR current service topology
                |
                v
assessment -> breach/failure-mode findings -> recommendations
                |
                v
reviewed IaC + alarms + SOPs + controlled fault tests
                |
                v
measured recovery/data loss -> reassess -> accept residual risk

Every arrow needs evidence. Resilience Hub does not operate the application for you, approve generated changes, guarantee business continuity or replace game days, backup restores, security reviews and incident management.

Model the workload before assessing it

Start with a user journey such as “place an order.” Trace its entry point, authentication, synchronous calls, queues, databases, object stores, DNS, secrets, certificates, observability and third-party dependencies. Separate:

  • hard dependencies whose failure stops the journey;
  • soft dependencies with degradation, cache or manual fallback;
  • sources of truth from disposable caches and derived indexes;
  • zonal, Regional, cross-Region and external failure boundaries;
  • shared resources whose owner or blast radius crosses services;
  • control-plane dependencies needed only during recovery.

Existing applications and AppComponents

Resources can be imported from supported sources such as CloudFormation stacks, resource groups/tags, Terraform state and EKS. Resilience Hub groups resources that fail and recover together into AppComponents. Automatic grouping is a starting hypothesis. Incorrect components produce misleading estimates, so the application owner must review missing, extra, unsupported and shared resources.

An application draft is not assessable evidence. Publish the application to create a version, then assess that published version. If resources change after assessment, record drift and publish/reassess rather than presenting old results.

Next-generation systems, journeys and services

Next-generation resource discovery can use CloudFormation, Terraform state in S3, tags and EKS. A tag source containing several tags matches them with AND; separate tag sources are combined with OR. Broad tags can silently import unrelated resources, while inconsistent tags omit dependencies.

A failure-mode assessment needs a valid topology, including at least two resources connected by a data-flow relationship, and an invoker role. Service creation can consider up to five Regions. Record exact source versions because a mutable stack, state file or tag query can produce a different model later.

Dependency discovery uses observed DNS/query behavior and a documented 35-day lookback. It can identify internal AWS and third-party dependencies, but it is not omniscient. Direct IP calls, inactive paths and shared-VPC attribution may be missed. A discovered edge is evidence to review, not proof of business criticality.

Define policy from business impact

Existing resiliency policy

The existing experience compares estimated workload recovery against target RTO and RPO for four disruption classes:

  1. application/software;
  2. infrastructure/hardware;
  3. Availability Zone;
  4. Region.

A zero target is accepted, but an estimated RTO or RPO cannot be exactly zero, so that objective will be breached. “Not applicable” needs a business reason; it must not hide an unsupported or unmodeled dependency. Policy selection should come from business-impact analysis, not the architecture's current capability.

Next-generation modular policies

The new model separates availability SLO, disaster recovery and data recovery policy components. Apply the components relevant to each service and journey. This allows different services to carry different objectives without pretending that every resource has the same criticality.

For either experience, record the objective's owner, measurement window, assumptions, scope and exception expiry. A policy is a requirement, not a control.

What an assessment actually does

Existing assessment

An assessment evaluates the imported application version against its attached policy. It estimates RTO/RPO from supported resource configurations and generates recommendations for resilience, CloudWatch alarms, Systems Manager standard operating procedures (SOPs) and AWS Fault Injection Service (FIS) tests.

Compliance/status can include Assessed, Not assessed, Policy breached and Changes detected. Read the component-level reason and recommendation; do not reduce the result to one color. Excluding a recommendation changes score only after reassessment, and exclusions must have an owner and rationale.

Classic resilience score

The maximum classic score weights recommendation coverage approximately as:

CategoryMaximum contribution
policy compliance40 points
implemented alarms20 points
implemented SOPs20 points
implemented tests20 points

A score of 100 means the service recognized the recommended coverage. It does not prove an alarm fired, a runbook succeeded, a fault remained contained, recovery met RTO/RPO, data was correct or operators were available. Preserve raw details, assessment version/time and measured test evidence beside the score.

Next-generation failure-mode assessment

The next-generation assessment analyzes topology and configuration to identify failure modes and findings. Its lifecycle commonly moves from PENDING to IN_PROGRESS, then SUCCESS or FAILED; AWS documents a typical duration of about 5–15 minutes. GenAI-produced reasoning must be validated against the actual architecture, service quotas, supported features, data classification and threat model. Only assessed resources influence results; inspect which discovered resources were omitted or unsupported before accepting coverage.

Operational recommendations are proposals

CloudWatch alarms

An alarm recommendation is useful only after checking metric namespace, dimensions, statistic, period, evaluation periods, missing-data treatment and threshold against a known steady state. Test both positive and negative paths: it must detect the condition and avoid persistent false alarms. Verify routing, deduplication, escalation, dashboard, retention and owner response time.

Systems Manager SOPs

An SOP recommendation can map to an SSM Automation runbook. Review every step, input, branch, timeout, retry, rollback and required IAM permission. Separate the operator who starts automation, the automation service role and roles assumed in other accounts. Test success, partial failure, idempotent retry and manual escape. A document that exists but has never run is not recovery evidence.

Fault Injection Service experiments

FIS recommendations can turn a failure hypothesis into a controlled experiment. Define exact targets, selection mode, blast radius, steady-state metrics, stop conditions, duration, rollback and emergency owner before execution. Use tags and account/Region checks to prevent targeting the wrong resources. Start in a safe environment and never treat generated templates as production authorization. AWS236 covers experiment design and safety in depth.

Resiliency recommendations

Configuration recommendations can improve redundancy, backup or recovery, but may increase cost, complexity or correlated failure. Validate supported engine/ Region behavior, KMS and network paths, quotas, deployment order and rollback. Implement through reviewed infrastructure as code where practical, then publish the new model and reassess.

Console walkthrough: read-only evidence

Console labels differ by experience. Confirm account and Region first.

Existing experience

  1. Open AWS Resilience Hub → Applications and select an owned application.
  2. Record application ARN/status, assessment schedule, policy and current

published version. Do not select Publish or Start assessment.

  1. Review AppComponents and resources. Compare them with the workload diagram

and record missing, extra, unsupported and incorrectly grouped resources.

  1. Open the newest assessment. Record version, time, compliance, estimated

disruption RTO/RPO, score and each recommendation category.

  1. Inspect implemented/excluded recommendations and drift. Record the evidence

that would be needed to validate each implementation independently.

Next-generation experience

  1. Open the next-generation Resilience Hub experience and inspect an owned

system, its user journeys and services.

  1. Record source types/versions, Regions, topology, dependency-discovery status,

policy components and invoker-role ownership.

  1. Inspect assessed versus discovered resources and data-flow relationships.
  2. Open the latest failure-mode assessment and record ID, lifecycle, scope,

failure modes, reasoning, findings and recommendations.

  1. Mark each recommendation accept, change, reject or needs evidence;

do not deploy or run anything in this lesson.

Return to the inventory and prove that no resource or assessment was created.

AWS CLI: read-only inventory

Start with private caller and Region checks:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --output json
aws configure list
aws --version

Never assume ap-south-1 supports every feature shown in another Region. Confirm the current regional table before planning deployment.

Existing resiliencehub API

aws resiliencehub list-apps --output json
aws resiliencehub list-resiliency-policies --output json

export APP_ARN="replace-with-owned-application-arn"
aws resiliencehub describe-app --app-arn "$APP_ARN" --output json
aws resiliencehub list-app-versions --app-arn "$APP_ARN" --output json
aws resiliencehub list-app-assessments --app-arn "$APP_ARN" --output json
aws resiliencehub list-alarm-recommendations \
  --assessment-arn "replace-with-assessment-arn" --output json
aws resiliencehub list-sop-recommendations \
  --assessment-arn "replace-with-assessment-arn" --output json
aws resiliencehub list-test-recommendations \
  --assessment-arn "replace-with-assessment-arn" --output json

Use describe-app-assessment for the selected assessment and inspect resource, component and recommendation APIs exposed by the installed CLI version. Follow pagination; a first page is not a complete inventory. Do not share full ARNs.

Next-generation resiliencehubv2 API

The exact operations available depend on the installed AWS CLI model. Discover them before copying commands from an older lesson:

aws resiliencehubv2 help
aws resiliencehubv2 list-systems --output json
aws resiliencehubv2 list-services --output json
aws resiliencehubv2 list-services --generate-cli-skeleton input

Then use the documented read-only get/list operations for the selected system, service, resources, relationships, policies and assessments. If the namespace is unknown, update the AWS CLI through an approved process; do not silently fall back to classic APIs and claim next-generation coverage. Never run a create, update, publish, start, import or delete operation in this lesson.

IAM, cross-account and data boundaries

Resilience Hub needs permissions to inspect modeled resources and, depending on the experience, discover topology or invoke assessment capabilities. Apply least privilege and separate these actors:

  • learner/read-only inventory identity;
  • service-linked or service role used by Resilience Hub;
  • next-generation invoker/discovery role;
  • cross-account roles in member/workload accounts;
  • deployment role for approved recommendations;
  • SSM Automation and FIS experiment roles.

Trust policy answers who may assume the role; permissions policy answers what the session may do. Restrict account, organization, external ID/source conditions and resources where supported. Monitor CloudTrail for role assumption and mutating calls. AWS Organizations integration can centralize next-generation visibility, but it does not transfer workload ownership or approval authority.

Protect exported assessments and generated templates as architecture-sensitive data. Set an owner, encryption, retention and deletion policy for S3, logs and local evidence. Do not embed secrets in tags, diagrams, SOP parameters or IaC.

Cost model and cleanup

Pricing changes, so calculate from the current regional pricing page before approval. As reviewed on September 15, 2026:

  • the existing experience prices assessed applications per month after its

documented introductory allowance;

  • next-generation pricing starts per assessed service after its first failure-

mode assessment, includes a limited assessment/resource allowance, and can add charges for larger or additional assessments;

  • optional dependency discovery has a per-service monthly charge;
  • FIS experiments and the workload resources, telemetry, storage, data transfer,

backups and recovery capacity are billed separately.

The cost owner must count services/applications, Regions/accounts, assessment frequency, assessed resource count, dependency discovery, FIS action duration, CloudWatch metrics/logs/alarms and resources added by recommendations. Budgets and tags do not stop charges automatically.

This practical creates no AWS resources. Cleanup proof is the before/after read-only inventory plus local artifact inventory. In a real implementation also remove obsolete generated templates, temporary test resources and stale evidence according to retention approval; never delete an unknown application model.

Troubleshooting from evidence

SymptomLikely causesEvidence and safe correction
no applications/serviceswrong account/Region, permissions or experiencecaller, Region, both namespaces, CloudTrail denial
assessment cannot startunpublished classic app, invalid topology, missing invoker role/policyversion, data-flow edge, role trust and validation message
assessment failspermission, unsupported resource, quota or transient service issuestatus reason, CloudTrail, service events, support matrix
result misses a dependencywrong source/version/tag logic, inactive DNS path, direct IP/shared VPCsource diff, 35-day observation window, flow/DNS/application evidence
unexpected resources includedbroad tags, shared stack/state or stale importsource query, owner and resource-level diff
policy breachedtarget stricter than estimated configurationcomponent estimate and business objective; improve design or accept risk, never weaken target silently
score did not changechange not recognized, exclusion not reassessed, stale app versionimplementation status, publish/version and new assessment ID
high score but game day failsscore measured recommendation coverage, not runtime outcomealarm history, SOP execution, FIS timeline, actual RTO/RPO
recommendation is unsafeincomplete context, unsupported design, excessive privilege/blast radiusarchitecture/security/cost review; change or reject with reason
next-gen CLI unknownold AWS CLI or unavailable regional endpointaws --version, approved update, current regional documentation

Diagnose from the first divergent layer: caller/Region → model experience → input source/version → topology → policy → IAM → assessment status → recommendation → implementation → measured test. Randomly editing resources destroys evidence.

Practical assessment: P05 supplied evidence

Do not use a live workload. Complete every section of the workbook using the P05 architecture and supplied course evidence.

  1. Choose classic or next-generation modeling and explain why. Also map its terms

to the other experience so an operator cannot confuse APIs.

  1. Trace one user journey and list hard/soft dependencies, source of truth,

failure boundaries and owner. Include one dependency-discovery blind spot.

  1. Define justified RTO/RPO and, where applicable, availability/DR/data-recovery

policy components. Do not copy objectives from current architecture.

  1. Record exact resource source/version, missing/unsupported/shared resources and

grouping or relationship corrections.

  1. Review at least three findings: accept one, modify one and reject one. For each,

document assumptions, blast radius, IAM, cost, owner and evidence required.

  1. Design an alarm, SSM SOP and FIS hypothesis for one failure mode. Include

negative test, idempotent retry, stop condition, rollback and approval boundary.

  1. Define measured RTO/RPO proof and a publish/reassessment sequence. State what a

high score or successful assessment still cannot prove.

  1. Finish with residual risk, exception expiry and exact no-create evidence.

Knowledge check

  1. Why is an RTO not the same as instance restart time?

Because it includes detection through business validation and service return.

  1. Why must a classic application be published before assessment?

The assessment evaluates a specific published application version, not an unversioned draft.

  1. What do AppComponents represent?

Resources expected to fail and recover together; owners must verify grouping.

  1. Why can a score of 100 coexist with a failed recovery exercise?

The score measures recognized recommendation coverage, not measured outcomes.

  1. What is required for a next-generation failure-mode assessment topology?

At least two resources connected by a data-flow relationship, plus valid scope and invoker permissions.

  1. Why may dependency discovery miss a critical service?

It relies on observed behavior; direct IP, inactive and shared-VPC paths can be blind spots.

  1. Are multiple tags in one next-generation tag source OR conditions?

No. Tags inside one source are AND; separate tag sources are OR.

  1. What is wrong with changing the policy target merely to clear a breach?

It hides business risk unless impact analysis and the policy owner approve it.

  1. When is an SOP implemented?

Not merely when a document exists; permissions, success, failure, retry and outcome must be tested and evidenced.

  1. Why reassess after changes?

To bind findings to the new model/configuration and reveal remaining drift.

Lesson acceptance

The lesson is complete only when the learner submits:

  • a filled workbook naming experience, account/Region scope and owners;
  • a journey diagram with resources, relationships, dependencies and blind spots;
  • business-owned RTO/RPO/SLO or modular policy rationale;
  • exact source/version and included/missing/unsupported-resource evidence;
  • assessment/finding register with reasoned accept/change/reject decisions;
  • alarm, SOP and FIS designs with roles, tests, stop conditions and rollback;
  • measured-recovery and reassessment plan that does not misuse the score;
  • current cost calculation inputs and explicit no-create/cleanup proof;
  • no secrets, full account IDs, internal endpoints or customer data.

Official sources

Advertisement