Lesson 105 · AWS Learning Path

AWS 105: Multi-AZ self-healing application design

· Published · 9 min read

A load balancer sends requests to healthy servers while bypassing one failed server and adding capacity

The real problem

A team recognizes the name Multi-AZ self-healing application design but has not connected the feature to a real requirement, identity boundary, network or data path, failure mode, price dimension, and cleanup owner. A plausible configuration could still fail the workload.

Final outcome

The learner will produce a requirement-led artifact for Multi-AZ self-healing application design, inspect the matching AWS control plane in the Management Console, run a matching CloudShell or AWS CLI query, interpret the output, diagnose one failure, defend one architecture choice, and prove cleanup or approved retained state.

The practical outcome is not a command transcript. It must show what was expected, what happened, what the result proves, what it does not prove, and which evidence would change the decision.

Learning objectives

By the end of this lesson, the learner can:

  • explain multi-az is end to end;
  • explain self-healing boundary;
  • explain private application tier;
  • explain bootstrap independence;
  • explain state placement;
  • connect control-plane state to the real data, network, identity, or application behavior;
  • identify cost and cleanup ownership before any optional mutation;
  • troubleshoot from evidence without opening broad access or adding broad permissions.

Relationship model

Requirement
   |
   v
Identity and policy -> AWS configuration -> network or data path -> workload behavior
        |                    |                      |                    |
        +--------------------+----------------------+--------------------+
                                      |
                                      v
                         monitoring, cost, recovery, cleanup

Use this model to separate an AWS object that exists from a result that actually works. Every arrow is a verification boundary.

Prerequisites, permissions, Region, and safety

  • Learning baseline: This sequence assumes practical Linux knowledge but no prior cloud-computing or AWS knowledge. Cloud, networking, security, data, automation, and architecture concepts must come from completed earlier lessons. If a prerequisite checkpoint is incomplete, return to its linked lesson before continuing.
  • Confirm a non-root caller with aws sts get-caller-identity and keep the account number private.
  • Use ap-south-1 unless this lesson explicitly names a second Region.
  • Confirm the intended profile and Region with aws configure list before interpreting an empty result.
  • Use read-only List, Get, and Describe permissions for the named services. Design exercises run locally and require no resource-creation permission.
  • This is a no-create lesson. Console and CLI work is read-only, and every design artifact is created locally.
  • Never publish account IDs, public addresses, ARNs containing private account data, session IDs, presigned URLs, object data, credentials, or KMS material.
  • Do not use root, world-open SSH or RDP, disabled TLS verification, unowned resources, or irreversible retention controls in a training exercise.

Core model

ConceptWhat the learner must understand
Multi-AZ is end to endPlacing instances in two AZs is insufficient when the load balancer, state store, network egress, deployment, or monitoring still has a single-AZ dependency.
Self-healing boundaryALB health checks, ASG replacement, immutable launch configuration, and stateless compute can replace failed instances. They do not repair corrupt shared data or a bad deployment image.
Private application tierPublic ALB nodes can route to private target addresses through local VPC routing. Application instances do not need public IPv4 addresses or inbound administration ports.
Bootstrap independenceA private no-NAT lab must not depend on package repositories or external APIs during boot. Production needs resilient egress or private endpoints for required dependencies.
State placementSessions, uploads, secrets, and database state should live in services designed for their durability and availability requirements, not on replaceable instance disks.
Failure testingTest one target failure, one AZ-capacity reduction scenario, bad launch-template version, failed health path, and cleanup. Define expected degradation before injecting anything.

How it works

The P05 lab deliberately separates a low-cost training implementation from production. It uses an HTTP ALB and the Python runtime already present in Amazon Linux 2023 so private instances require no NAT or live package installation. This choice is about network and cost isolation, not a substitute for the learner's Apache or Nginx administration skills. A production design adds owned-domain TLS, an approved web or application runtime and patch path, observability, data services, deployment controls, and tested recovery.

Read the result in layers:

  1. Scope: account, Region, VPC, bucket, AZ, endpoint, principal, object version, or resource ARN.
  2. Control plane: the requested configuration exists and reached an expected state.
  3. Behavior: the request, connection, health check, replication, restore, or application result meets the requirement.
  4. Operations: monitoring, failure owner, cost, retention, rollback, and cleanup are known.

Control-plane success is necessary but not sufficient. A resource can be available while policy, routing, DNS, health, data, or application behavior remains wrong.

Architecture decision table

RequirementPreferred directionWhy
Public HTTP entry with private stateless targetsInternet-facing ALB plus private ASG subnetsOnly the ALB receives public traffic.
No package or management egress in the labSelf-contained user data and no SSHThe sample avoids NAT and interface-endpoint hourly charges.
Production administration and patchingSession Manager endpoints or resilient egress by requirementThe lab shortcut is not a production operating model.
One target fails health checksALB stops routing and ASG replaces itHealth integration restores group capacity from the immutable template.

Professional questions normally contain several valid services. State the requirement that selects one option, why the nearest alternative fails it, and what changed requirement would reverse the choice.

AWS Management Console guided practice

Before opening a service page, write the expected account, Region, starting state, and evidence. Do not choose Create, Save, Purchase, Lock, or Delete unless the lesson explicitly authorizes the live track.

  1. Review the P05 diagram and exact value table: nw-p05-vpc 10.50.0.0/16, two public /24s, two private /24s, ALB, target group, launch template, and ASG.
  2. Open VPC and EC2 dashboards in ap-south-1 and prove that no conflicting Project=NitWings-P05 resources exist before the paid lab.
  3. Open AWS Pricing Calculator or current service pricing and approve the maximum 90-minute ALB, public IPv4, EC2, and EBS cost envelope before AWS 106.

For each step, capture the field name and value in text. A screenshot may support the record but does not replace the explanation. Console labels can evolve, so use the service search and current documentation if a navigation label differs.

CloudShell and AWS CLI practice

CloudShell is the default browser-based command environment taught in AWS 028. AWS 029 and AWS 030 cover local CLI installation and authentication. This lesson therefore does not assume that an unconfigured local shell is ready.

Start every session with:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account portion of the ARN before sharing. Then perform the topic query:

Review the final planned resource graph and prove every dependency has a creation, verification, failure, and deletion step.

aws ec2 describe-subnets --filters Name=tag:Project,Values=NitWings-P05 --query 'Subnets[].{Name:Tags[?Key==`Name`]|[0].Value,CIDR:CidrBlock,AZ:AvailabilityZone,PublicIP:MapPublicIpOnLaunch}' --output table
aws elbv2 describe-load-balancers --names nw-p05-alb --output json
aws autoscaling describe-auto-scaling-groups --auto-scaling-group-names nw-p05-web-asg --output json

Expected interpretation:

Before AWS 106, empty load-balancer and ASG results are expected because this is a design gate. Existing resources with those names require ownership review before proceeding.

Replace every replace-with-... sample value before running its command, and use only an explicitly owned resource. Explain each option first. These queries are read-only; a successful response does not authorize a later create or delete operation.

Practical work

Submit p05-final-build-plan.md with exact CIDRs 10.50.10.0/24, 10.50.20.0/24, 10.50.110.0/24, and 10.50.120.0/24; route tables; IGW; nw-p05-alb-sg; nw-p05-app-sg; nw-p05-web-tg on port 8080; nw-p05-web-lt; and nw-p05-web-asg across private A and B. Include creation order, HTTP verification, one-target failure, cleanup order, price evidence, timer, and owner approval. No resource is created.

The evidence package must contain:

  • the problem and final requirement in the learner's own words;
  • caller type and Region with private identifiers redacted;
  • exact planned values, ownership, and cost class;
  • one Console observation and matching CLI or API evidence;
  • one behavior result or supplied data-plane record;
  • one denied, failed, or counterexample result and evidence-led diagnosis;
  • one architecture choice plus the rejected alternative;
  • cleanup proof or explicit retained-state owner, expiry, and next lesson.

Verification standard

Use expected state before observed state. Record timestamps in UTC and preserve the original failure before changing anything. A passing submission answers all four questions:

  1. What exact requirement was tested?
  2. Which evidence proves the AWS configuration?
  3. Which evidence proves the workload behavior?
  4. What remains unproven or requires later monitoring?

If AWS returns no rows, verify account, Region, permission, filters, pagination, resource type, and deletion state before concluding that nothing exists.

Common failures and troubleshooting

SymptomEvidence firstLikely boundarySmallest safe response
object appears missingcaller, Region, filters, pagination, tagsscope or read permissionalign scope before creating a duplicate
state remains pending or unavailableservice state, events, dependencies, quotasdependency or capacitycorrect the named dependency and wait with a bound
AccessDeniedprincipal, action, resource, explicit-deny contextidentity, resource, endpoint, organization, or KMS policychange only the proven policy layer
configuration exists but behavior failsroute, DNS, security, listener, health, logs, object versiondata path or applicationtest the next boundary and change one control
bill is higher than expectedhours, bytes, requests, AZs, addresses, retentioncost model or retained resourcestop optional work and reconcile the ledger
cleanup is blockeddependency inventory and owning servicedeletion order or immutable stateremove owned dependants in reviewed reverse order

Do not troubleshoot by attaching administrator access, opening administration ports to the internet, disabling encryption, retrying uncontrolled creation, deleting unknown resources, or weakening retention.

Cost, cleanup, and retained state

No AWS resource is created. Close CloudShell and remove or redact downloaded evidence.

Cleanup evidence requires terminal state and an after-inventory. Search related ENIs, public IPv4 addresses, EBS volumes and snapshots, load balancers, target groups, Auto Scaling instances, endpoints, logs, S3 versions and delete markers, backup recovery points, and global IAM roles when they apply. Billing data can lag, so schedule a later review.

Architecture and certification decisions

  • Certification coverage: SAA-C03; SOA-C03; SAP-C02; DOP-C02.
  • Exam mapping: SAA D2-D4.
  • Explain service scope, failure boundary, consistency, recovery, security, operations, and price rather than matching a keyword.
  • Treat availability and durability, encryption and authorization, routing and filtering, health and lifecycle, backup and replication, and discount and capacity as separate concepts.
  • Do not reproduce protected certification questions.

Knowledge check

  1. Do private targets require a NAT gateway for ALB health checks?

Expected direction: No. ALB nodes reach target private addresses through VPC routing.

  1. Why avoid package installation in this no-NAT bootstrap?

Expected direction: Private instances have no external repository path.

  1. Does two-AZ compute protect a single-AZ database?

Expected direction: No. The dependency remains a single-AZ failure.

  1. What makes instance replacement reproducible?

Expected direction: A versioned launch template, deterministic bootstrap or image, and externalized state.

Completion gate and assessment

AreaPointsPassing evidence
Requirement and model15Correct scope, terminology, and final outcome
Console evidence15Current path and interpreted fields
CLI or API evidence15Scoped command, expected result, and limitations
Behavior or decision exercise20Reproducible result or defensible architecture reasoning
Troubleshooting15Original symptom, hypothesis, one change, retest, rollback
Security and cost10Least privilege, data protection, current price dimensions
Cleanup and handoff10Terminal-state proof or approved retained-state record

Pass at 80 out of 100 with no critical safety failure. A missing practical artifact, unexplained output, unsafe access, destructive action outside the owned scope, unplanned billed resource, or false cleanup claim requires remediation and a changed retest.

Official sources

Advertisement