AWS 187: Pilot light strategy
Why this lesson matters
Keep critical data and core services ready in a recovery Region while most application capacity remains off or minimal until an event.
What you will be able to do
By the end, you can:
- explain pilot light strategy in plain language;
- locate the current service controls in the AWS Management Console;
- run the matching CloudShell or AWS CLI queries and explain every important field;
- draw the identity, network, data, failure, and monitoring path;
- choose the service from requirements and reject it when those requirements are absent;
- diagnose a failed or misleading result from evidence;
- state the cost owner and prove cleanup or a no-create result.
Before you start
- Use a personal AWS account only when its owner has approved the lesson. Do not use the root user for daily work.
- CloudShell is the default command environment. AWS028 explains CloudShell; AWS029 and AWS030 explain local AWS CLI installation and profiles.
- The course example Region is
ap-south-1. Global services and services with a required control Region are called out in their commands. - Run
aws sts get-caller-identityprivately. Redact the account number before sharing evidence. - Never paste access keys, passwords, secret values, private object data, presigned URLs, or full account-specific ARNs into a submission.
- This is a no-create lesson. Every Console action and AWS CLI command is read-only. Create the practical artifact locally.
- Console wording can change. Use the Console service search if a menu label has moved, then confirm the current field in the official documentation.
The core model
| Question | What it means in this lesson |
|---|---|
| Purpose | Keep critical data and core services ready in a recovery Region while most application capacity remains off or minimal until an event. |
| Scope and boundary | The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Pilot light disaster recovery. |
| Evidence of success | Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Pilot light disaster recovery. |
| Cost model | Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design. |
| Safe rejection rule | Avoid leaving critical dependencies, quotas, keys, images, or IaC unavailable in the recovery Region. |
How the request flows
+----------------------------+
| Primary workload changes |
+----------------------------+
|
v
+----------------------------------------------+
| Replicated data and core recovery services |
+----------------------------------------------+
|
v
+------------------------------------------+
| Event triggers infrastructure scale-up |
+------------------------------------------+
|
v
+-----------------------------------------+
| Validation, DNS cutover, and failback |
+-----------------------------------------+
What remains lit in a pilot light
Pilot light is active/passive DR. The recovery Region continuously holds current or near-current core data and enough foundation to launch the application, but most serving capacity is not running. “Stopped servers everywhere” is not the model; prefer reproducible IaC, artifacts and launch configuration.
| Always prepared | Created or scaled during recovery |
|---|---|
| recovery account, VPC/routes/endpoints, roles, KMS keys, secrets path, DNS/certificates | production application compute and concurrency |
| replicated data or continuous block replication, plus point-in-time backups | database promotion/restore and production-sized capacity |
| copied AMIs/images/packages and versioned IaC/configuration | load balancers/targets if absent, autoscaling desired capacity and workers |
| quotas, capacity plan, health checks, runbook and access | traffic shift, partner allow-list changes and final validation |
Pilot light costs more than backup/restore and can recover faster because data and foundations exist. It is slower and riskier than warm standby because provisioning, promotion and scale occur during the incident.
Activation sequence
- Declare disaster and freeze conflicting automation.
- Determine the last consistent checkpoint and fence primary writers to avoid split brain.
- Promote or restore recovery data and validate consistency.
- Deploy the approved application release with regional configuration and secrets.
- Scale database, compute, queues and quotas to tested load.
- Run technical and business transactions with outbound side effects initially disabled.
- Shift traffic with Route 53 or Global Accelerator and watch errors, latency, saturation and correctness.
- Establish backup protection in the active recovery Region and communicate achieved RPO/RTO.
Automatic health failover must not promote an incomplete environment. Low DNS TTL does not guarantee immediate client change, and existing connections can persist.
Elastic Disaster Recovery and data corruption
AWS Elastic Disaster Recovery continuously replicates supported source servers into a low-cost staging area and launches recovery instances for drill/failover. It is a server-recovery implementation, not a complete application plan: dependencies, database consistency, DNS, identity, quotas, testing, failback and licensing remain yours.
Replication can lose recent asynchronous writes and can copy deletion, corruption or ransomware. Keep isolated versioned point-in-time backups, monitor lag and define when to use replicated state versus an older clean point. Measure achieved RPO from timestamps, not product targets.
Architecture decision table
| Situation | Direction | Reason |
|---|---|---|
| Requirement matches | Use pilot light when recovery speed must beat backup-and-restore but full standby cost is not justified. | Select only after scope, behavior, security, recovery, operations, and price evidence agree. |
| Requirement does not match | Avoid leaving critical dependencies, quotas, keys, images, or IaC unavailable in the recovery Region. | Rejecting an attractive service is a valid architecture result. |
| No create permission or cost approval | Use supplied evidence and local design work | Learning does not depend on creating an hourly resource. |
| Existing resource is unknown or unowned | Inspect only, then stop | Never change or delete a resource merely because it resembles a course example. |
AWS Management Console, step by step
Sign in with the normal non-root learning identity. Write the expected starting state before opening the service.
- Use the Console service search and open RDS or DynamoDB replication, AMIs, ECR, S3, CloudFormation, Route 53, and Backup read-only inventories; confirm the account and Region before reading the page.
- Inspect the supplied or owned resource's status, configuration, permissions, networking, encryption, monitoring, tags, and dependencies without changing it.
- Open the related metrics, logs, events, or history view and record one timestamped signal that would prove or disprove the expected behavior.
- Return to the resource list, clear filters, and record the final inventory. On the read-only track, do not choose Create, Save, or Delete.
CloudShell and AWS CLI, step by step
Start with a known caller and Region:
export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list
Redact the account part of the ARN in shared evidence. Now run the topic queries:
aws rds describe-db-instances --query 'DBInstances[].{Id:DBInstanceIdentifier,Replicas:ReadReplicaDBInstanceIdentifiers}' --output json
aws s3api get-bucket-replication --bucket replace-with-owned-bucket
aws ecr describe-replication-configuration --output json
aws cloudformation list-stacks --stack-status-filter CREATE_COMPLETE UPDATE_COMPLETE --output table
Expected interpretation
Pilot-light evidence must show current data, deployable infrastructure, dependencies, quotas, credentials, DNS cutover, scale-up time, tests, and failback. Templates alone are not recovery.
Practical work
Design a pilot light for P05 in a second Region. Mark always-on data, replicated images and configuration, secrets strategy, infrastructure deployment, capacity ramp, DNS change, RPO/RTO, quarterly test, and failback.
Diagnose this topic from its own evidence
- Data cannot decrypt: inspect destination key, policy/grants and service role.
- Launch misses RTO: separate quota, image, instance-type, IaC drift, dependency and approval delays.
- Traffic shifts but clients fail: inspect endpoint readiness, TLS, health checks, DNS caching and origin allow lists.
- Replication is healthy but data is wrong: compare checkpoint/lag, consistency and corruption timeline.
- Failback risks split brain: establish single-writer fencing and reverse synchronization first.
Cost and cleanup
Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Knowledge check
- What operational purpose is this lesson solving?
Expected direction: Keep critical data and core services ready in a recovery Region while most application capacity remains off or minimal until an event.
- Which scope or ownership boundary must be proved first?
Expected direction: The learner must identify the account and Region scope, resource boundary, identity path, data or network path, failure behavior, observability, and cleanup ownership for Pilot light disaster recovery.
- What evidence is strong enough to accept the result?
Expected direction: Success means the Console fields, CLI result, workload behavior, monitoring evidence, and architecture claim agree. An available state alone is not enough for Pilot light disaster recovery.
- Which tempting design or shortcut must be rejected?
Expected direction: Avoid leaving critical dependencies, quotas, keys, images, or IaC unavailable in the recovery Region.
- Which cost dimensions and retained resources need an owner?
Expected direction: Requests, running capacity, storage, logs, data transfer, retained state, and optional features must be priced for the exact design.
Lesson acceptance
- Identify exactly what runs continuously and what launches during recovery.
- Produce a dependency-ordered activation/traffic runbook with rollback gates.
- Explain Elastic Disaster Recovery's scope and remaining application duties.
- Combine replication with isolated point-in-time backup.
- Demonstrate measured RPO/RTO, quota readiness and safe failback.