Lesson 061 · AWS Learning Path

AWS 061: EC2 concepts and instance lifecycle

· Published · 15 min read

An EC2 instance moves through pending, running, stopped, hibernated, and terminated states with storage and address changes shown

The problem

A team uses EC2 concepts and instance lifecycle as a label or Console setting without proving the identity, scope, behavior, failure boundary, cost, or operational result. The configuration appears complete but the real requirement remains untested.

Final outcome

The learner will inspect, explain, test, and document EC2 concepts and instance lifecycle. The submission must connect the Console and CLI view to the same AWS control, state what the evidence proves, diagnose one likely failure, and make a requirement-based architecture decision.

Learning outcomes

You will be able to:

  1. explain what Amazon EC2 is and which connected services form a working application;
  2. distinguish an instance, AMI, instance type, network interface, address, security group, volume, snapshot, load balancer, and certificate;
  3. explain why lifecycle state is not health;
  4. explain why stop and start can move the host;
  5. explain why reboot normally keeps the instance in place;
  6. explain why termination is final for the instance;
  1. distinguish successful configuration from successful workload behavior;
  2. preserve redacted evidence and complete the stated cleanup or retention decision.

Mental model

requirement
   -> identity and permission
   -> account, Region, and resource scope
   -> configuration or request
   -> observable state
   -> workload result
   -> failure evidence
   -> cleanup or controlled retention

Never begin with a create button or a copied command. Start with the result that must be proven and the boundary that must remain protected.

Start with the complete EC2 application picture

Amazon Elastic Compute Cloud, shortened to Amazon EC2, provides virtual computers called instances. “Elastic” means capacity can be changed or replaced as requirements change. “Compute” means processor and memory used to run an operating system and applications. “Cloud” means the capability is requested from AWS through service interfaces instead of installing the physical server yourself.

An EC2 instance is only the compute part of an application. A beginner must see the surrounding system before studying each component separately:

User's browser
      |
      | HTTPS request for app.example.com
      v
DNS record
      |
      v
Load balancer + TLS certificate
      |
      | listener rule selects a target group
      v
Load-balancer security group
      |
      v
EC2 network interface + EC2 security group
      |
      v
EC2 instance: operating system + web application
      |
      +---- EBS volume: durable block storage in one AZ
      |
      +---- Instance store: temporary disks on the physical host, when supported

EBS volume --backup operation--> EBS snapshot

Every box answers a different question. Treating them as interchangeable creates both exam mistakes and production failures.

ComponentBeginner meaningRelationship to EC2Failure question
AMIA launch image containing an operating-system starting point and block-device mappingsEC2 uses it when creating an instanceDid the instance start from the intended image and architecture?
Instance typeA named hardware allocation such as a combination of vCPU, memory, network, and storage capabilitiesDetermines the instance's capacity and supported featuresIs the workload exceeding CPU, memory, network, or EBS limits?
Network interface, ENIA virtual network adapter with private addresses and security-group attachmentsConnects the instance to a subnet in a VPCDoes the interface have the expected subnet, address, route path, and security groups?
Public IPv4 addressA publicly routable address that AWS can assign automaticallyCan change after stop/startDid the address change, and should clients use DNS instead?
Elastic IP addressA static public IPv4 allocation owned by the account until releasedCan be associated with an ENI or supported resourceIs it associated correctly, and is an idle public IPv4 allocation being billed?
Security groupA stateful virtual firewall attached to network interfacesControls allowed inbound and outbound flowsDoes the required source, protocol, and port have an allow rule?
EBS volumeDurable block storage presented like a diskCommonly stores the root filesystem and application dataIs it in the same AZ, attached, formatted, mounted, full, encrypted, and performing within limits?
Instance storeTemporary block storage physically attached to the host for supported instance typesUseful for replaceable cache or scratch dataWas required data incorrectly stored on media that disappears after stop or host loss?
EBS snapshotA point-in-time block backup stored and managed independently of one EBS volumeCreates replacement volumes for recoveryWas the snapshot completed, retained, copied where needed, and actually restored in a test?
Load balancerA separate managed service that receives connections and distributes them to healthy targetsPlaces stable service endpoints and health-based routing in front of instancesAre the listener, rule, target registration, health check, routes, and security groups consistent?
ACM certificateA certificate managed by AWS Certificate Manager for supported integrated servicesCommonly enables TLS on a load-balancer listener; it is not installed automatically merely because EC2 existsDoes the certificate cover the hostname, validate successfully, exist in the required Region, and remain associated?
Auto Scaling groupA desired-capacity controller that launches and replaces instances from a templateMaintains a fleet instead of depending on one serverCan it launch, pass health checks, scale, and replace failed capacity?

Terms that are related but not contained inside EC2

  • EBS is a separate storage service. EC2 attaches an EBS volume as a block device.
  • Elastic Load Balancing is a separate service. A load balancer sends traffic to registered targets, which can include EC2 instances.
  • ACM is a separate certificate service. An Application Load Balancer can use an ACM certificate for an HTTPS listener.
  • An Elastic IP is an account allocation associated with a network interface or supported service; it is not permanent merely because an instance once used it.
  • A security group attaches to a network interface. Describing it only as an “EC2 firewall” hides why load balancers, databases, and interface endpoints can also use security groups.

How this EC2 sequence builds capability

The pages are not independent articles. They form one cumulative build:

Existing lessonsCapability added to the same mental model
AWS 053 to AWS 060Plan and build the VPC, subnets, routes, internet gateway, security group, and network ACL before launching compute
AWS 061 to AWS 065Understand the instance lifecycle, select an AMI and instance type, compare processor architectures, and launch deliberately
AWS 066 to AWS 068Reach the instance safely, bootstrap it, and use instance metadata through IMDSv2
AWS 069 to AWS 073Distinguish EBS from instance store; select a volume type; attach, format, mount, snapshot, restore, and verify persistence
AWS 074 to AWS 080Understand interfaces and public addressing; monitor, break, diagnose, recover, and clean up the first EC2 project
AWS 091 to AWS 098Choose placement, tenancy, purchasing, fleet, launch-template, image-lifecycle, and monitoring controls
AWS 099 to AWS 102Put the application behind the correct load-balancer type; configure listeners, rules, target groups, health checks, TLS, and ACM
AWS 103 to AWS 107Build an Auto Scaling fleet across Availability Zones, test unhealthy-target replacement, and diagnose scaling failures

Later lessons may deepen a component, but this lesson must establish enough meaning that the learner knows where that component fits before seeing its Console page.

EC2 automation is larger than Auto Scaling

Automation occurs at several layers, and the course must distinguish them:

Automation layerPurposeExisting lesson location
User data and cloud-initPerform first-boot configurationAWS 067 and AWS 072
AMI lifecycleCapture a controlled machine-image baseline and replace outdated instancesAWS 062 and AWS 097
Launch templatesVersion the EC2 launch specification used by fleets and Auto Scaling groupsAWS 096 and AWS 103
EC2 Fleet and mixed capacityRequest capacity across instance types and purchase optionsAWS 093 to AWS 095
EC2 Auto ScalingMaintain desired capacity, replace unhealthy instances, and adjust capacity to demandAWS 103 to AWS 107
Systems ManagerInventory, patch, configure, run commands, and automate operational proceduresAWS 210 and AWS 223 to AWS 228
Infrastructure as codeCreate repeatable EC2 and network stacks from reviewed templatesAWS 219 to AWS 222 and AWS 372 to AWS 384
Event-driven remediationReact to state changes, alarms, and findings through EventBridge, Lambda, or Automation runbooksAWS 206, AWS 230, AWS 392 to AWS 404
Deployment pipelineBuild a new artifact or image and roll it through tested environmentsAWS 355 to AWS 371

One automation method does not replace all the others. For example, an Auto Scaling group can replace a failed instance, but it cannot repair a broken launch template. User data can configure a new instance, but it does not maintain the desired number of instances. Systems Manager can patch running instances, while an immutable-image design may instead build a new AMI and perform an instance refresh.

Auto Scaling coverage contract for AWS 103 to AWS 107

The later Auto Scaling pages must teach and test all of these connected concepts:

  • the difference between vertical scaling (changing resource size) and horizontal scaling (changing resource count);
  • an Auto Scaling group, its launch template, selected subnets, Availability Zones, attached target groups, and service-linked role;
  • minimum, desired, maximum, and current capacity, including why desired capacity must remain within the configured limits;
  • manual scaling, scheduled actions, predictive scaling, target tracking, step scaling, and the legacy simple-scaling model;
  • CloudWatch metrics and alarms, including choosing a metric that changes in a useful direction when capacity changes;
  • default instance warmup, cooldown, health-check grace period, estimated instance warmup, and why these timers solve different problems;
  • EC2, Elastic Load Balancing, VPC Lattice, EBS, and custom health evidence where supported, without treating a process that merely started as a healthy application;
  • replacement of unhealthy instances and the difference between maintaining desired capacity and reacting to increased demand;
  • multi-AZ balancing, impaired-AZ behavior, Availability Zone distribution, and capacity shortages;
  • scale-in protection, standby, detach and attach operations, and how manual changes affect desired capacity;
  • default and custom termination policies, including why AZ balance is considered before ordinary termination preference;
  • lifecycle hooks for launch and termination, heartbeat timeouts, completion results, and abandoned or timed-out hooks;
  • load-balancer registration, connection draining or deregistration delay, graceful application shutdown, and state externalization;
  • warm pools, reuse policy, prepared instance state, and the cost-versus-startup-time trade-off;
  • mixed instances, instance weights, On-Demand and Spot allocation, Spot interruption handling, and Capacity Rebalancing;
  • instance maintenance policies and the availability-versus-cost effect of launch-before-terminate and terminate-and-launch behavior;
  • instance refresh, launch-template versions, minimum and maximum healthy percentages, checkpoints, checkpoint delay, bake time, skip matching, alarms, cancellation, and rollback;
  • maximum instance lifetime and controlled replacement of aging or outdated instances;
  • scaling activities, CloudTrail, CloudWatch, EventBridge events, Auto Scaling metrics, target health, application logs, and cost evidence;
  • IAM permissions, tags that propagate at launch, encrypted volumes, IMDSv2, secrets delivery, network controls, and least-privilege instance roles;
  • quotas, insufficient capacity, invalid launch templates, unavailable instance types, failed health checks, blocked lifecycle hooks, alarm mistakes, scaling oscillation, and runaway cost;
  • dependency-aware deletion of scaling policies, scheduled actions, lifecycle hooks, load-balancer relationships, the Auto Scaling group, launch-template versions, instances, ENIs, volumes, snapshots, alarms, logs, and retained images.

The guided lab must prove more than group creation. It must establish a healthy baseline, create load, observe scale-out, remove load, observe scale-in, deliberately make one instance unhealthy, observe replacement, deploy a changed template through an instance refresh, test rollback, and reconcile the final resource inventory and billable state.

Follow one request and one write

For a web request, the browser resolves a DNS name, negotiates TLS with the load balancer, reaches a listener, matches a rule, enters a target group, passes network controls, and finally reaches the application process on an instance. A running instance proves only that EC2 reached a lifecycle state. It does not prove that DNS, TLS, the load balancer, the target health check, the operating system, or the application works.

For a disk write, the application asks the operating system to write through a mounted filesystem to a block device. If that device is EBS, the volume has its own attachment, AZ, encryption, performance, snapshot, retention, and billing state. If it is instance store, the application must tolerate losing that data when the supporting host lifecycle ends.

These two paths explain why troubleshooting starts with the observed symptom and tests one boundary at a time.

Core facts

ConceptWhat it means in practice
EC2 is resizable virtual computeAn instance combines an AMI, instance type, network interfaces, storage, IAM role, placement, purchasing option, and launch configuration in one Availability Zone.
Lifecycle state is not healthPending, running, stopping, stopped, shutting-down, and terminated describe lifecycle. System, instance, EBS, load-balancer, and application checks answer different health questions.
Stop and start can move the hostAn EBS-backed instance normally keeps its instance ID, private IPv4, EBS volumes, and ENIs, but moves to new hardware and loses an automatically assigned public IPv4 unless an Elastic IP is used.
Reboot keeps the instance in placeA reboot is an operating-system restart on the same host with the same addressing and attached volumes. It does not perform the host replacement of a stop and start.
Termination is final for the instanceTermination deletes the instance and any EBS volume marked DeleteOnTermination. Retained EBS volumes and snapshots continue to exist and can continue to cost money.
Hibernate saves RAM to EBSSupported encrypted EBS-backed instances can write memory to the root volume and later resume, but hibernation has prerequisites and does not preserve instance-store data.

How to reason about it

Design applications so an instance can be replaced rather than repaired forever. Store durable state outside the host, bootstrap from a versioned image and configuration, monitor application health, and use Auto Scaling or orchestration to replace unhealthy capacity across Availability Zones.

Use the following decision table as a starting point, then change the answer when the scenario changes.

RequirementPreferred directionWhy
Operating system is stuck but host is healthyReboot, then investigate guest logsA guest restart is the least disruptive recovery attempt.
Underlying host problem affects EBS-backed instanceAutomatic recovery or stop and startRecovery moves or restores the instance on healthy infrastructure.
Stateless service instance fails health checkReplace through Auto ScalingReplacement avoids dependence on one long-lived server.
Need rapid resume with in-memory process stateEvaluate hibernationRAM is persisted to encrypted root EBS under supported conditions.

Prerequisites, permissions, Region, and cost

Run this lesson as the approved non-root course identity. Begin with aws sts get-caller-identity, confirm the private account record, and set the fixed project Region before any regional query. IAM resources are global within an account, while STS endpoints and the services reached by an identity can be regional.

Use only the read or change actions required for this lesson. An AccessDenied result is evidence to analyse, not permission to switch to root or attach AdministratorAccess. Record the action, resource, principal type, request context, and smallest justified correction.

The listed cost tier is T0 - no resource creation. Free Tier eligibility and credits are account-specific. Before a mutating lab, identify every resource that can charge, estimate its duration, start a timer, and write the reverse cleanup order. A budget reports cost after billing data arrives and is not a real-time stop control.

AWS Management Console method

  1. Open EC2 Instances and select an instance.
  2. Inspect lifecycle state, status checks, Availability Zone, addressing, root device, termination protection, and scheduled events.
  3. Use Instance state actions to compare reboot, stop, hibernate, and terminate, but perform destructive actions only on a disposable learning instance.

For every step record the service page, selected account and Region, exact object, visible state, and why that state matters. A Console label or green status is not enough unless it is tied to the final workload outcome.

AWS CLI or API evidence

Read lifecycle, launch time, addressing, root type, and protection flags together.

aws ec2 describe-instances --instance-ids i-0123456789abcdef0 --query 'Reservations[0].Instances[0].{State:State.Name,AZ:Placement.AvailabilityZone,Root:RootDeviceType,Private:PrivateIpAddress,Public:PublicIpAddress,Launch:LaunchTime}'
aws ec2 describe-instance-attribute --instance-id i-0123456789abcdef0 --attribute disableApiTermination

Termination protection blocks selected API termination, not shutdown from the guest, Auto Scaling replacement, or every failure path. Back up important data independently.

Before running the command, replace every placeholder, explain each option, and decide whether the operation is read-only or mutating. Capture the exit code immediately. Redact account IDs, ARNs, public addresses, request identifiers, and personal data before sharing.

The CLI and Console are clients of AWS APIs. Matching state across them increases confidence, but neither substitutes for data-plane or application verification.

Practical work

Create p04-instance-lifecycle.md for the future nw-p04-web-1. Trace pending, running, stopping, stopped, shutting-down, and terminated states. For each transition record compute billing direction, EBS retention, public IPv4 behavior, host-memory loss, instance ID behavior, and the operator evidence. No instance is launched.

The evidence package must include:

  • non-root principal type, account verified privately, and Region;
  • exact intended and observed state;
  • one Console observation and the matching CLI or API result;
  • one successful result and one denied, failed, or counterexample result;
  • what each result does not prove;
  • cost state and elapsed lab time;
  • cleanup evidence or a named retained-state owner and next lesson.

Verification standard

Use three levels of proof:

  1. Control plane: the object or policy exists with the intended configuration.
  2. Data plane or behavior: the request, packet, session, storage path, or application does what the requirement states.
  3. Operations: monitoring, failure diagnosis, cost, ownership, and cleanup are known.

A control-plane response can precede final readiness. Use waiters or state polling where supported, then test the actual behavior. If a request times out, do not assume it failed. Inspect state and use documented idempotency before retrying a mutation.

Troubleshooting method

SymptomEvidence firstSmallest safe response
command cannot authenticatecredential source, expiry, caller preflightrestore approved temporary login
access is deniedaction, resource, principal, all policy layerscorrect only the missing or conflicting control
object appears missingaccount, Region, filters, pagination, permissionalign scope before creating anything
configured state exists but behavior failsroute, identity, dependency, logs, service statetest the next boundary in the path
cleanup is blockeddependency inventory and owning serviceremove dependants in reviewed reverse order

Keep the original symptom and timestamp. State one hypothesis, make one reversible change, repeat the original test, and record rollback. Never open a management port to the world, expose credentials, disable TLS verification, format an unknown disk, or add broad permissions as a generic fix.

Architecture and certification decisions

Professional-level questions provide competing valid features. Identify the requirement that decides between them: human or workload identity, same-account or cross-account access, regional or zonal scope, stateful or stateless filtering, durable or ephemeral data, latency, RTO/RPO, cost, or operational ownership.

Explain why the selected option fits and why each plausible alternative fails one stated requirement. Do not rely on feature memorization or reproduce protected certification questions.

Knowledge check

  1. Does stopping an EBS-backed instance delete its EBS volumes?

Expected direction: Normally no, but storage charges continue and instance-store data is lost.

  1. Which operation normally changes the underlying host: reboot or stop/start?

Expected direction: Stop and start.

  1. Does a running state prove the web application is healthy?

Expected direction: No. Use status checks, target health, and application monitoring.

  1. What determines whether a volume is deleted at termination?

Expected direction: The block-device mapping DeleteOnTermination setting for that attachment.

Cost, cleanup, and retained state

Retain only the state explicitly required by the next P04 lesson and record it in the private resource ledger.

Cleanup evidence includes the final state query, not only a successful delete response. Search related resources, other Regions used by the lab, retained storage, public IPv4 addresses, logging destinations, and service-managed dependencies. Schedule a later billing review because cost data can lag.

Completion gate

Pass when the practical artifact explains the real problem, matches Console and CLI evidence, answers every knowledge check, diagnoses one failure without broadening access, records cost, and proves cleanup or approved retention. The learner must defend one decision orally and name the requirement that would change it.

Official sources

Advertisement