Lesson 116 · AWS Learning Path

AWS 116: Amazon EFS

· Published · 12 min read

An EC2 instance connects to persistent EBS volumes on one side and fast temporary host-local instance-store disks on the other

The real problem

A team recognizes the name Amazon EFS but has not connected the feature to a real requirement, identity boundary, network or data path, failure mode, price dimension, and cleanup owner. A plausible configuration could still fail the workload.

Final outcome

The learner will produce a requirement-led artifact for Amazon EFS, inspect the matching AWS control plane in the Management Console, run a matching CloudShell or AWS CLI query, interpret the output, diagnose one failure, defend one architecture choice, and prove cleanup or approved retained state.

The practical outcome is not a command transcript. It must show what was expected, what happened, what the result proves, what it does not prove, and which evidence would change the decision.

Learning objectives

By the end of this lesson, the learner can:

  • explain managed nfs;
  • explain mount targets;
  • explain regional and one zone;
  • explain performance and throughput;
  • explain access points;
  • connect control-plane state to the real data, network, identity, or application behavior;
  • identify cost and cleanup ownership before any optional mutation;
  • troubleshoot from evidence without opening broad access or adding broad permissions.

Relationship model

Requirement
   |
   v
Identity and policy -> AWS configuration -> network or data path -> workload behavior
        |                    |                      |                    |
        +--------------------+----------------------+--------------------+
                                      |
                                      v
                         monitoring, cost, recovery, cleanup

Use this model to separate an AWS object that exists from a result that actually works. Every arrow is a verification boundary.

Prerequisites, permissions, Region, and safety

  • Learning baseline: This sequence assumes practical Linux knowledge but no prior cloud-computing or AWS knowledge. Cloud, networking, security, data, automation, and architecture concepts must come from completed earlier lessons. If a prerequisite checkpoint is incomplete, return to its linked lesson before continuing.
  • Confirm a non-root caller with aws sts get-caller-identity and keep the account number private.
  • Use ap-south-1 unless this lesson explicitly names a second Region.
  • Confirm the intended profile and Region with aws configure list before interpreting an empty result.
  • Use read-only List, Get, and Describe permissions for the named services. Design exercises run locally and require no resource-creation permission.
  • This is a no-create lesson. Console and CLI work is read-only, and every design artifact is created locally.
  • Never publish account IDs, public addresses, ARNs containing private account data, session IDs, presigned URLs, object data, credentials, or KMS material.
  • Do not use root, world-open SSH or RDP, disabled TLS verification, unowned resources, or irreversible retention controls in a training exercise.

Core model

ConceptWhat the learner must understand
Managed NFSAmazon EFS provides elastic shared file storage using NFSv4 for supported clients. Multiple instances can mount the same filesystem and use normal hierarchical file semantics.
Mount targetsClients reach EFS through mount-target ENIs. A resilient Regional design creates a mount target in each client Availability Zone and permits NFS TCP 2049 from client security groups.
Regional and One ZoneRegional EFS storage classes store data across multiple AZs. EFS One Zone stores within one AZ and trades resilience for cost.
Performance and throughputPerformance mode and throughput mode affect latency and throughput behavior. Elastic throughput scales with activity; provisioned and bursting choices depend on workload and Region support.
Access pointsEFS access points provide an application-specific entry path and POSIX identity enforcement. IAM authorization and TLS can be used with the EFS mount helper.
Lifecycle and backupLifecycle policies can move files among EFS storage classes according to access. AWS Backup and tested restore address recovery beyond filesystem availability.

How it works

EFS availability includes DNS, mount target per AZ, routing, NACL, security group, NFS client, mount options, POSIX UID/GID, access-point root, throughput, connection quota, backup, and application locking behavior.

EFS exists because many Linux applications expect a shared hierarchical filesystem rather than an object API or one-host block device. It is a Regional service unless One Zone is deliberately chosen. A Regional file system stores data redundantly across multiple AZs, while the mount target is the zonal network entry point: an ENI with an IP address in a subnet. Create one mount target in every AZ containing clients so DNS returns a local path and an AZ/network failure does not force cross-AZ dependency. A mount target is not a replica containing a separate copy of files.

The data path is process -> Linux VFS/NFS client -> EFS mount helper/TLS tunnel -> DNS -> mount-target ENI -> distributed EFS storage. The control plane creates file systems, policies, access points and mount targets; NFS is the data plane. IAM may authorize mount operations when the mount helper uses iam, but normal POSIX ownership/mode checks still apply. Root squashing and the file-system resource policy can deny an apparently valid OS user. An access point can force a root directory and POSIX UID/GID, giving each application a consistent view without creating a separate filesystem.

Files, metadata and application semantics

EFS supports NFSv4.x file operations, directories, permissions and file locking, but a network filesystem is not a local disk. Thousands of tiny metadata-heavy files, serial operations and chatty locking can bottleneck even when byte throughput is low. Applications must handle transient NFS errors and must not assume that adding EC2 instances makes a single-file serial workload faster. Test representative file counts, operation mix, client count, mount options, latency and failover - not only a sequential dd result.

General Purpose is the recommended performance mode for current new designs; Max I/O is a previous-generation mode with higher per-operation latency and is incompatible with some current options. Throughput mode is separate: Elastic follows activity, Provisioned reserves throughput independent of stored size, and Bursting earns/uses credits based on storage. Read current Regional limits before using a numeric maximum because quotas and client versions affect them. Monitor PercentIOLimit, MeteredIOBytes, throughput utilization, client connections and burst credits where relevant.

Lifecycle management can transition files based on access among Standard, Infrequent Access and Archive classes where supported, and can return accessed files according to policy. Access charges and minimum-storage-duration effects can make aggressive transitions expensive. EFS Replication creates a read-only destination file system for disaster-recovery workflows; it is asynchronous and is not a backup against every logical deletion. AWS Backup recovery points plus an application-consistent restore test address recovery. Regional storage, replication and backup solve different failure classes.

Read the result in layers:

  1. Scope: account, Region, VPC, bucket, AZ, endpoint, principal, object version, or resource ARN.
  2. Control plane: the requested configuration exists and reached an expected state.
  3. Behavior: the request, connection, health check, replication, restore, or application result meets the requirement.
  4. Operations: monitoring, failure owner, cost, retention, rollback, and cleanup are known.

Control-plane success is necessary but not sufficient. A resource can be available while policy, routing, DNS, health, data, or application behavior remains wrong.

Architecture decision table

RequirementPreferred directionWhy
Linux shared content across AZsRegional EFSNFS semantics and multi-AZ storage fit shared file access.
Single-AZ re-creatable development shareEFS One Zone may fitThe workload accepts loss of that AZ.
Per-application directory and identity boundaryEFS access pointIt standardizes path and POSIX identity.
Windows SMB and Active Directory requirementEvaluate FSx for Windows File ServerEFS is NFS-oriented rather than managed Windows SMB.

Professional questions normally contain several valid services. State the requirement that selects one option, why the nearest alternative fails it, and what changed requirement would reverse the choice.

AWS Management Console guided practice

Before opening a service page, write the expected account, Region, starting state, and evidence. Do not choose Create, Save, Purchase, Lock, or Delete unless the lesson explicitly authorizes the live track.

  1. Open EFS File systems and inspect type, lifecycle state, encryption, performance mode, throughput mode, lifecycle policy, backup state, and tags.
  2. Open Network and map every mount target to subnet, AZ, IP, and security group; compare against the client AZ list.
  3. Open Access points and supplied mount evidence, then diagnose one DNS, TCP 2049, POSIX identity, or missing-AZ problem.

For each step, capture the field name and value in text. A screenshot may support the record but does not replace the explanation. Console labels can evolve, so use the service search and current documentation if a navigation label differs.

CloudShell and AWS CLI practice

CloudShell is the default browser-based command environment taught in AWS 028. AWS 029 and AWS 030 cover local CLI installation and authentication. This lesson therefore does not assume that an unconfigured local shell is ready.

Start every session with:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account portion of the ARN before sharing. Then perform the topic query:

Inventory file systems, encryption, lifecycle, mount targets, AZ coverage, and endpoint security groups.

NW_EFS_ID="replace-with-owned-file-system-id"
aws efs describe-file-systems --query 'FileSystems[].{Id:FileSystemId,State:LifeCycleState,Encrypted:Encrypted,Mode:PerformanceMode,Throughput:ThroughputMode,Bytes:SizeInBytes.Value}' --output table
aws efs describe-mount-targets --file-system-id "$NW_EFS_ID" --output table

Expected interpretation:

Available lifecycle and mount targets support control-plane readiness. They do not prove NFS reachability, POSIX permission, application locking, throughput, or restore.

Replace every replace-with-... sample value before running its command, and use only an explicitly owned resource. Explain each option first. These queries are read-only; a successful response does not authorize a later create or delete operation.

Practical work

Create p06-efs-decision.md for a two-AZ Linux content share. Specify Regional storage, encryption, mount targets in both private app subnets, client-SG to EFS-SG TCP 2049, TLS mount helper, access point path and UID/GID, lifecycle, throughput mode, backup plan, restore test, monitoring, and cost. Do not create a filesystem.

Add two workload comparisons: (1) 100 web servers reading shared assets and writing rare uploads; (2) one database requiring low-latency block-level writes. Select EFS only for the first and explain why S3 plus deployment artifacts might still be simpler for immutable assets and why EBS/RDS - not EFS - is the database starting point. Draw normal, mount-target-AZ failure, IAM-denied, POSIX-denied and restore paths. Calculate monthly storage by class, metered throughput where applicable, lifecycle access, replication, backup and cross-AZ/client transfer assumptions from dated pricing inputs.

Using supplied mount output, interpret nfs4 options, endpoint, access point and TLS state. Design these controlled tests: create/read/rename/lock a file from two clients; deny TCP 2049; use an incorrect UID; remove one zonal mount path; exhaust a stated throughput boundary; restore a deleted test directory to a separate path. For each, write expected symptom, metric/log/evidence, smallest correction and rollback.

The evidence package must contain:

  • the problem and final requirement in the learner's own words;
  • caller type and Region with private identifiers redacted;
  • exact planned values, ownership, and cost class;
  • one Console observation and matching CLI or API evidence;
  • one behavior result or supplied data-plane record;
  • one denied, failed, or counterexample result and evidence-led diagnosis;
  • one architecture choice plus the rejected alternative;
  • cleanup proof or explicit retained-state owner, expiry, and next lesson.

Verification standard

Use expected state before observed state. Record timestamps in UTC and preserve the original failure before changing anything. A passing submission answers all four questions:

  1. What exact requirement was tested?
  2. Which evidence proves the AWS configuration?
  3. Which evidence proves the workload behavior?
  4. What remains unproven or requires later monitoring?

If AWS returns no rows, verify account, Region, permission, filters, pagination, resource type, and deletion state before concluding that nothing exists.

Common failures and troubleshooting

SymptomEvidence firstLikely boundarySmallest safe response
object appears missingcaller, Region, filters, pagination, tagsscope or read permissionalign scope before creating a duplicate
state remains pending or unavailableservice state, events, dependencies, quotasdependency or capacitycorrect the named dependency and wait with a bound
AccessDeniedprincipal, action, resource, explicit-deny contextidentity, resource, endpoint, organization, or KMS policychange only the proven policy layer
configuration exists but behavior failsroute, DNS, security, listener, health, logs, object versiondata path or applicationtest the next boundary and change one control
bill is higher than expectedhours, bytes, requests, AZs, addresses, retentioncost model or retained resourcestop optional work and reconcile the ledger
cleanup is blockeddependency inventory and owning servicedeletion order or immutable stateremove owned dependants in reviewed reverse order
mount hangs or times outDNS result, client AZ, route/NACL and SG TCP 2049no data-plane path or missing mount targettrace the zonal path; do not open NFS to the internet
access denied by serverfile-system policy, IAM mount options, access pointauthorization policyidentify explicit deny/principal/action before changing policy
mount works but files return permission deniednumeric UID/GID, modes, root squash, access-point POSIX identityPOSIX layeralign intended identity; broad chmod 777 is not diagnosis
latency rises with many tiny filesoperation mix, metadata concurrency, PercentIOLimitworkload/performance modelparallelize safely, reduce metadata chatter, benchmark alternatives
unexpected chargebytes by class, metered I/O, lifecycle access, replication/backupincomplete cost modelreconcile each class and retained copy before changing mode

Do not troubleshoot by attaching administrator access, opening administration ports to the internet, disabling encryption, retrying uncontrolled creation, deleting unknown resources, or weakening retention.

Cost, cleanup, and retained state

No AWS resource is created. Close CloudShell and remove or redact downloaded evidence.

Cleanup evidence requires terminal state and an after-inventory. Search related ENIs, public IPv4 addresses, EBS volumes and snapshots, load balancers, target groups, Auto Scaling instances, endpoints, logs, S3 versions and delete markers, backup recovery points, and global IAM roles when they apply. Billing data can lag, so schedule a later review.

Architecture and certification decisions

  • Certification coverage: SAA-C03; SOA-C03; SAP-C02; DOP-C02.
  • Exam mapping: SAA D1-D4.
  • Explain service scope, failure boundary, consistency, recovery, security, operations, and price rather than matching a keyword.
  • Treat availability and durability, encryption and authorization, routing and filtering, health and lifecycle, backup and replication, and discount and capacity as separate concepts.
  • Do not reproduce protected certification questions.

Knowledge check

  1. Which port does NFS use for EFS?

Expected direction: TCP 2049.

  1. Why create a mount target per client AZ?

Expected direction: To keep a zonally local resilient network path.

  1. Does EFS replace backup?

Expected direction: No. Availability and recoverability are separate.

  1. When is EFS a poor fit?

Expected direction: When object, block, Windows SMB, or specialized high-performance filesystem semantics are required.

Completion gate and assessment

AreaPointsPassing evidence
Requirement and model15Correct scope, terminology, and final outcome
Console evidence15Current path and interpreted fields
CLI or API evidence15Scoped command, expected result, and limitations
Behavior or decision exercise20Reproducible result or defensible architecture reasoning
Troubleshooting15Original symptom, hypothesis, one change, retest, rollback
Security and cost10Least privilege, data protection, current price dimensions
Cleanup and handoff10Terminal-state proof or approved retained-state record

Pass at 80 out of 100 with no critical safety failure. A missing practical artifact, unexplained output, unsafe access, destructive action outside the owned scope, unplanned billed resource, or false cleanup claim requires remediation and a changed retest.

Official sources

Advertisement