Lesson 117 · AWS Learning Path

AWS 117: Amazon FSx file systems

· Published · 11 min read

An EC2 instance connects to persistent EBS volumes on one side and fast temporary host-local instance-store disks on the other

The real problem

A team recognizes the name Amazon FSx file systems but has not connected the feature to a real requirement, identity boundary, network or data path, failure mode, price dimension, and cleanup owner. A plausible configuration could still fail the workload.

Final outcome

The learner will produce a requirement-led artifact for Amazon FSx file systems, inspect the matching AWS control plane in the Management Console, run a matching CloudShell or AWS CLI query, interpret the output, diagnose one failure, defend one architecture choice, and prove cleanup or approved retained state.

The practical outcome is not a command transcript. It must show what was expected, what happened, what the result proves, what it does not prove, and which evidence would change the decision.

Learning objectives

By the end of this lesson, the learner can:

  • explain fsx family;
  • explain windows file server;
  • explain lustre;
  • explain netapp ontap;
  • explain openzfs;
  • connect control-plane state to the real data, network, identity, or application behavior;
  • identify cost and cleanup ownership before any optional mutation;
  • troubleshoot from evidence without opening broad access or adding broad permissions.

Relationship model

Requirement
   |
   v
Identity and policy -> AWS configuration -> network or data path -> workload behavior
        |                    |                      |                    |
        +--------------------+----------------------+--------------------+
                                      |
                                      v
                         monitoring, cost, recovery, cleanup

Use this model to separate an AWS object that exists from a result that actually works. Every arrow is a verification boundary.

Prerequisites, permissions, Region, and safety

  • Learning baseline: This sequence assumes practical Linux knowledge but no prior cloud-computing or AWS knowledge. Cloud, networking, security, data, automation, and architecture concepts must come from completed earlier lessons. If a prerequisite checkpoint is incomplete, return to its linked lesson before continuing.
  • Confirm a non-root caller with aws sts get-caller-identity and keep the account number private.
  • Use ap-south-1 unless this lesson explicitly names a second Region.
  • Confirm the intended profile and Region with aws configure list before interpreting an empty result.
  • Use read-only List, Get, and Describe permissions for the named services. Design exercises run locally and require no resource-creation permission.
  • This is a no-create lesson. Console and CLI work is read-only, and every design artifact is created locally.
  • Never publish account IDs, public addresses, ARNs containing private account data, session IDs, presigned URLs, object data, credentials, or KMS material.
  • Do not use root, world-open SSH or RDP, disabled TLS verification, unowned resources, or irreversible retention controls in a training exercise.

Core model

ConceptWhat the learner must understand
FSx familyAmazon FSx provides managed file systems based on Windows File Server, Lustre, NetApp ONTAP, and OpenZFS. These are different products, not performance tiers of one interchangeable filesystem.
Windows File ServerFSx for Windows File Server provides SMB, Windows permissions, Active Directory integration, and Windows-compatible features for managed Windows file workloads.
LustreFSx for Lustre targets high-performance parallel workloads such as HPC, machine learning, media processing, and analytics, with optional integration to S3 datasets.
NetApp ONTAPFSx for NetApp ONTAP supports multiprotocol access and ONTAP data-management capabilities for workloads that require that ecosystem.
OpenZFSFSx for OpenZFS provides managed ZFS-based file storage with NFS access and ZFS data-management behavior for compatible Linux and migration workloads.
Deployment detailsAvailability model, protocol, throughput capacity, SSD or HDD options, backups, replication, networking, directory dependency, and minimum provisioned capacity vary by FSx type.

How it works

FSx is a family, not one interchangeable filesystem. AWS manages infrastructure and lifecycle around four established engines, while the learner still owns protocol access, directory integration, client behavior, capacity/performance choices, backup/restore and application compatibility:

FamilyNative fitImportant resources and boundaries
FSx for Windows File ServerWindows SMB shares, NTFS semantics, Active Directory integrationfile system, Windows file servers, DNS aliases, shares, AWS Managed Microsoft AD or self-managed AD; Single-AZ or Multi-AZ choices
FSx for Lustreparallel high-performance compute, ML, rendering and S3-linked datasetsdeployment type, storage type/capacity, throughput, data repository association/tasks and Lustre clients
FSx for NetApp ONTAPmultiprotocol enterprise storage, ONTAP features and migrationsfile system/HA pairs, storage virtual machines, volumes, junctions, NFS/SMB/iSCSI, snapshots, tiering and SnapMirror
FSx for OpenZFSmanaged OpenZFS with NFS, snapshots and clonesfile system, volumes, NFS exports, record/compression settings, snapshots and clones; deployment type affects failure boundary

The request path differs. Windows is commonly domain client -> DNS/Kerberos/SMB -> file-server endpoint -> NTFS/share ACL -> storage; Lustre is compute client -> Lustre network endpoint -> metadata/storage servers, optionally moving data to/from S3; ONTAP adds an SVM and protocol endpoint before the volume; OpenZFS presents NFS exports from filesystem endpoints. “The security group allows the port” proves only one layer. AD trust/time/DNS, Kerberos, share ACLs, export policy, UNIX UID mapping, client packages and mount options can still fail.

Capacity, performance and lifecycle

Choose storage capacity, SSD/HDD where supported, throughput capacity, deployment type, backups and update window together. A large namespace does not guarantee required throughput; a high-throughput configuration does not repair a serial application. Model throughput, IOPS, latency, metadata operations, client concurrency, working-set size and growth. Validate current minimums, increments, supported updates and Region availability rather than memorizing values that change.

Lustre data repository associations are data-movement relationships, not a magic POSIX view of S3: imports/exports have task state, metadata behavior, cost and consistency considerations. ONTAP capacity pools and tiering require working-set and access-cost analysis. Windows shadow copies, ONTAP/OpenZFS snapshots and AWS Backup/native backups are useful recovery mechanisms but do not replace application-consistent restore tests. Multi-AZ deployment protects specified infrastructure failures; it does not prevent deletion, bad ACL changes, malware or regional loss.

Encryption at rest, KMS permissions, TLS/in-transit protocol support, security groups, directory credentials, filesystem/export/share policies and OS permissions are distinct. Keep endpoints private and allow only exact client security groups/CIDRs and required directory/DNS flows. Monitor free capacity, throughput/IOPS utilization, latency, network, metadata pressure, failover events, backup age and client-visible errors.

Choose the filesystem implementation from protocol and workload semantics first. Then validate AZ design, client compatibility, identity, throughput, IOPS, latency, capacity growth, backup, replication, failover, monitoring, maintenance, and cost.

Read the result in layers:

  1. Scope: account, Region, VPC, bucket, AZ, endpoint, principal, object version, or resource ARN.
  2. Control plane: the requested configuration exists and reached an expected state.
  3. Behavior: the request, connection, health check, replication, restore, or application result meets the requirement.
  4. Operations: monitoring, failure owner, cost, retention, rollback, and cleanup are known.

Control-plane success is necessary but not sufficient. A resource can be available while policy, routing, DNS, health, data, or application behavior remains wrong.

Architecture decision table

RequirementPreferred directionWhy
Windows home directories with SMB and ADFSx for Windows File ServerManaged Windows filesystem and identity behavior fit.
Parallel processing of a large S3 datasetFSx for LustreHigh-throughput parallel access and S3 integration fit.
Multiprotocol enterprise NAS using ONTAP featuresFSx for NetApp ONTAPThe workload depends on ONTAP semantics and management.
Migrate NFS ZFS workflowsFSx for OpenZFSManaged OpenZFS behavior aligns with the source and client needs.

Professional questions normally contain several valid services. State the requirement that selects one option, why the nearest alternative fails it, and what changed requirement would reverse the choice.

AWS Management Console guided practice

Before opening a service page, write the expected account, Region, starting state, and evidence. Do not choose Create, Save, Purchase, Lock, or Delete unless the lesson explicitly authorizes the live track.

  1. Open Amazon FSx and compare the four filesystem creation choices without starting a wizard past the review stage.
  2. Inspect supplied filesystems for type, deployment, storage, throughput, network, DNS, directory, backup, maintenance, and monitoring details.
  3. Complete a workload matrix for Windows SMB, HPC scratch, multiprotocol NAS, and ZFS migration and defend one rejection per option.

For each step, capture the field name and value in text. A screenshot may support the record but does not replace the explanation. Console labels can evolve, so use the service search and current documentation if a navigation label differs.

CloudShell and AWS CLI practice

CloudShell is the default browser-based command environment taught in AWS 028. AWS 029 and AWS 030 cover local CLI installation and authentication. This lesson therefore does not assume that an unconfigured local shell is ready.

Start every session with:

export AWS_DEFAULT_REGION="ap-south-1"
aws sts get-caller-identity --query Arn --output text
aws configure list

Redact the account portion of the ARN before sharing. Then perform the topic query:

Inventory FSx filesystem type, lifecycle, deployment, capacity, throughput, network, and DNS without creating one.

aws fsx describe-file-systems --query 'FileSystems[].{Id:FileSystemId,Type:FileSystemType,State:Lifecycle,StorageGiB:StorageCapacity,StorageType:StorageType,Throughput:ThroughputCapacity,Vpc:VpcId,Subnets:SubnetIds,DNS:DNSName}' --output table

Expected interpretation:

An AVAILABLE state supports control-plane readiness. It does not prove client protocol, directory trust, mount permission, application performance, failover, or restore.

Replace every replace-with-... sample value before running its command, and use only an explicitly owned resource. Explain each option first. These queries are read-only; a successful response does not authorize a later create or delete operation.

Practical work

Create p06-fsx-selection.md. Compare all four FSx services for protocol, clients, identity, AZ model, storage, throughput, latency, S3 integration, snapshots or backups, replication, failover, maintenance, monitoring, minimum footprint, and current ap-south-1 availability and price. Do not create a filesystem.

Apply the comparison to four cases: Windows departmental shares with AD and Multi-AZ requirements; a temporary 500-node genomics job whose input/output resides in S3; an ONTAP migration needing NFS and SMB plus snapshot replication; and a Linux application needing OpenZFS snapshots/clones. For each, choose exact family/deployment/protocol, draw identity and data paths, define benchmark and restore tests, list the nearest rejected option, and model idle capacity/throughput, backups, tiering/data movement and transfer.

The evidence package must contain:

  • the problem and final requirement in the learner's own words;
  • caller type and Region with private identifiers redacted;
  • exact planned values, ownership, and cost class;
  • one Console observation and matching CLI or API evidence;
  • one behavior result or supplied data-plane record;
  • one denied, failed, or counterexample result and evidence-led diagnosis;
  • one architecture choice plus the rejected alternative;
  • cleanup proof or explicit retained-state owner, expiry, and next lesson.

Verification standard

Use expected state before observed state. Record timestamps in UTC and preserve the original failure before changing anything. A passing submission answers all four questions:

  1. What exact requirement was tested?
  2. Which evidence proves the AWS configuration?
  3. Which evidence proves the workload behavior?
  4. What remains unproven or requires later monitoring?

If AWS returns no rows, verify account, Region, permission, filters, pagination, resource type, and deletion state before concluding that nothing exists.

Common failures and troubleshooting

SymptomEvidence firstLikely boundarySmallest safe response
object appears missingcaller, Region, filters, pagination, tagsscope or read permissionalign scope before creating a duplicate
state remains pending or unavailableservice state, events, dependencies, quotasdependency or capacitycorrect the named dependency and wait with a bound
AccessDeniedprincipal, action, resource, explicit-deny contextidentity, resource, endpoint, organization, or KMS policychange only the proven policy layer
configuration exists but behavior failsroute, DNS, security, listener, health, logs, object versiondata path or applicationtest the next boundary and change one control
bill is higher than expectedhours, bytes, requests, AZs, addresses, retentioncost model or retained resourcestop optional work and reconcile the ledger
cleanup is blockeddependency inventory and owning servicedeletion order or immutable stateremove owned dependants in reviewed reverse order
Windows share cannot authenticateAD health, DNS, time, trust, SPN/Kerberos and SMB errordirectory/identity before storagerepair the proven dependency; do not switch to anonymous access
Lustre mount or job stallsclient version/modules, endpoint path, server/metadata metricsclient/network/performancevalidate supported client and trace metadata versus data pressure
S3-linked file is absent or stalerepository association, task state/error, path and metadatadata movementinspect import/export task evidence; do not assume synchronous mirroring
ONTAP/OpenZFS NFS permission deniedSVM/endpoint, export policy, UNIX identity and modeprotocol/authorizationchange only the failed export/identity layer
backup exists but app cannot recoverrestore resource, mount/share/export, permissions and app checkrecovery validationrestore separately and run application integrity tests

Do not troubleshoot by attaching administrator access, opening administration ports to the internet, disabling encryption, retrying uncontrolled creation, deleting unknown resources, or weakening retention.

Cost, cleanup, and retained state

No AWS resource is created. Close CloudShell and remove or redact downloaded evidence.

Cleanup evidence requires terminal state and an after-inventory. Search related ENIs, public IPv4 addresses, EBS volumes and snapshots, load balancers, target groups, Auto Scaling instances, endpoints, logs, S3 versions and delete markers, backup recovery points, and global IAM roles when they apply. Billing data can lag, so schedule a later review.

Architecture and certification decisions

  • Certification coverage: SAA-C03; SOA-C03; SAP-C02; DOP-C02.
  • Exam mapping: SAA D1-D4.
  • Explain service scope, failure boundary, consistency, recovery, security, operations, and price rather than matching a keyword.
  • Treat availability and durability, encryption and authorization, routing and filtering, health and lifecycle, backup and replication, and discount and capacity as separate concepts.
  • Do not reproduce protected certification questions.

Knowledge check

  1. Which FSx service fits managed Windows SMB?

Expected direction: FSx for Windows File Server.

  1. Which FSx service fits parallel HPC access?

Expected direction: FSx for Lustre.

  1. Are FSx products merely storage classes?

Expected direction: No. They implement different filesystem technologies and protocols.

  1. What should decide before price?

Expected direction: Protocol, application semantics, identity, availability, and performance requirements.

Completion gate and assessment

AreaPointsPassing evidence
Requirement and model15Correct scope, terminology, and final outcome
Console evidence15Current path and interpreted fields
CLI or API evidence15Scoped command, expected result, and limitations
Behavior or decision exercise20Reproducible result or defensible architecture reasoning
Troubleshooting15Original symptom, hypothesis, one change, retest, rollback
Security and cost10Least privilege, data protection, current price dimensions
Cleanup and handoff10Terminal-state proof or approved retained-state record

Pass at 80 out of 100 with no critical safety failure. A missing practical artifact, unexplained output, unsafe access, destructive action outside the owned scope, unplanned billed resource, or false cleanup claim requires remediation and a changed retest.

Official sources

Advertisement