Lesson 011 · Linux Administration Learning Path

RHEL RAID, Device Failure and Recovery

· Published · 7 min read

Labelled RHEL Linux administration path highlighting RAID members array state degraded service replacement resynchronization backup and acceptance

RAID changes availability and performance characteristics; it does not replace backup. A mirrored array can reproduce deletion or corruption immediately, and a rebuild stresses every remaining member. Device replacement begins with exact serial and slot identity, current array metadata, application impact and a tested recovery path. Guessing from /dev/sdX can remove the healthy member.

Separate array redundancy from data recovery

RAID level, member count, spare policy, metadata version and workload determine the failure tolerance. “Degraded” means expected redundancy is absent, not that data is safe indefinitely. Before rebuild, inspect device health, controller/path errors, kernel logs and backups. A path or cable problem can make a healthy device appear absent.

Resynchronization progress alone does not prove filesystem or application integrity. Monitor latency, errors and remaining members, then validate the upper layers after the array returns to the expected state.

Trace a failed I/O from workload to member

LayerQuestion to answerEvidence
ApplicationWhich service and consistency model depend on the array?Error time, workload state and backup
Filesystem/LVMWhich upper layers map to md device?findmnt, lsblk and LVM mapping
ArrayWhich level, state and action are active?mdadm --detail and /proc/mdstat
Member/pathWhich serial, WWN, slot and path failed?udev, platform and kernel evidence
RecoveryCan data be restored if another member fails?Independent backup and tested restore

Replace one failed lab member deliberately

  1. Confirm the array, workload impact, backup and console access.
  2. Map every member path to serial, WWN and physical/virtual slot.
  3. Capture mdadm --detail, /proc/mdstat and relevant journal before action.
  4. Fail and remove only the verified lab member if it is not already absent.
  5. Replace the exact device, reproduce the approved partition layout and add the intended member.
  6. Monitor rebuild errors, latency and progress; do not reboot merely to speed recovery.
  7. Verify expected array state, upper-layer filesystem and application data after completion.

Inspect mdraid without changing it

cat /proc/mdstat
mdadm --detail /dev/md0
mdadm --examine /dev/sdb1
lsblk -o NAME,PATH,TYPE,SIZE,MODEL,SERIAL,WWN,FSTYPE,MOUNTPOINTS
udevadm info --query=property --name=/dev/sdb
journalctl -k --since '-2 hours' --no-pager
findmnt -S /dev/md0
smartctl -x /dev/sdb 2>/dev/null
watch -n 5 cat /proc/mdstat
  • Replace sample devices only after independent identity checks.
  • mdadm --examine reads member metadata; compare array UUID and role before adding anything.
  • SMART is one evidence source and may be unavailable or incomplete behind controllers.
  • Retain the first kernel errors because later rebuild activity can obscure the initial failure.

Evidence and acceptance criteria

EvidenceHealthy resultFailure meaning
Array identityExpected UUID, level, members and stateWrong array or metadata conflict
Member identitySerial/WWN/slot agree across layersHealthy device could be removed
BackupIndependent restore point passes validationSecond failure can become data loss
RebuildProgress advances without member or path errorsRemaining media/path cannot sustain recovery
AcceptanceArray, filesystem and application checks passRedundancy restored but data/service remains damaged

Worked scenario: the apparent failed disk is a path failure

A member disappears and the array degrades. A technician replaces the disk named by its current device letter, but the missing member was on another slot and the real fault is a controller path. The replacement removes the remaining valid mirror and turns a recoverable incident into restore work.

The controlled process maps md member UUID to WWN and enclosure slot, correlates transport resets in the kernel journal and verifies the device through the platform. The path is repaired, the correct member is reintroduced under the documented procedure and the independent backup remains protected until application acceptance completes.

Practical how-to cases

Case 1: Create RAID1 in a lab

Mirror two equal disposable partitions and watch initial synchronization. Use stable device identity, confirm backups, and remember RAID availability is not backup.

sudo mdadm --create /dev/md0 --level=1 --raid-devices=2 /dev/vdb1 /dev/vdc1
watch -n 2 cat /proc/mdstat
sudo mdadm --detail /dev/md0
sudo mkfs.xfs /dev/md0
CheckpointWhat to establish
Expected resultBoth members are active, synchronization completes and the array reports clean.
If it failsA member with prior signatures or different size needs investigation before forced assembly.
Safe recoveryStop only the lab array when required, restore the saved mdadm layout, and recover files from an independent tested backup.

Case 2: Simulate a member failure

Fail and remove one lab member while proving data remains available. Use stable device identity, confirm backups, and remember RAID availability is not backup.

sudo mdadm /dev/md0 --fail /dev/vdc1
sudo mdadm /dev/md0 --remove /dev/vdc1
cat /proc/mdstat
sudo mdadm --detail /dev/md0
findmnt -S /dev/md0
CheckpointWhat to establish
Expected resultThe array is degraded with one active member and the mounted test data remains readable.
If it failsContinuing degraded increases risk; identify the physical device before replacement.
Safe recoveryStop only the lab array when required, restore the saved mdadm layout, and recover files from an independent tested backup.

Case 3: Add a replacement

Copy the reviewed partition layout to a verified empty replacement, add it and monitor rebuild. Use stable device identity, confirm backups, and remember RAID availability is not backup.

sudo sfdisk --dump /dev/vdb > /root/vdb-layout.sfdisk
sudo sfdisk /dev/vdd < /root/vdb-layout.sfdisk
sudo mdadm /dev/md0 --add /dev/vdd1
watch -n 5 cat /proc/mdstat
sudo mdadm --detail /dev/md0
CheckpointWhat to establish
Expected resultThe new member changes from rebuilding to active and the final array is clean.
If it failsI/O errors, size mismatch or replacement of the wrong physical device requires stopping.
Safe recoveryStop only the lab array when required, restore the saved mdadm layout, and recover files from an independent tested backup.

Case 4: Examine without assembling

Read superblock identity before deciding whether an inactive array is safe to assemble. Use stable device identity, confirm backups, and remember RAID availability is not backup.

sudo mdadm --examine /dev/vdb1 /dev/vdc1
sudo mdadm --examine --scan
lsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,MOUNTPOINTS
journalctl -k -b | grep -i md
CheckpointWhat to establish
Expected resultArray UUID, role, event counters and member identity support one unambiguous assembly decision.
If it failsDifferent UUIDs or event counts can indicate stale members; forced assembly risks choosing older data.
Safe recoveryStop only the lab array when required, restore the saved mdadm layout, and recover files from an independent tested backup.

Independent practice tasks

  1. Build RAID1 and store checksum-verified test data.
  2. Replace one failed member and retain before/after detail output.
  3. Explain RAID0, RAID1, RAID5, RAID6 and RAID10 failure tolerance.
  4. Recover an inactive lab array without using force.

For this lesson on RHEL RAID and Device Recovery, complete each task without copying the worked command sequence. Record the initial state, exact change, verification, negative test and recovery command. A task is unfinished if it works now but does not survive a reboot where persistence is required.

Troubleshooting by symptom

SymptomInspect firstDefensible next action
Array degradedDetail, mdstat, kernel errors and member identityProtect workload and backup, then repair exact member/path
Rebuild stallsKernel I/O errors, remaining member health and controllerReduce workload and escalate hardware path; preserve recovery options
Member reports foreign arrayArray UUID, event counter and source historyDo not force-add; determine authoritative data set
Array clean but filesystem failsUpper-layer logs, mount and filesystem healthUse filesystem-specific recovery with backup
Repeated member flapsPower, cable, controller, multipath and device logsCorrect shared path rather than replacing disks repeatedly

Unsafe operations and recovery boundaries

  • Unsafe: mdadm --create on an existing array can overwrite metadata and destroy assembly evidence.
  • Unsafe: forced assembly or adding a member with uncertain event history can select stale data.
  • Unsafe: removing a device by Linux letter without serial/WWN/slot correlation can remove the healthy member.

Rewritten knowledge checks

Does RAID replace backup?
No. It does not protect against deletion, corruption, compromise or array-wide failure.
What does degraded mean?
The array is operating without its expected redundancy or membership.
Why record array UUID?
It distinguishes membership and prevents assembling unrelated devices by path alone.
What is resynchronization?
Copying/reconciling array data and parity so the intended member state is restored.
Why monitor latency during rebuild?
Rebuild consumes I/O and may harm the workload or expose failing media.
Can a device path equal physical identity?
No. Device enumeration can change; use serial, WWN and slot evidence.
What must be checked after rebuild?
Array state, errors, filesystem/LVM and application integrity.
What is the safest first response to a degraded array?
Preserve evidence, protect workload and verify backup and exact member identity before changing membership.

Guided lab and acceptance test

  1. Create two disposable loop devices and a small RAID1 array.
  2. Format and mount it, write checksummed sample data and record array/member UUIDs.
  3. Fail and remove one exact loop member; observe degraded state.
  4. Add a replacement loop device and monitor synchronization.
  5. Verify checksums and filesystem behavior after recovery.
  6. Unmount, stop and remove only the lab array and loop mappings.

Primary references

Advertisement