RHEL RAID, Device Failure and Recovery
RAID changes availability and performance characteristics; it does not replace backup. A mirrored array can reproduce deletion or corruption immediately, and a rebuild stresses every remaining member. Device replacement begins with exact serial and slot identity, current array metadata, application impact and a tested recovery path. Guessing from /dev/sdX can remove the healthy member.
Separate array redundancy from data recovery
RAID level, member count, spare policy, metadata version and workload determine the failure tolerance. “Degraded” means expected redundancy is absent, not that data is safe indefinitely. Before rebuild, inspect device health, controller/path errors, kernel logs and backups. A path or cable problem can make a healthy device appear absent.
Resynchronization progress alone does not prove filesystem or application integrity. Monitor latency, errors and remaining members, then validate the upper layers after the array returns to the expected state.
Trace a failed I/O from workload to member
| Layer | Question to answer | Evidence |
|---|---|---|
| Application | Which service and consistency model depend on the array? | Error time, workload state and backup |
| Filesystem/LVM | Which upper layers map to md device? | findmnt, lsblk and LVM mapping |
| Array | Which level, state and action are active? | mdadm --detail and /proc/mdstat |
| Member/path | Which serial, WWN, slot and path failed? | udev, platform and kernel evidence |
| Recovery | Can data be restored if another member fails? | Independent backup and tested restore |
Replace one failed lab member deliberately
- Confirm the array, workload impact, backup and console access.
- Map every member path to serial, WWN and physical/virtual slot.
- Capture
mdadm --detail,/proc/mdstatand relevant journal before action. - Fail and remove only the verified lab member if it is not already absent.
- Replace the exact device, reproduce the approved partition layout and add the intended member.
- Monitor rebuild errors, latency and progress; do not reboot merely to speed recovery.
- Verify expected array state, upper-layer filesystem and application data after completion.
Inspect mdraid without changing it
cat /proc/mdstat
mdadm --detail /dev/md0
mdadm --examine /dev/sdb1
lsblk -o NAME,PATH,TYPE,SIZE,MODEL,SERIAL,WWN,FSTYPE,MOUNTPOINTS
udevadm info --query=property --name=/dev/sdb
journalctl -k --since '-2 hours' --no-pager
findmnt -S /dev/md0
smartctl -x /dev/sdb 2>/dev/null
watch -n 5 cat /proc/mdstat- Replace sample devices only after independent identity checks.
mdadm --examinereads member metadata; compare array UUID and role before adding anything.- SMART is one evidence source and may be unavailable or incomplete behind controllers.
- Retain the first kernel errors because later rebuild activity can obscure the initial failure.
Evidence and acceptance criteria
| Evidence | Healthy result | Failure meaning |
|---|---|---|
| Array identity | Expected UUID, level, members and state | Wrong array or metadata conflict |
| Member identity | Serial/WWN/slot agree across layers | Healthy device could be removed |
| Backup | Independent restore point passes validation | Second failure can become data loss |
| Rebuild | Progress advances without member or path errors | Remaining media/path cannot sustain recovery |
| Acceptance | Array, filesystem and application checks pass | Redundancy restored but data/service remains damaged |
Worked scenario: the apparent failed disk is a path failure
A member disappears and the array degrades. A technician replaces the disk named by its current device letter, but the missing member was on another slot and the real fault is a controller path. The replacement removes the remaining valid mirror and turns a recoverable incident into restore work.
The controlled process maps md member UUID to WWN and enclosure slot, correlates transport resets in the kernel journal and verifies the device through the platform. The path is repaired, the correct member is reintroduced under the documented procedure and the independent backup remains protected until application acceptance completes.
Practical how-to cases
Case 1: Create RAID1 in a lab
Mirror two equal disposable partitions and watch initial synchronization. Use stable device identity, confirm backups, and remember RAID availability is not backup.
sudo mdadm --create /dev/md0 --level=1 --raid-devices=2 /dev/vdb1 /dev/vdc1
watch -n 2 cat /proc/mdstat
sudo mdadm --detail /dev/md0
sudo mkfs.xfs /dev/md0| Checkpoint | What to establish |
|---|---|
| Expected result | Both members are active, synchronization completes and the array reports clean. |
| If it fails | A member with prior signatures or different size needs investigation before forced assembly. |
| Safe recovery | Stop only the lab array when required, restore the saved mdadm layout, and recover files from an independent tested backup. |
Case 2: Simulate a member failure
Fail and remove one lab member while proving data remains available. Use stable device identity, confirm backups, and remember RAID availability is not backup.
sudo mdadm /dev/md0 --fail /dev/vdc1
sudo mdadm /dev/md0 --remove /dev/vdc1
cat /proc/mdstat
sudo mdadm --detail /dev/md0
findmnt -S /dev/md0| Checkpoint | What to establish |
|---|---|
| Expected result | The array is degraded with one active member and the mounted test data remains readable. |
| If it fails | Continuing degraded increases risk; identify the physical device before replacement. |
| Safe recovery | Stop only the lab array when required, restore the saved mdadm layout, and recover files from an independent tested backup. |
Case 3: Add a replacement
Copy the reviewed partition layout to a verified empty replacement, add it and monitor rebuild. Use stable device identity, confirm backups, and remember RAID availability is not backup.
sudo sfdisk --dump /dev/vdb > /root/vdb-layout.sfdisk
sudo sfdisk /dev/vdd < /root/vdb-layout.sfdisk
sudo mdadm /dev/md0 --add /dev/vdd1
watch -n 5 cat /proc/mdstat
sudo mdadm --detail /dev/md0| Checkpoint | What to establish |
|---|---|
| Expected result | The new member changes from rebuilding to active and the final array is clean. |
| If it fails | I/O errors, size mismatch or replacement of the wrong physical device requires stopping. |
| Safe recovery | Stop only the lab array when required, restore the saved mdadm layout, and recover files from an independent tested backup. |
Case 4: Examine without assembling
Read superblock identity before deciding whether an inactive array is safe to assemble. Use stable device identity, confirm backups, and remember RAID availability is not backup.
sudo mdadm --examine /dev/vdb1 /dev/vdc1
sudo mdadm --examine --scan
lsblk -o NAME,SIZE,MODEL,SERIAL,FSTYPE,MOUNTPOINTS
journalctl -k -b | grep -i md| Checkpoint | What to establish |
|---|---|
| Expected result | Array UUID, role, event counters and member identity support one unambiguous assembly decision. |
| If it fails | Different UUIDs or event counts can indicate stale members; forced assembly risks choosing older data. |
| Safe recovery | Stop only the lab array when required, restore the saved mdadm layout, and recover files from an independent tested backup. |
Independent practice tasks
- Build RAID1 and store checksum-verified test data.
- Replace one failed member and retain before/after detail output.
- Explain RAID0, RAID1, RAID5, RAID6 and RAID10 failure tolerance.
- Recover an inactive lab array without using force.
For this lesson on RHEL RAID and Device Recovery, complete each task without copying the worked command sequence. Record the initial state, exact change, verification, negative test and recovery command. A task is unfinished if it works now but does not survive a reboot where persistence is required.
Troubleshooting by symptom
| Symptom | Inspect first | Defensible next action |
|---|---|---|
| Array degraded | Detail, mdstat, kernel errors and member identity | Protect workload and backup, then repair exact member/path |
| Rebuild stalls | Kernel I/O errors, remaining member health and controller | Reduce workload and escalate hardware path; preserve recovery options |
| Member reports foreign array | Array UUID, event counter and source history | Do not force-add; determine authoritative data set |
| Array clean but filesystem fails | Upper-layer logs, mount and filesystem health | Use filesystem-specific recovery with backup |
| Repeated member flaps | Power, cable, controller, multipath and device logs | Correct shared path rather than replacing disks repeatedly |
Unsafe operations and recovery boundaries
- Unsafe:
mdadm --createon an existing array can overwrite metadata and destroy assembly evidence. - Unsafe: forced assembly or adding a member with uncertain event history can select stale data.
- Unsafe: removing a device by Linux letter without serial/WWN/slot correlation can remove the healthy member.
Rewritten knowledge checks
Guided lab and acceptance test
- Create two disposable loop devices and a small RAID1 array.
- Format and mount it, write checksummed sample data and record array/member UUIDs.
- Fail and remove one exact loop member; observe degraded state.
- Add a replacement loop device and monitor synchronization.
- Verify checksums and filesystem behavior after recovery.
- Unmount, stop and remove only the lab array and loop mappings.