ZFS Pool Shows Degraded But No Drive Is Marked Failed: What’s Going On?
You see zpool status reporting DEGRADED, but every drive in the pool shows ONLINE with no FAULTED or UNAVAIL flags. Your instinct says “replace a drive,” but nothing is clearly failing. This is a confusing and stressful position. The short answer: ZFS sets the pool to DEGRADED when it detects data integrity issues — often accumulating checksum errors — even if no single drive has met the software’s threshold to be marked failed. This guide walks you through the exact causes, diagnostic steps, and a concrete decision checklist so you know whether you need to buy a new drive or fix a cable. This guide breaks down zfs pool degraded no drive failed in practical terms.
Why ZFS Shows Degraded Without a Failed Drive
ZFS is unusually sensitive to read errors. Unlike traditional RAID controllers that may silently correct a few bad sectors and move on, ZFS tracks every checksum mismatch on every drive. When the total number of uncorrectable read errors or checksum mismatches exceeds a certain threshold, the pool state flips from ONLINE to DEGRADED — even if no single drive has been flagged as failed.
ZFS does not automatically mark a drive as FAULTED after a single read error. It accumulates errors per device. A drive with hundreds of checksum errors can still show ONLINE while the pool reports DEGRADED. The pool state reflects overall data integrity risk, not per-drive hardware health.
Common scenarios that trigger this:
- Checksum errors on one drive — the drive is readable but some blocks returned bad data that ZFS corrected from parity or mirror.
- Read errors across multiple drives — a marginal cable or controller can cause scattered errors on several disks at once, degrading the pool without a single obvious culprit.
- Transient errors during a scrub or heavy I/O — a one-time glitch from a loose SATA cable or power dip can spike error counters.
zpool Status DEGRADED Troubleshooting: Step-by-Step
Here is the exact sequence to follow when you see a degraded pool with no failed drive. Do not skip steps.
Run zpool status -v. Look at the CKSUM, READ, and WRITE columns. A drive with zero errors in all columns is likely fine. A drive with dozens or hundreds of checksum errors is suspect, even if marked ONLINE.
Execute zpool scrub poolname. The scrub will attempt to read every block and repair any data it can from parity or mirrors. After completion, check zpool status -v again. If the scrub repaired all errors and error counters are now zero, the issue was likely transient.
Use smartctl -a /dev/sdX on each disk. Focus on Reallocated_Sector_Ct, Current_Pending_Sector, and UDMA_CRC_Error_Count. A high CRC error count (hundreds or thousands) points to a cable or backplane issue, not the drive itself.
If multiple drives show CRC errors, reseat all SATA/SAS cables and power connectors. If using an HBA, check for firmware updates. A failing HBA can corrupt data across the entire pool without any single drive failing.
Log the output of zpool status -v before and after the scrub. If error counters grow again within a few days, you have an active problem. If they stay at zero, the issue was transient.
Non-Drive Causes That Mimic Drive Failure
Before you order a replacement drive, rule out these common culprits. They are far more common than actual drive failure in a pool with no single failed disk.
| Cause | Symptoms | Fix |
|---|---|---|
| Bad SATA/SAS cable | CRC errors on one or two drives, intermittent disconnects | Replace cable; check for bent pins |
| Failing HBA or expander | Errors across multiple drives on the same controller port | Reseat card, update firmware, try a different PCIe slot |
| Marginal PSU | Errors during high I/O, drives dropping offline briefly | Check 12V rail voltage under load; replace PSU if unstable |
| Faulty RAM (ECC or non-ECC) | Silent data corruption, phantom errors, system crashes | Run memtest86 for at least 24 hours |
Do not ignore RAM errors. A single bit flip in a ZFS metadata write can corrupt the entire pool. If you suspect RAM, stop all writes and test immediately.
zpool Clear vs. Replace Disk: When Each Is Safe
The zpool clear command resets error counters on a device or the entire pool. It does not fix underlying problems. Use this table to decide which action to take.
Safe to Use zpool clear
- Scrub completed with zero errors after the clear
- SMART shows no reallocated or pending sectors
- CRC errors were resolved by reseating cables
- Error counters were low (single digits) and not growing
Replace the Drive Instead
- Scrub failed to repair some blocks (permanent errors)
- SMART shows increasing reallocated sector count
- Error counters grow again within 24 hours after clear
- Multiple drives show errors on the same controller simultaneously
ZFS Checksum Errors No Failed Disk: Decision Checklist
Use this checklist to decide with confidence whether to replace a drive, fix a cable, or do nothing.
- Run
zpool status -vand note which drives have checksum errors. - Run a full scrub. After completion, check if errors were repaired or remain.
- Check SMART attributes: reallocated sectors, pending sectors, CRC errors.
- If CRC errors are high on multiple drives, inspect cables and backplane.
- If one drive has reallocated sectors and the scrub could not repair some blocks, replace that drive.
- If all error counters are zero after scrub and SMART is clean, run
zpool clear poolnameand monitor for 7 days. - If errors return within 7 days, replace the drive or HBA.
In a RAIDZ2 or mirrored pool, you can safely replace a suspect drive without downtime. For RAIDZ1, you are vulnerable during rebuild — ensure you have a good backup before replacing. Our guide on RAID 5 vs RAID 6 explains why RAIDZ2 is safer for large drives.
How ZFS Snapshots Help After a Degraded Event
If a degraded pool led to data corruption that the scrub could not fully repair, your last line of defense is a clean snapshot. Restoring from a snapshot taken before the corruption occurred is far faster and more reliable than trying to piece together damaged files. For a full walkthrough, see our guide on ZFS Snapshots Explained: A Homelab Guide to Instant Rollbacks.
ZFS snapshots are instantaneous and consume no space until data changes. If you are not running automated snapshots on your pool, start now — they are your best recovery tool after a scrub fails.
Bottom Line: When to Replace vs. When to Wait
If your pool is DEGRADED but no drive shows FAULTED, do not panic. Start with a scrub and a SMART check on every disk. If the scrub completes with zero errors and SMART is clean, run zpool clear and monitor for a week. If errors return, replace the drive with the highest reallocated sector count or the one showing the most checksum errors. If multiple drives show CRC errors, fix cables or the HBA first. Only replace a drive when you have clear evidence — growing error counters after a clear, reallocated sectors, or uncorrectable scrub errors. In RAIDZ2 or mirrors, you have room to test. In RAIDZ1, be more conservative: replace at the first sign of reallocated sectors.
Frequently Asked Questions
Does DEGRADED always mean a drive is about to fail?
No. DEGRADED means ZFS detected data integrity issues — typically checksum errors — that could be caused by a bad cable, a failing HBA, marginal RAM, or a transient power glitch. A drive can have zero hardware defects and still trigger DEGRADED if a cable introduces errors. Always run a scrub and check SMART data before concluding the drive is failing. In many cases, reseating cables and running zpool clear resolves the issue permanently.
Can bad cables or a bad controller cause a ZFS pool to go degraded?
Yes, absolutely. A loose SATA cable, a damaged SAS cable, or a failing HBA can introduce CRC errors on multiple drives simultaneously. This often shows as UDMA_CRC_Error_Count in SMART data across several disks. If your pool went degraded and multiple drives show CRC errors, the cable or controller is the likely culprit. Replacing a $10 cable can fix a problem that looks like a $200 drive failure.
Is it safe to just run zpool clear and move on?
It is safe only if you have confirmed the underlying cause is resolved. Run a full scrub first. If the scrub completes with zero errors and SMART is clean, then zpool clear is appropriate. If you clear errors without diagnosing the cause, you risk masking a failing drive that will eventually cause data loss. Always verify that error counters do not return within 7 days after clearing.
How do I tell if my degraded pool is a real problem or a transient glitch?
Run a scrub and compare error counters before and after. If the scrub repaired all errors and counters are zero, the issue was likely transient — possibly a one-time power fluctuation or cable wiggle. If errors persist after the scrub, or if SMART shows reallocated sectors, you have an active hardware problem. Monitor the pool for 7 days after clearing. If errors return, the problem is real and needs component replacement.
Last verified: July 10, 2026. Specifications cross-checked against OpenZFS documentation, Oracle Solaris ZFS Administration Guide, and community-tested procedures from Proxmox and TrueNAS forums.
🛡 Shop Recommended Hardware
Prices and stock verified regularly by our affiliate partners. As an affiliate, HomeLabCost may earn a commission on qualifying purchases at no extra cost to you.
Browse Hardware Picks →