What does a shingled (SMR) hard drive failure look like in logs, and what does SMART miss?
Published 2026-10-11
The public SMR failures are mostly not worn-out media but a write-cache overflow: a drive-managed SMR drive absorbs writes in a small conventional area and, once that fills, slows by one to three orders of magnitude, and in some reports it then returns errors and is dropped from the array. Measured slowdowns: 13.2 MiB/s against 209.3 MiB/s for 32 KiB writes on a 4 TB WD Red EFAX (Ars Technica, 15.9:1), nearly 230 h against under 17 h for a 4-drive RAIDZ resilver (ServeTheHome), and about 0.1 MB/s for more than 25 min in a prototype host-aware SMR drive (University of Minnesota paper). In the one failing resilver whose logs were published (OpenZFS issue 10214) SMART overall health said PASSED and the reallocated, pending, uncorrectable and CRC counters were all 0, while the ATA error log counted 7,487 errors (the 24 it still showed were all IDNF) and ZFS counted 47.7 K write errors. Only one lab (iXsystems) and some users saw the failed state; ServeTheHome and Ars, testing the same drive model, saw slowness only, and a CMR drive on a shared controller logged the same error text.
- Question
- When a drive-managed SMR disk is overloaded with writes, which public signals (SMART, ATA error log, kernel log, filesystem counters, timing) show it, and which show nothing?
- Drives
- WD Red 4 TB WD40EFAX (DM-SMR, firmware 82.00A82) in four independent tests or reports; one prototype Seagate 8 TB host-aware SMR drive in a university benchmark.
- Failure layers
- Persistent-cache exhaustion and blocking cleaning; ATA IDNF errors and SCSI 'logical block address out of range'; array software dropping a slow or erroring member.
- Not covered
- SMR failure rates, a ranking of SMR against CMR, host-managed zoned drives in production, and any recovery of data from affected drives.
How far a drive-managed SMR drive slows once its cache is full
| Measurement | SMR drive result | Comparison result | Ratio | Source |
|---|---|---|---|---|
| 32 KiB sequential writes, imitating one ZFS RAIDZ member (WD Red 4 TB EFAX vs Seagate IronWolf 12 TB) | 13.2 MiB/s | 209.3 MiB/s | 15.9:1 slower (as published) | Ars Technica, 2020 |
| 4-drive RAIDZ resilver of an array about 60% full, with network load (WD40EFAX vs three CMR drives) | nearly 230 h (9 days 14 h) | under 17 h for each CMR drive | more than 13:1 slower (derived, 230 / 17) | ServeTheHome, 2020 |
| Peak latency of one 1 MiB random write (WD Red 4 TB EFAX vs IronWolf) | 1.3 s | 108 ms | about 12:1 slower (derived, 1,300 / 108) | Ars Technica, 2020 |
| ZFS resilver of a WD40EFAX with write cache and look-ahead switched off (one user, via interview) | 8 days at under 6 MB/s | 24 h, which the user calls the more usual time | about 8:1 slower (derived, 8 days / 24 h) | Blocks & Files, 2020 |
| Blocking cache cleaning in a prototype host-aware SMR drive, 256 KB non-sequential writes for 2 h | about 0.1 MB/s for more than 25 min (128 zones) and more than 37 min (256 zones) | over 100 MB/s before cleaning starts | about 1,000:1 slower (derived, 100 / 0.1) | Wu et al., HA-SMR evaluation |
What each signal showed during one failing ZFS resilver
| Signal | Value in the report | What it shows | Source |
|---|---|---|---|
| SMART overall-health self-assessment | PASSED | No warning while the drive was failing the resilver | OpenZFS issue 10214, 2020 |
| SMART attributes 5, 197, 198, 199 (raw values, count) | 0, 0, 0, 0 (reallocated sectors, pending sectors, offline uncorrectable, UDMA CRC errors) | No media or link-integrity counter moved | OpenZFS issue 10214, 2020 |
| ATA Extended Comprehensive SMART Error Log (count) | Device Error Count 7,487; shown errors are 'IDNF at LBA' on WRITE FPDMA QUEUED commands; the log keeps only the latest 24 | The only drive-side record of the failure; a commenter says it needs smartctl -x, because -a does not print the extended log | OpenZFS issue 10214, 2020 |
| ZFS zpool status WRITE column (count) | 47.7 K write errors on the new disk at 27.61% resilvered, 4.69 TB issued of 17.0 TB | Filesystem-level view of the same failure; the poster later suspects the counter wraps at about 65.5 K | OpenZFS issue 10214, 2020 |
| Linux kernel log, SCSI sense | Sense Key: Illegal Request; Add. Sense: Logical block address out of range; Write(16) commands; 'critical target error' | The host-side translation of the drive's IDNF reply | OpenZFS issue 10214, 2020 |
| SATA PHY event counters (count) | CRC-error counters 0; PhyRdy to PhyNRdy transitions 82 | No CRC errors on the link, but 82 link-ready drops were recorded in that log | OpenZFS issue 10214, 2020 |
Who saw a failed state, and who saw only slowness
| Report | Drive and workload | Outcome | Source |
|---|---|---|---|
| iXsystems lab test (2020) | WD40EFAX, firmware 82.00A82, heavy write loads including resilvering | Drive entered a faulty state, returned IDNF, became unusable and was treated as failed by ZFS; about 100 Minis shipped with 2 TB or 6 TB DM-SMR drives and one issue reported | iXsystems / TrueNAS notice, 2020 |
| OpenZFS issue 10214 (user, 16 April 2020) | WD40EFAX, zpool replace into a 5-disk raidz1 | Stalled at 27.61% with IOPS falling to zero and 47.7 K write errors; the same disk later resilvered 7.26 G in 4 min 24 s with 2 write errors after a sequential copy | OpenZFS issue 10214, 2020 |
| Blocks & Files interview (15 April 2020, one university network manager) | WD40EFAX, ZFS resilver | About 100 MB/s for about 40 min, then the drives 'die'; waiting about 1 h restarted the cycle; ZFS reported dozens of delays over 60 s and one pause of 3 min | Blocks & Files, 2020 |
| ServeTheHome test (28 May 2020) | Two WD40EFAX drives, firmware 82.00A82, 4-drive RAIDZ resilver | Finished in nearly 230 h with no errors; the failed state was not reproduced | ServeTheHome, 2020 |
| Ars Technica test (5 June 2020) | One WD40EFAX, 8-disk mdraid RAID6 rebuild, array 75% full | Rebuilt without trouble, with the same time whether the drive was new or already full | Ars Technica, 2020 |
| OpenZFS issue 10214 (commenter, 4 June and 6 August 2020) | 4 WD30EFRX drives, which the commenter checked were not SMR, in a RAID10 on a SATA controller, later an LSI HBA | Same IDNF and 'out of range' log text on 1 to 4 drives at once; none for 9 days after switching the SATA port from AHCI to IDE mode, then again after moving to the HBA; cause not found | OpenZFS issue 10214, 2020 |
Reading the numbers
The measured behaviour is a cliff, not a decline. Ars's 32 KiB test, ServeTheHome's resilver and the University of Minnesota benchmark all show a drive that works at normal speed until a conventional write area is used up, and then collapses. The Minnesota authors found that each cleaning pass dropped throughput from over 100 MB/s to about 0.1 MB/s, and that a workload spread over the whole drive made the first cleaning last so long that their test program broke after about 100 min; they guess the drive throttled writes long enough to cause timeouts in the Linux stack. That paper is about a prototype host-aware drive, so it supports the mechanism, not the size of the effect on a retail WD Red. Western Digital's own post says the same mechanism in plain words: sustained random writes during ZFS resilvering leave no idle time for the drive's background work.
The logs show the errors only after the cliff has turned into failed commands, and the usual health signals stay quiet. In the one published log set, SMART overall health said PASSED, and reallocated, pending, uncorrectable and CRC counters were all 0 at the same time as an ATA error count of 7,487 (the latest 24 shown were IDNF). A health check built on those four attributes would have reported a healthy drive. What did record the event was the ATA extended error log, the kernel's 'logical block address out of range' sense text and the filesystem's own write-error counter. None of the opened sources shows a SMART attribute that tracks the slow phase before errors start; timing seen from the host (peak latency, pauses over 60 s) is the only early sign in the evidence, and no source gives a threshold for it.
The error text does not identify SMR by itself. The same IDNF and 'logical block address out of range' lines appeared in a report on four drives that the commenter checked were not SMR, ; those errors began after the pool grew from 2 to 4 drives, paused for 9 days after an AHCI to IDE change and returned on a different controller, and the commenters suggested controller or power-supply causes. The ordinary link counters (CRC errors, PHY ready drops) are separate evidence and, in the SMR log, showed 0 CRC errors. The two readings can only be told apart by the drive model, firmware and the write pattern at the time, which the logs of the CMR report do not settle.
What is unknown is larger than what is measured. Three testers used the same model and firmware: the iXsystems lab and one user saw the failed state, while ServeTheHome and Ars saw slowness only. The differences in workload (replacement into a degraded vdev, array size, other load, prior writes) are described but not isolated by any source. Western Digital says non-ZFS rebuilds on Synology or QNAP systems take as long as CMR or slightly longer, and Ars's mdraid result is consistent with that for one drive. Whether the failed state is a firmware bug or the drive's normal behaviour under an unusual access order is argued by the interviewee and commenters but not shown by the vendor. No source opened here gives a failure rate for SMR drives, so this page does not claim SMR drives fail more often; it documents what an overloaded one looks like in logs.
Related: HDD mechanical failure points, SCSI sense and kernel log signatures, SATA and SAS link failures and the SMART failure-signal note.
Method
Everything here was read on 2026-10-11 from pages and documents opened that day: the iXsystems/TrueNAS notice on WD Red SMR drives and ZFS; the OpenZFS GitHub issue 10214 (opened 16 April 2020, closed 22 December 2020) with its pasted smartctl and kernel logs; Ars Technica's test of 5 June 2020; ServeTheHome's test of 28 May 2020 (both pages); the Blocks & Files interview of 15 April 2020; Western Digital's blog post 'On WD Red NAS Drives' (entries of 20 April, 22 April and 23 June 2020); and the PDF 'Performance Evaluation of Host Aware Shingled Magnetic Recording (HA-SMR) Drives' by Wu, Fan, Yang, Zhang, Ge and Du, hosted on a University of Minnesota course page (no publication date appears in the opened text). Numbers are copied from those texts. Ratios marked derived are my own division of two published figures, with the rounding shown. A claim is only stated as fact where the source reports it as a measurement or a log line; beliefs and guesses of interviewees and commenters are labelled as such.
Limits
All measurements come from one drive family (WD Red 2 to 6 TB EFAX, firmware 82.00A82 where stated) tested in 2020, except the University of Minnesota paper, which measured a prototype Seagate 8 TB host-aware SMR drive (ST8000AS0022, prototype firmware SN03), not a drive-managed retail drive. The page therefore cannot say how other drive-managed SMR models behave and does not rank SMR against CMR on failure rate: no source opened here gives a failure rate for SMR drives. The failed-state evidence is thin: one lab confirmation (iXsystems), one published log set from one user, and anecdotes from an interviewee and commenters; iXsystems itself calls the event rare, with one reported issue among about one hundred FreeNAS Mini systems shipped with 2 TB or 6 TB DM-SMR drives. ServeTheHome's two tested drives and Ars's mdraid rebuild did not reach the failed state, so the cause of the errors is not settled: the interviewee and an issue commenter call it a Western Digital firmware bug, one commenter believes IDNF is a normal SMR event that the drive should mask, and Western Digital's posts do not mention IDNF at all. The OpenZFS maintainer closed the issue saying the reports sounded hardware-caused. The conventional-cache size of the WD drives is not published; the 'few tens of GB up to 100 GB' figure is one interviewee's inference. The SMART excerpt is from one drive at 540 power-on hours, so it shows what that log looked like, not what every SMR drive reports. No source gives a latency threshold that separates cache exhaustion from a failing drive.
Sources
- 01iXsystems / TrueNAS Documentation Hub: WD Red SMR Drive Compatibility with ZFS (page last modified 2024-08-27; original notice date not shown) · accessed 2026-10-11
- 02OpenZFS GitHub issue 10214: WD WDx0EFAX drive unable to resilver (opened 2020-04-16, closed 2020-12-22) · accessed 2026-10-11
- 03Ars Technica: We put Western Digital's dreaded SMR Red drive to the test (5 June 2020) · accessed 2026-10-11
- 04ServeTheHome: WD Red SMR vs CMR Tested Avoid Red SMR (28 May 2020, pages 1 and 2) · accessed 2026-10-11
- 05Blocks & Files: Shingled hard drives have non-shingled zones for caching writes (15 April 2020) · accessed 2026-10-11
- 06Western Digital blog: On WD Red NAS Drives (entries of 20 April, 22 April and 23 June 2020) · accessed 2026-10-11
- 07Wu, Fan, Yang, Zhang, Ge, Du: Performance Evaluation of Host Aware Shingled Magnetic Recording (HA-SMR) Drives (PDF on a University of Minnesota course page; publication date not shown in the opened text) · accessed 2026-10-11