Skip to content

Stat analysis

What does a shingled (SMR) hard drive failure look like in logs, and what does SMART miss?

Published 2026-10-11

The public SMR failures are mostly not worn-out media but a write-cache overflow: a drive-managed SMR drive absorbs writes in a small conventional area and, once that fills, slows by one to three orders of magnitude, and in some reports it then returns errors and is dropped from the array. Measured slowdowns: 13.2 MiB/s against 209.3 MiB/s for 32 KiB writes on a 4 TB WD Red EFAX (Ars Technica, 15.9:1), nearly 230 h against under 17 h for a 4-drive RAIDZ resilver (ServeTheHome), and about 0.1 MB/s for more than 25 min in a prototype host-aware SMR drive (University of Minnesota paper). In the one failing resilver whose logs were published (OpenZFS issue 10214) SMART overall health said PASSED and the reallocated, pending, uncorrectable and CRC counters were all 0, while the ATA error log counted 7,487 errors (the 24 it still showed were all IDNF) and ZFS counted 47.7 K write errors. Only one lab (iXsystems) and some users saw the failed state; ServeTheHome and Ars, testing the same drive model, saw slowness only, and a CMR drive on a shared controller logged the same error text.

Question
When a drive-managed SMR disk is overloaded with writes, which public signals (SMART, ATA error log, kernel log, filesystem counters, timing) show it, and which show nothing?
Drives
WD Red 4 TB WD40EFAX (DM-SMR, firmware 82.00A82) in four independent tests or reports; one prototype Seagate 8 TB host-aware SMR drive in a university benchmark.
Failure layers
Persistent-cache exhaustion and blocking cleaning; ATA IDNF errors and SCSI 'logical block address out of range'; array software dropping a slow or erroring member.
Not covered
SMR failure rates, a ranking of SMR against CMR, host-managed zoned drives in production, and any recovery of data from affected drives.

How far a drive-managed SMR drive slows once its cache is full

Published figures with units. 'Derived' ratios are my own division of two published numbers. MiB/s = mebibytes per second, MB/s = megabytes per second, h = hours, s = seconds, ms = milliseconds.
MeasurementSMR drive resultComparison resultRatioSource
32 KiB sequential writes, imitating one ZFS RAIDZ member (WD Red 4 TB EFAX vs Seagate IronWolf 12 TB)13.2 MiB/s209.3 MiB/s15.9:1 slower (as published)Ars Technica, 2020
4-drive RAIDZ resilver of an array about 60% full, with network load (WD40EFAX vs three CMR drives)nearly 230 h (9 days 14 h)under 17 h for each CMR drivemore than 13:1 slower (derived, 230 / 17)ServeTheHome, 2020
Peak latency of one 1 MiB random write (WD Red 4 TB EFAX vs IronWolf)1.3 s108 msabout 12:1 slower (derived, 1,300 / 108)Ars Technica, 2020
ZFS resilver of a WD40EFAX with write cache and look-ahead switched off (one user, via interview)8 days at under 6 MB/s24 h, which the user calls the more usual timeabout 8:1 slower (derived, 8 days / 24 h)Blocks & Files, 2020
Blocking cache cleaning in a prototype host-aware SMR drive, 256 KB non-sequential writes for 2 habout 0.1 MB/s for more than 25 min (128 zones) and more than 37 min (256 zones)over 100 MB/s before cleaning startsabout 1,000:1 slower (derived, 100 / 0.1)Wu et al., HA-SMR evaluation

What each signal showed during one failing ZFS resilver

All values are from the logs pasted into OpenZFS issue 10214 for one WD40EFAX (firmware 82.00A82, 540 power-on hours, 22 days 12 h of power-on time). The resilver replaced a disk in a 5-disk raidz1 and stalled at 27.61% done. Counts are dimensionless event counts.
SignalValue in the reportWhat it showsSource
SMART overall-health self-assessmentPASSEDNo warning while the drive was failing the resilverOpenZFS issue 10214, 2020
SMART attributes 5, 197, 198, 199 (raw values, count)0, 0, 0, 0 (reallocated sectors, pending sectors, offline uncorrectable, UDMA CRC errors)No media or link-integrity counter movedOpenZFS issue 10214, 2020
ATA Extended Comprehensive SMART Error Log (count)Device Error Count 7,487; shown errors are 'IDNF at LBA' on WRITE FPDMA QUEUED commands; the log keeps only the latest 24The only drive-side record of the failure; a commenter says it needs smartctl -x, because -a does not print the extended logOpenZFS issue 10214, 2020
ZFS zpool status WRITE column (count)47.7 K write errors on the new disk at 27.61% resilvered, 4.69 TB issued of 17.0 TBFilesystem-level view of the same failure; the poster later suspects the counter wraps at about 65.5 KOpenZFS issue 10214, 2020
Linux kernel log, SCSI senseSense Key: Illegal Request; Add. Sense: Logical block address out of range; Write(16) commands; 'critical target error'The host-side translation of the drive's IDNF replyOpenZFS issue 10214, 2020
SATA PHY event counters (count)CRC-error counters 0; PhyRdy to PhyNRdy transitions 82No CRC errors on the link, but 82 link-ready drops were recorded in that logOpenZFS issue 10214, 2020

Who saw a failed state, and who saw only slowness

Each row is one report; none is a fleet statistic. Durations and counts are as published.
ReportDrive and workloadOutcomeSource
iXsystems lab test (2020)WD40EFAX, firmware 82.00A82, heavy write loads including resilveringDrive entered a faulty state, returned IDNF, became unusable and was treated as failed by ZFS; about 100 Minis shipped with 2 TB or 6 TB DM-SMR drives and one issue reportediXsystems / TrueNAS notice, 2020
OpenZFS issue 10214 (user, 16 April 2020)WD40EFAX, zpool replace into a 5-disk raidz1Stalled at 27.61% with IOPS falling to zero and 47.7 K write errors; the same disk later resilvered 7.26 G in 4 min 24 s with 2 write errors after a sequential copyOpenZFS issue 10214, 2020
Blocks & Files interview (15 April 2020, one university network manager)WD40EFAX, ZFS resilverAbout 100 MB/s for about 40 min, then the drives 'die'; waiting about 1 h restarted the cycle; ZFS reported dozens of delays over 60 s and one pause of 3 minBlocks & Files, 2020
ServeTheHome test (28 May 2020)Two WD40EFAX drives, firmware 82.00A82, 4-drive RAIDZ resilverFinished in nearly 230 h with no errors; the failed state was not reproducedServeTheHome, 2020
Ars Technica test (5 June 2020)One WD40EFAX, 8-disk mdraid RAID6 rebuild, array 75% fullRebuilt without trouble, with the same time whether the drive was new or already fullArs Technica, 2020
OpenZFS issue 10214 (commenter, 4 June and 6 August 2020)4 WD30EFRX drives, which the commenter checked were not SMR, in a RAID10 on a SATA controller, later an LSI HBASame IDNF and 'out of range' log text on 1 to 4 drives at once; none for 9 days after switching the SATA port from AHCI to IDE mode, then again after moving to the HBA; cause not foundOpenZFS issue 10214, 2020

Reading the numbers

The measured behaviour is a cliff, not a decline. Ars's 32 KiB test, ServeTheHome's resilver and the University of Minnesota benchmark all show a drive that works at normal speed until a conventional write area is used up, and then collapses. The Minnesota authors found that each cleaning pass dropped throughput from over 100 MB/s to about 0.1 MB/s, and that a workload spread over the whole drive made the first cleaning last so long that their test program broke after about 100 min; they guess the drive throttled writes long enough to cause timeouts in the Linux stack. That paper is about a prototype host-aware drive, so it supports the mechanism, not the size of the effect on a retail WD Red. Western Digital's own post says the same mechanism in plain words: sustained random writes during ZFS resilvering leave no idle time for the drive's background work.

The logs show the errors only after the cliff has turned into failed commands, and the usual health signals stay quiet. In the one published log set, SMART overall health said PASSED, and reallocated, pending, uncorrectable and CRC counters were all 0 at the same time as an ATA error count of 7,487 (the latest 24 shown were IDNF). A health check built on those four attributes would have reported a healthy drive. What did record the event was the ATA extended error log, the kernel's 'logical block address out of range' sense text and the filesystem's own write-error counter. None of the opened sources shows a SMART attribute that tracks the slow phase before errors start; timing seen from the host (peak latency, pauses over 60 s) is the only early sign in the evidence, and no source gives a threshold for it.

The error text does not identify SMR by itself. The same IDNF and 'logical block address out of range' lines appeared in a report on four drives that the commenter checked were not SMR, ; those errors began after the pool grew from 2 to 4 drives, paused for 9 days after an AHCI to IDE change and returned on a different controller, and the commenters suggested controller or power-supply causes. The ordinary link counters (CRC errors, PHY ready drops) are separate evidence and, in the SMR log, showed 0 CRC errors. The two readings can only be told apart by the drive model, firmware and the write pattern at the time, which the logs of the CMR report do not settle.

What is unknown is larger than what is measured. Three testers used the same model and firmware: the iXsystems lab and one user saw the failed state, while ServeTheHome and Ars saw slowness only. The differences in workload (replacement into a degraded vdev, array size, other load, prior writes) are described but not isolated by any source. Western Digital says non-ZFS rebuilds on Synology or QNAP systems take as long as CMR or slightly longer, and Ars's mdraid result is consistent with that for one drive. Whether the failed state is a firmware bug or the drive's normal behaviour under an unusual access order is argued by the interviewee and commenters but not shown by the vendor. No source opened here gives a failure rate for SMR drives, so this page does not claim SMR drives fail more often; it documents what an overloaded one looks like in logs.

Related: HDD mechanical failure points, SCSI sense and kernel log signatures, SATA and SAS link failures and the SMART failure-signal note.

Method

Everything here was read on 2026-10-11 from pages and documents opened that day: the iXsystems/TrueNAS notice on WD Red SMR drives and ZFS; the OpenZFS GitHub issue 10214 (opened 16 April 2020, closed 22 December 2020) with its pasted smartctl and kernel logs; Ars Technica's test of 5 June 2020; ServeTheHome's test of 28 May 2020 (both pages); the Blocks & Files interview of 15 April 2020; Western Digital's blog post 'On WD Red NAS Drives' (entries of 20 April, 22 April and 23 June 2020); and the PDF 'Performance Evaluation of Host Aware Shingled Magnetic Recording (HA-SMR) Drives' by Wu, Fan, Yang, Zhang, Ge and Du, hosted on a University of Minnesota course page (no publication date appears in the opened text). Numbers are copied from those texts. Ratios marked derived are my own division of two published figures, with the rounding shown. A claim is only stated as fact where the source reports it as a measurement or a log line; beliefs and guesses of interviewees and commenters are labelled as such.

Limits

All measurements come from one drive family (WD Red 2 to 6 TB EFAX, firmware 82.00A82 where stated) tested in 2020, except the University of Minnesota paper, which measured a prototype Seagate 8 TB host-aware SMR drive (ST8000AS0022, prototype firmware SN03), not a drive-managed retail drive. The page therefore cannot say how other drive-managed SMR models behave and does not rank SMR against CMR on failure rate: no source opened here gives a failure rate for SMR drives. The failed-state evidence is thin: one lab confirmation (iXsystems), one published log set from one user, and anecdotes from an interviewee and commenters; iXsystems itself calls the event rare, with one reported issue among about one hundred FreeNAS Mini systems shipped with 2 TB or 6 TB DM-SMR drives. ServeTheHome's two tested drives and Ars's mdraid rebuild did not reach the failed state, so the cause of the errors is not settled: the interviewee and an issue commenter call it a Western Digital firmware bug, one commenter believes IDNF is a normal SMR event that the drive should mask, and Western Digital's posts do not mention IDNF at all. The OpenZFS maintainer closed the issue saying the reports sounded hardware-caused. The conventional-cache size of the WD drives is not published; the 'few tens of GB up to 100 GB' figure is one interviewee's inference. The SMART excerpt is from one drive at 540 power-on hours, so it shows what that log looked like, not what every SMR drive reports. No source gives a latency threshold that separates cache exhaustion from a failing drive.

Sources

  1. 01iXsystems / TrueNAS Documentation Hub: WD Red SMR Drive Compatibility with ZFS (page last modified 2024-08-27; original notice date not shown) · accessed 2026-10-11
  2. 02OpenZFS GitHub issue 10214: WD WDx0EFAX drive unable to resilver (opened 2020-04-16, closed 2020-12-22) · accessed 2026-10-11
  3. 03Ars Technica: We put Western Digital's dreaded SMR Red drive to the test (5 June 2020) · accessed 2026-10-11
  4. 04ServeTheHome: WD Red SMR vs CMR Tested Avoid Red SMR (28 May 2020, pages 1 and 2) · accessed 2026-10-11
  5. 05Blocks & Files: Shingled hard drives have non-shingled zones for caching writes (15 April 2020) · accessed 2026-10-11
  6. 06Western Digital blog: On WD Red NAS Drives (entries of 20 April, 22 April and 23 June 2020) · accessed 2026-10-11
  7. 07Wu, Fan, Yang, Zhang, Ge, Du: Performance Evaluation of Host Aware Shingled Magnetic Recording (HA-SMR) Drives (PDF on a University of Minnesota course page; publication date not shown in the opened text) · accessed 2026-10-11