How likely is a read error during a RAID rebuild, and why do the datasheet math and the field data disagree?
Published 2026-10-11
Taken literally, the datasheet rate says a rebuild that reads five 12 TB disks (60 TB, 4.8e14 bits) at the WD Red Plus limit of fewer than 1 error in 10^14 bits would expect up to 4.8 unrecoverable read errors, a 99.2% chance of at least one (derived, Poisson). The same read volume at the Seagate Exos X20 rating of 1 sector per 10^15 bits is 0.8 expected errors, a 55.1% chance for five 20 TB survivors (derived). The field data say the real picture is lumpier: in a NetApp study of 1.53 million disks over 32 months only 3.45% of disks ever developed a latent sector error, but a disk that had one was far more likely to get more (0.671 against 0.018 for a disk without one), so errors cluster on a few disks rather than arriving at a steady per-bit rate. A single unrecoverable read during a single-parity rebuild is what costs data; a double-parity layout such as RAIDZ2 can absorb one per stripe, and OpenZFS cannot rebuild RAIDZ sequentially, so a RAIDZ resilver walks the block tree and checks checksums as it goes.
- Question
- What does the datasheet unrecoverable-read-error rate imply for a RAID or RAIDZ rebuild, what do two field studies of latent sector errors say about that model, and what does OpenZFS actually read and verify when it resilvers?
- Evidence types
- Two vendor data sheets, two peer-reviewed field studies from the same NetApp data, one OpenZFS documentation page and one OpenZFS source file
- Out of scope
- SSD rebuilds, RAID write hole (covered elsewhere), ZFS checksum counters (see the filesystem checksum note), and any attempt to recover data from other people's drives
- Access date
- All sources opened on 2026-10-11
What two data sheets promise
| Drive family (data sheet date) | Printed rate (errors per bits read) | Capacity range (TB) | Read volume per expected error (TB, derived) | Source |
|---|---|---|---|---|
| WD Red Plus, CMR, 3.5-inch NAS (September 2025) | <1 in 10^14 | 2 to 12 | 12.5 at the limit; a drive better than the limit reads more per error, but the sheet gives only the bound | WD Red Plus data sheet |
| Seagate Exos X20, CMR, helium, 3.5-inch (November 2021) | 1 sector per 10E15 | 18 and 20 | 125 | Seagate Exos X20 data sheet |
| WD Red Plus rated workload | 180 TB/year | 2 to 12 | Up to 14.4 expected errors per year at the full rated workload if the rate sits at the limit (derived: 180 / 12.5) | WD Red Plus data sheet |
| WD Red Plus, one full read of a 12 TB disk | <1 in 10^14 | 12 | up to 0.96 expected errors per full read at the limit (derived: 12 / 12.5) | WD Red Plus data sheet |
| Seagate Exos X20, one full read of a 20 TB disk | 1 sector per 10E15 | 20 | 0.16 expected sectors per full read (derived: 20 / 125) | Seagate Exos X20 data sheet |
What the spec implies for one rebuild, if errors were independent
| Scenario | Survivors read (disks x TB) | Bits read | Expected errors (count) | P(at least one) (%) |
|---|---|---|---|---|
| 6-disk array, 12 TB disks, rate 1 in 10^14 | 5 x 12 = 60 TB | 4.8e14 | 4.80 | 99.2 |
| 4-disk array, 12 TB disks, rate 1 in 10^14 | 3 x 12 = 36 TB | 2.88e14 | 2.88 | 94.4 |
| 6-disk array, 20 TB disks, rate 1 in 10^15 | 5 x 20 = 100 TB | 8.0e14 | 0.80 | 55.1 |
| 4-disk array, 20 TB disks, rate 1 in 10^15 | 3 x 20 = 60 TB | 4.8e14 | 0.48 | 38.1 |
| 6-disk array of 12 TB disks, only half the capacity is read (an assumed half-full pool), rate 1 in 10^14 | 2.5 x 12 = 30 TB | 2.4e14 | 2.40 | 90.9 |
What 1.53 million disks actually did
| Measurement | Value | Unit | Note in the source | Source |
|---|---|---|---|---|
| Disks that developed at least one latent sector error in 32 months | 3.45 (53,820 of 1.53 million) | % of disks | Nearline 8.5%, enterprise 1.9% | Bairavasundaram et al., SIGMETRICS 2007 |
| Median / mean errors per disk that had any | 3 / 19.7 | errors per error disk | Mode is 1 error (30% of error disks); 0.2% of error disks had more than 1,000 and were left out of the mean | Bairavasundaram et al., SIGMETRICS 2007 |
| P(another error by month 18 | an error by month 12) against the unconditional chance | 0.671 against 0.018 | probability (unitless) | Ratio about 37 (derived); errors are not independent | Bairavasundaram et al., SIGMETRICS 2007 |
| Errors found by media scrubbing | more than 60 | % of errors | Media scrubs typically completed within 2 weeks; verify operations found 86.6% of nearline and 61.5% of enterprise errors | Bairavasundaram et al., SIGMETRICS 2007 |
| Error bursts that are a single sector | 90 to 98 | % of bursts | Geometric model deviated more than 13 times more than a Pareto fit (chi-squared) | Schroeder et al., FAST 2010 |
| Drives whose errors all fall in one 2-week window | 55 to 85 | % of drives | Poisson was a poor fit; Pareto fit well | Schroeder et al., FAST 2010 |
| Chance at least one of 5 survivors ever had an error (derived) | 35.9 (nearline), 9.1 (enterprise) | % over 32 months | 1 - (1 - 0.085)^5 and 1 - (1 - 0.019)^5, assuming independent disks; contrast only, not a rebuild probability | Bairavasundaram et al., SIGMETRICS 2007 |
What an OpenZFS resilver reads and checks
| Behaviour | Value or statement | Applies to | Consequence stated | Source |
|---|---|---|---|---|
| Healing resilver | Reads blocks via the block tree and verifies each checksum as it reads | mirror, raidz, draid | Repairs on the spot, but the I/O pattern is random, which 'substantially increases the time required' | OpenZFS vdev_rebuild.c |
| Sequential reconstruction | Rebuilds in LBA order, no checksum verification; a scrub then starts automatically | mirror and draid only | Restores redundancy faster; 'not supported for raidz' because RAIDZ has variable stripe width | OpenZFS: Scrub and Resilver |
| Sequential read segment and in-flight limit | 1 MiB per data disk; 64 MiB per leaf vdev | sequential rebuild | 64 MiB was observed to give the best performance | OpenZFS vdev_rebuild.c |
| Measured rebuild rate quoted in a code comment | 1.2 GB/s to a distributed spare on a 106-drive draid2:11d:106c HDD pool | draid | One test, one pool; a comment, not a benchmark report | OpenZFS vdev_rebuild.c |
| Scrub frequency suggested | Monthly for consumer disks, quarterly for enterprise | all | Described as a common baseline, not a measured optimum | OpenZFS: Scrub and Resilver |
Reading the numbers
The answer. Read literally, the printed rates make a full-capacity rebuild of a large single-parity array look close to hopeless at 1 in 10^14 (99.2% for five 12 TB survivors, derived) and a coin toss at 1 in 10^15 (55.1% for five 20 TB survivors, derived). Those numbers come from the spec treated as a steady per-bit rate. The two field studies opened here show that real latent errors are not steady: they hit few disks, and then hit them in bursts.
What the field data change. In the 2007 study 3.45% of disks developed any latent sector error in 32 months, yet a disk with one error was 37 times as likely (0.671 against 0.018, derived ratio) to get another soon after. The 2010 analysis of the same data found that 55 to 85% of affected drives had all their errors in one 2-week window and that a geometric or Poisson model fit poorly. A per-bit Poisson model therefore tends to spread errors thinly over every disk, whereas the field data put many errors on a few disks, so the datasheet math says little about whether the next rebuild will meet one. It does not say the risk is small: the study counted nearline disks at 8.5% over 32 months, and a read error during a single-parity rebuild is a data-loss event.
Why this does not simply cancel out. Both papers also show what helps. More than 60% of errors in the NetApp systems were found by media scrubbing, so an error that scrubbing found and repaired before a disk failed is not there to hit the rebuild; a disk that has recently logged errors is the one to treat carefully, since errors arrive within a month of the first for 54.8% of nearline and 62.0% of enterprise error disks. The 2007 paper's own repair suggestion was to accelerate rebuilds when the survivors are over a year old or have logged an error in the last 1,000 minutes.
What ZFS adds, and what it cannot. A RAIDZ resilver is a healing resilver: the OpenZFS documentation says sequential reconstruction is not supported for raidz, so the resilver verifies each block's checksum as it reads and repairs on the spot, at the cost of a slower, more random I/O pattern. Redundancy is not restored until it finishes, so the risk window is longer than a sequential rebuild would be. A mirror or dRAID can rebuild sequentially, faster, but without checksum verification until the automatic scrub completes. Whether a given pool would meet an unrecoverable read at all depends on how much of it holds data: the source comment says sequential reconstruction resilvers only allocated capacity, so the 90.9% row in the math table is an assumed half-full case, not a property of any particular pool.
What not to conclude. None of the sources opened here measures how often rebuilds fail. A pessimistic reading of the spec (nearly every large single-parity rebuild hits an error) is not supported by the field data, and an optimistic reading (3.45% means a rebuild is nearly always safe) is not supported either, because the rate is for a 32-month lifetime on old drives and depends strongly on model, age and capacity: the 2007 paper reports the fraction of affected disks rising with capacity within a family, yet no consistent rise in errors per gigabyte. Datasheet and field figures are different quantities; compare them only with those definitions in hand.
Related: the SMART failure-signal note, SCSI sense and kernel log signatures SATA and SAS link failure signatures and filesystem checksum error detection.
Method
Everything here was read on 2026-10-11 from documents opened that day: the Seagate Exos X20 data sheet DS2080-2111US (November 2021); the Western Digital WD Red Plus data sheet (September 2025); the SIGMETRICS 2007 paper 'An Analysis of Latent Sector Errors in Disk Drives' (Bairavasundaram, Goodson, Pasupathy, Schindler), PDF from the University of Wisconsin ADSL; the FAST 2010 paper 'Understanding latent sector errors and how to protect against them' (Schroeder, Damouras, Gill), PDF from USENIX; the OpenZFS documentation page 'Scrub and Resilver'; and the OpenZFS source file module/zfs/vdev_rebuild.c (master). Numbers in the 'derived' rows are my own arithmetic: bits read = survivors x capacity (TB, decimal, as the vendors define it) x 10^12 x 8; expected errors = bits read x the datasheet rate; probability of at least one error = 1 - exp(-expected errors), which assumes errors arrive independently at a fixed per-bit rate (a Poisson model). That assumption is exactly what the two field papers test and reject, so the derived probabilities are what the spec would imply, not predictions. Where a figure comes from a paper or datasheet it is quoted with its own unit.
Limits
The datasheet rates are specification limits, not measurements: WD prints '<1 in 10^14' and Seagate prints '1 sector per 10E15', so the true rate of any one drive is not given and may be much lower. I found no source opened here that measures the rate of unrecoverable reads during actual RAID rebuilds, so the rebuild probabilities above are model outputs, not observed frequencies. The field studies are old and narrow: NetApp systems from January 2004 for 32 months, nearline ATA and enterprise Fibre Channel disks of that era, with nearline write errors hidden by automatic remapping (the paper says its rates are lower bounds); no source opened here gives latent-error rates for current 12 to 20 TB helium or CMR drives. The 35.9% and 9.1% figures in the field table assume the five survivors are independent and apply a 32-month, whole-life rate to a single rebuild window, which overstates a single rebuild (errors found by scrubbing earlier would have been repaired) and ignores that errors are correlated across disks of one batch; they are an order-of-magnitude contrast, not a rebuild probability. I did not model RAID6 or RAIDZ2 data-loss probability: that needs the chance of two bad sectors in the same stripe, which neither paper gives. The OpenZFS pages describe behaviour of the current master and documentation, not of older releases; I did not open the RAIDZ documentation page, so nothing on RAIDZ geometry or expansion is taken from it. Nothing here covers SSD rebuilds, erasure-coded clouds, or any recovery of other people's media.
Sources
- 01Seagate Exos X20 data sheet DS2080-2111US, November 2021 · accessed 2026-10-11
- 02Western Digital WD Red Plus data sheet, September 2025 · accessed 2026-10-11
- 03Bairavasundaram, Goodson, Pasupathy, Schindler: An Analysis of Latent Sector Errors in Disk Drives, SIGMETRICS 2007 · accessed 2026-10-11
- 04Schroeder, Damouras, Gill: Understanding latent sector errors and how to protect against them, FAST 2010 · accessed 2026-10-11
- 05OpenZFS documentation: Scrub and Resilver · accessed 2026-10-11
- 06OpenZFS source: module/zfs/vdev_rebuild.c (master) · accessed 2026-10-11