0.86% of nearline disks and 0.065% of enterprise disks developed a checksum mismatch in 41 months
Published 2026-10-10
In Bairavasundaram et al. (FAST 2008), 3,088 of about 358,000 nearline disks (0.86%) and 767 of 1.17 million enterprise-class disks (0.065%) developed at least one silent checksum mismatch over 41 months, in field logs covering 1.53 million disks. That is about 400,000 mismatched blocks in all, but the burden is lopsided: the median affected disk had 3 mismatches, the mean was 104, and the worst disk had 33,000. The paper's rates are per disk and come from one vendor's support logs, so they do not say how often a file is corrupted. An independent 2007 CERN report found 22 mismatching files among 33,700 checked, about 1 in 1,500, using a different method and a different population.
- Population
- 1.53 million disks: about 358,000 nearline (SATA behind adapters) and 1.17 million enterprise-class (Fibre Channel)
- Window
- 41 months starting January 2004; disks shipped after January 2004; 14 disk families, 31 models
- Source of data
- Support logs from tens of thousands of production and development storage systems at hundreds of customer sites
- Measure
- Disks with at least one checksum mismatch: a 4 KB block that fails its stored checksum when read. Detected corruption only
- Cross-check
- CERN IT, January and February 2007: pattern probes on more than 3,000 nodes, RAID verification, and file checksum comparison
Numbers
| Measure | Value | Unit | Population or scope | Source |
|---|---|---|---|---|
| Disks in the sample | 1.53 million (about 358,000 NL; 1.17 million ES) | disks | 41 months from January 2004; 14 families; 31 models | FAST 2008 (PDF) |
| Checksum mismatches observed | about 400,000 | 4 KB blocks | Whole sample, 41 months | FAST 2008 (HTML) |
| Nearline disks with at least one mismatch | 0.86 (3,088 disks) | percent of disks | About 358,000 nearline disks, 41 months | FAST 2008 (PDF) |
| Enterprise disks with at least one mismatch | 0.065 (767 disks) | percent of disks | 1.17 million enterprise disks, 41 months | FAST 2008 (PDF) |
| First 17 months in the field, NL and ES | 0.66 and 0.06 | percent of disks | Weighted class averages | FAST 2008 (PDF) |
| Mismatches per corrupt disk: median, mean, maximum | 3, 104, 33,000 | blocks per disk | 3,855 corrupt disks, 41 months | FAST 2008 (PDF) |
| Share of all mismatches from the top 1% of corrupt disks | more than 50 | percent of mismatches | Corrupt disks only | FAST 2008 (HTML) |
| Chance of more mismatches after one, NL | about 0.6 versus 0.0066 for a first | probability, unitless | First 17 months in the field | FAST 2008 (HTML) |
| Corrupt blocks with a corrupt neighbor block, NL | more than 50 | percent of corrupt blocks | Disks with 2 to 10 mismatches | FAST 2008 (HTML) |
| Found by scrubs, NL and ES | about 49 and 73 | percent of mismatches | Weighted class averages | FAST 2008 (HTML) |
| Found during RAID reconstruction, NL | about 8 | percent of mismatches | Weighted class average | FAST 2008 (HTML) |
| Disks affected per year, checksum mismatches versus latent sector errors, NL | 0.466 versus 9.5 | percent of disks per year | Table 2 averages | FAST 2008 (HTML) |
| Disks affected per year, checksum mismatches versus latent sector errors, ES | 0.042 versus 1.4 | percent of disks per year | Table 2 averages | FAST 2008 (HTML) |
| Disks with identity discrepancies | 365 | disks | 1.53 million disks, whole study | FAST 2008 (PDF) |
| Files failing a checksum comparison (CERN) | 22 of 33,700, about 1 in 1,500 | files; about 8.7 TB checked | CERN disk pool, early 2007, checked against tape checksums | CERN 2007 |
| RAID verification block problems versus vendor-rate expectation (CERN) | about 300 versus about 850 | blocks in 4 weeks | 492 systems, about 1.5 PB | CERN 2007 |
| Pattern probe errors (CERN) | 500 errors on 100 nodes | errors in about 5 weeks | More than 3,000 nodes, probe every 2 hours | CERN 2007 |
This table is a short extract of printed figures, not a copy of the papers and not the underlying logs. Rows 1 to 14 are from Bairavasundaram et al. (HTML and PDF); rows 15 to 17 are from the CERN report and are a different measure. Rates marked per disk are shares of disks, not shares of blocks or files.
Method
The figures are copied from the USENIX HTML and PDF versions of Bairavasundaram, Goodson, Schroeder, Arpaci-Dusseau, and Arpaci-Dusseau, FAST 2008, not refit. The HTML conversion drops several numerals, so each of them was checked against the PDF. The authors mined the support logs of tens of thousands of production and development storage systems at hundreds of customer sites, for 41 months starting in January 2004, and kept only disks shipped after January 2004. The sample is 1.53 million disks in 14 families and 31 models. A checksum mismatch is a 4 KB block whose stored checksum fails when the RAID layer reads it, from any read, scrub, or rebuild. The paper separates this from identity discrepancies (found by file system block identity checks, from lost or misdirected writes) and parity inconsistencies (found only by scrubs). A corrupt disk is a disk with at least one corrupt block. The independent cross-check is the CERN IT report 'Data integrity' (Panzer-Steindel, draft 1.3, 8 April 2007), which ran its own probes in January and February 2007. It reports counts, not a fleet-wide rate per disk, and does not test the NetApp percentages.
Limits
Only corruption that the system detected is counted. Corruption nobody reads, and nobody scrubs, is invisible. The authors cannot verify that every disk was scrubbed, and a RAID group was scrubbed about once every two weeks on average. Not all customers enable logging, and some enable it only after a period of use. Disks that were removed are counted only up to removal, so error-prone disks may drop out of view. The paper says the counts mix causes: bit-level corruption, torn writes, and misdirected writes, and it names faulty adapters or shelf controllers for three disk models, so 'disk' here includes its adapter path, not only the platters. Workload data are coarse per-system weekly counts, so the finding of no correlation (below 0.1) may hide a per-drive effect. Disk models are anonymized. Nearline and enterprise disks use different on-disk formats (512-byte sectors with a segment after every eight, versus 520-byte sectors). The paper prints a mean of 104 mismatches per corrupt disk for the whole 41-month sample and a mean of 78 in its first-17-months analysis, and the text does not reconcile the two beyond the different windows. The CERN report is a draft written for internal discussion, from one site, with probe programs it designed itself, not peer-reviewed. Its expected-error comparison relies on a vendor bit error rate of 1 in 10 to the power 14 bits. Its counts are not comparable to per-disk fractions. Several of the NetApp results are shown only as charts, so this page uses only numbers printed in text or tables.
What the paper found
The answer. Silent corruption is rare per disk and common across a large fleet. Over 41 months, 3,855 of 1.53 million disks developed at least one checksum mismatch: 3,088 nearline disks (0.86%) and 767 enterprise disks (0.065%). Nearline disks and their adapters were about an order of magnitude more likely to be affected. In the first 17 months in the field the shares were 0.66% and 0.06%. Averaged over a year, the paper's Table 2 puts the share of disks affected at 0.466% for nearline and 0.042% for enterprise, against 9.5% and 1.4% for latent sector errors, which a drive detects itself. Checksum mismatches are therefore roughly an order of magnitude rarer than latent sector errors, but nothing in the drive reports them.
The mismatches cluster. Among the 3,855 corrupt disks, the median had 3 mismatches, the mean was 104, the mode was 1, and the maximum for one drive was 33,000. The top 1% of corrupt disks produced more than half of all mismatches. A nearline disk that had one mismatch had a probability of about 0.6 of developing more, against 0.0066 for a first mismatch in 17 months. For disks with 2 to 10 mismatches, more than half of corrupt blocks on nearline disks had a corrupt neighbor in the immediately adjacent block, and across drives with at least 2 mismatches an average of 3.4 consecutive blocks were affected. In one nearline system, 92 disks developed mismatches; the authors put the chance of that happening independently at below 1 in 10 to the power 12, which points to a shared component such as a shelf controller or adapter.
How it was found matters. Scrubs found about 49% of nearline mismatches and 73% of enterprise mismatches. RAID reconstruction found about 8% of nearline mismatches, so a corrupt block was discovered only when a disk had already failed, which is when it can cause real data loss. The paper argues for more aggressive scrubbing and double-parity protection. The two other corruption classes are rarer: 365 disks had identity discrepancies, and parity inconsistencies affected 4.4 times fewer nearline disks than checksum mismatches did (3.5 times fewer for enterprise disks). Both were found only because the system stores block identity and parity as well as checksums.
What the independent report adds. CERN's 2007 measurements show the same kind of silent error in a different setting: 22 of 33,700 files (about 8.7 TB) failed a checksum comparison, about 1 bad file in 1,500, and RAID verification on 492 systems (about 1.5 PB) found about 300 block problems in 4 weeks, where the vendors' quoted bit error rate would predict about 850. A pattern-writing probe on more than 3,000 nodes found 500 errors on 100 nodes in about 5 weeks, 80% of them 64 KB regions of corrupted data. These are counts from one site's own tests, not per-disk shares, so they agree on the existence and the clumpy character of corruption, not on a rate.
The medium behind these counts is thehard disk drivecard. Whole-drive replacement rates are compared with datasheet figures in thedatasheet versus field replacement note.
Sources
- 01An Analysis of Data Corruption in the Storage Stack, Bairavasundaram, Goodson, Schroeder, Arpaci-Dusseau, and Arpaci-Dusseau, FAST 2008 (USENIX HTML proceedings) · accessed 2026-10-10
- 02An Analysis of Data Corruption in the Storage Stack (USENIX PDF, used for numerals the HTML conversion drops) · accessed 2026-10-10
- 03Data integrity, Bernd Panzer-Steindel, CERN IT, draft 1.3, 8 April 2007 · accessed 2026-10-10