Skip to content

Paper note

0.86% of nearline disks and 0.065% of enterprise disks developed a checksum mismatch in 41 months

Published 2026-10-10

In Bairavasundaram et al. (FAST 2008), 3,088 of about 358,000 nearline disks (0.86%) and 767 of 1.17 million enterprise-class disks (0.065%) developed at least one silent checksum mismatch over 41 months, in field logs covering 1.53 million disks. That is about 400,000 mismatched blocks in all, but the burden is lopsided: the median affected disk had 3 mismatches, the mean was 104, and the worst disk had 33,000. The paper's rates are per disk and come from one vendor's support logs, so they do not say how often a file is corrupted. An independent 2007 CERN report found 22 mismatching files among 33,700 checked, about 1 in 1,500, using a different method and a different population.

Population
1.53 million disks: about 358,000 nearline (SATA behind adapters) and 1.17 million enterprise-class (Fibre Channel)
Window
41 months starting January 2004; disks shipped after January 2004; 14 disk families, 31 models
Source of data
Support logs from tens of thousands of production and development storage systems at hundreds of customer sites
Measure
Disks with at least one checksum mismatch: a 4 KB block that fails its stored checksum when read. Detected corruption only
Cross-check
CERN IT, January and February 2007: pattern probes on more than 3,000 nodes, RAID verification, and file checksum comparison

Numbers

Silent data corruption, as printed in Bairavasundaram et al., FAST 2008 (1.53 million disks, 41 months from January 2004, support logs) and, for rows 15 to 17, in a 2007 CERN report. A checksum mismatch is a 4 KB block that fails its stored checksum. NL is nearline and ES is enterprise class. Percent of disks is the share of disks with at least one event. Qualifiers such as 'about' and 'more than' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
Disks in the sample1.53 million (about 358,000 NL; 1.17 million ES)disks41 months from January 2004; 14 families; 31 modelsFAST 2008 (PDF)
Checksum mismatches observedabout 400,0004 KB blocksWhole sample, 41 monthsFAST 2008 (HTML)
Nearline disks with at least one mismatch0.86 (3,088 disks)percent of disksAbout 358,000 nearline disks, 41 monthsFAST 2008 (PDF)
Enterprise disks with at least one mismatch0.065 (767 disks)percent of disks1.17 million enterprise disks, 41 monthsFAST 2008 (PDF)
First 17 months in the field, NL and ES0.66 and 0.06percent of disksWeighted class averagesFAST 2008 (PDF)
Mismatches per corrupt disk: median, mean, maximum3, 104, 33,000blocks per disk3,855 corrupt disks, 41 monthsFAST 2008 (PDF)
Share of all mismatches from the top 1% of corrupt disksmore than 50percent of mismatchesCorrupt disks onlyFAST 2008 (HTML)
Chance of more mismatches after one, NLabout 0.6 versus 0.0066 for a firstprobability, unitlessFirst 17 months in the fieldFAST 2008 (HTML)
Corrupt blocks with a corrupt neighbor block, NLmore than 50percent of corrupt blocksDisks with 2 to 10 mismatchesFAST 2008 (HTML)
Found by scrubs, NL and ESabout 49 and 73percent of mismatchesWeighted class averagesFAST 2008 (HTML)
Found during RAID reconstruction, NLabout 8percent of mismatchesWeighted class averageFAST 2008 (HTML)
Disks affected per year, checksum mismatches versus latent sector errors, NL0.466 versus 9.5percent of disks per yearTable 2 averagesFAST 2008 (HTML)
Disks affected per year, checksum mismatches versus latent sector errors, ES0.042 versus 1.4percent of disks per yearTable 2 averagesFAST 2008 (HTML)
Disks with identity discrepancies365disks1.53 million disks, whole studyFAST 2008 (PDF)
Files failing a checksum comparison (CERN)22 of 33,700, about 1 in 1,500files; about 8.7 TB checkedCERN disk pool, early 2007, checked against tape checksumsCERN 2007
RAID verification block problems versus vendor-rate expectation (CERN)about 300 versus about 850blocks in 4 weeks492 systems, about 1.5 PBCERN 2007
Pattern probe errors (CERN)500 errors on 100 nodeserrors in about 5 weeksMore than 3,000 nodes, probe every 2 hoursCERN 2007

This table is a short extract of printed figures, not a copy of the papers and not the underlying logs. Rows 1 to 14 are from Bairavasundaram et al. (HTML and PDF); rows 15 to 17 are from the CERN report and are a different measure. Rates marked per disk are shares of disks, not shares of blocks or files.

Method

The figures are copied from the USENIX HTML and PDF versions of Bairavasundaram, Goodson, Schroeder, Arpaci-Dusseau, and Arpaci-Dusseau, FAST 2008, not refit. The HTML conversion drops several numerals, so each of them was checked against the PDF. The authors mined the support logs of tens of thousands of production and development storage systems at hundreds of customer sites, for 41 months starting in January 2004, and kept only disks shipped after January 2004. The sample is 1.53 million disks in 14 families and 31 models. A checksum mismatch is a 4 KB block whose stored checksum fails when the RAID layer reads it, from any read, scrub, or rebuild. The paper separates this from identity discrepancies (found by file system block identity checks, from lost or misdirected writes) and parity inconsistencies (found only by scrubs). A corrupt disk is a disk with at least one corrupt block. The independent cross-check is the CERN IT report 'Data integrity' (Panzer-Steindel, draft 1.3, 8 April 2007), which ran its own probes in January and February 2007. It reports counts, not a fleet-wide rate per disk, and does not test the NetApp percentages.

Limits

Only corruption that the system detected is counted. Corruption nobody reads, and nobody scrubs, is invisible. The authors cannot verify that every disk was scrubbed, and a RAID group was scrubbed about once every two weeks on average. Not all customers enable logging, and some enable it only after a period of use. Disks that were removed are counted only up to removal, so error-prone disks may drop out of view. The paper says the counts mix causes: bit-level corruption, torn writes, and misdirected writes, and it names faulty adapters or shelf controllers for three disk models, so 'disk' here includes its adapter path, not only the platters. Workload data are coarse per-system weekly counts, so the finding of no correlation (below 0.1) may hide a per-drive effect. Disk models are anonymized. Nearline and enterprise disks use different on-disk formats (512-byte sectors with a segment after every eight, versus 520-byte sectors). The paper prints a mean of 104 mismatches per corrupt disk for the whole 41-month sample and a mean of 78 in its first-17-months analysis, and the text does not reconcile the two beyond the different windows. The CERN report is a draft written for internal discussion, from one site, with probe programs it designed itself, not peer-reviewed. Its expected-error comparison relies on a vendor bit error rate of 1 in 10 to the power 14 bits. Its counts are not comparable to per-disk fractions. Several of the NetApp results are shown only as charts, so this page uses only numbers printed in text or tables.

What the paper found

The answer. Silent corruption is rare per disk and common across a large fleet. Over 41 months, 3,855 of 1.53 million disks developed at least one checksum mismatch: 3,088 nearline disks (0.86%) and 767 enterprise disks (0.065%). Nearline disks and their adapters were about an order of magnitude more likely to be affected. In the first 17 months in the field the shares were 0.66% and 0.06%. Averaged over a year, the paper's Table 2 puts the share of disks affected at 0.466% for nearline and 0.042% for enterprise, against 9.5% and 1.4% for latent sector errors, which a drive detects itself. Checksum mismatches are therefore roughly an order of magnitude rarer than latent sector errors, but nothing in the drive reports them.

The mismatches cluster. Among the 3,855 corrupt disks, the median had 3 mismatches, the mean was 104, the mode was 1, and the maximum for one drive was 33,000. The top 1% of corrupt disks produced more than half of all mismatches. A nearline disk that had one mismatch had a probability of about 0.6 of developing more, against 0.0066 for a first mismatch in 17 months. For disks with 2 to 10 mismatches, more than half of corrupt blocks on nearline disks had a corrupt neighbor in the immediately adjacent block, and across drives with at least 2 mismatches an average of 3.4 consecutive blocks were affected. In one nearline system, 92 disks developed mismatches; the authors put the chance of that happening independently at below 1 in 10 to the power 12, which points to a shared component such as a shelf controller or adapter.

How it was found matters. Scrubs found about 49% of nearline mismatches and 73% of enterprise mismatches. RAID reconstruction found about 8% of nearline mismatches, so a corrupt block was discovered only when a disk had already failed, which is when it can cause real data loss. The paper argues for more aggressive scrubbing and double-parity protection. The two other corruption classes are rarer: 365 disks had identity discrepancies, and parity inconsistencies affected 4.4 times fewer nearline disks than checksum mismatches did (3.5 times fewer for enterprise disks). Both were found only because the system stores block identity and parity as well as checksums.

What the independent report adds. CERN's 2007 measurements show the same kind of silent error in a different setting: 22 of 33,700 files (about 8.7 TB) failed a checksum comparison, about 1 bad file in 1,500, and RAID verification on 492 systems (about 1.5 PB) found about 300 block problems in 4 weeks, where the vendors' quoted bit error rate would predict about 850. A pattern-writing probe on more than 3,000 nodes found 500 errors on 100 nodes in about 5 weeks, 80% of them 64 KB regions of corrupted data. These are counts from one site's own tests, not per-disk shares, so they agree on the existence and the clumpy character of corruption, not on a rate.

The medium behind these counts is thehard disk drivecard. Whole-drive replacement rates are compared with datasheet figures in thedatasheet versus field replacement note.

Sources

  1. 01An Analysis of Data Corruption in the Storage Stack, Bairavasundaram, Goodson, Schroeder, Arpaci-Dusseau, and Arpaci-Dusseau, FAST 2008 (USENIX HTML proceedings) · accessed 2026-10-10
  2. 02An Analysis of Data Corruption in the Storage Stack (USENIX PDF, used for numerals the HTML conversion drops) · accessed 2026-10-10
  3. 03Data integrity, Bernd Panzer-Steindel, CERN IT, draft 1.3, 8 April 2007 · accessed 2026-10-10