Skip to content

Stat analysis

What fails most often: unreadable data on disks, uncorrectable reads on SSDs, and bus errors on the link

Published 2026-10-09

The most frequent field fault is not a dead drive but data that cannot be read: 3.45% of 1.53 million hard disks developed at least one latent sector error in 32 months, and 20.3–90.5% of Google's flash drives (ten models, first 4 years; 3 for eMLC) had at least one uncorrectable error and 19.1–62.7% a final read error, while 4.2–34.1% of Facebook's SSD platforms reported at least one uncorrectable error. Link faults show up separately as CRC counts (about 2% of Google's disks), and the public counters miss a large share of failed drives: over 56% of failed drives in the Google study and 23.3% in Backblaze's 2016 log showed no signal on the SMART attributes each used.

Hard disks (latent errors)
1.53 million enterprise and nearline disks, 32 months, production storage systems
SSDs, Google
Ten drive models (MLC, SLC, eMLC), many millions of drive days, first four years in the field (three for eMLC)
SSDs, Facebook
Majority of flash-based SSDs in Facebook data centers over nearly four years; platform-level incidence
SMART blind spots
Google: over 100,000 consumer-grade ATA disks; Backblaze: 67,814 drives, five SMART attributes
Link
CRC counts in the Google disk study; error classification in the Linux libata documentation

How often each fault appears

Share of devices that reported at least one event of the kind named, in each study's own window. Rows are not additive and not ranked against each other. Windows: disks 32 months; Google flash first 4 years; Facebook SSDs as reported per platform.
Fault (as the source names it)MediumValueUnitPopulation and windowSource
Latent sector error, at least oneHDD3.45percent of disks1.53 million disks, 32 monthsSIGMETRICS 2007
Latent sector error, nearline disksHDD8.5percent of disksnearline subset, 32 monthsSIGMETRICS 2007
Latent sector error, enterprise disksHDD1.9percent of disksenterprise subset, 32 monthsSIGMETRICS 2007
final read error, at least oneSSD (Google flash)19.1–62.7percent of drivesten models, first 4 years (eMLC: 3); Table 2; paper text states 20–63FAST 2016
uncorrectable error, at least oneSSD (Google flash)20.3–90.5percent of drivesten models, first 4 years (eMLC: 3); Table 2FAST 2016
Uncorrectable error, at least oneSSD (Facebook)4.2–34.1percent of SSDsper platform, about 4 years of dataSIGMETRICS 2015
Bad block developed in the fieldSSD (Google flash)30–80percent of drivesten models, first 4 years (eMLC: 3)FAST 2016
Bad chipSSD (Google flash)2–7percent of drivesten models, first 4 years (eMLC: 3)FAST 2016
CRC errorsHDD (link)about 2percent of disksover 100,000 ATA disks, 9 months of SMART dataFAST 2007

What the public counters did not see

Failed drives with no non-zero value on the attributes each study used. The attribute sets differ, so the two rows are not comparable.
StudyAttributes checkedFailed drives with no signalUnitOperational drives with a signalSource
Google, 2005–2006 windowscan errors, reallocation, offline reallocation, probational countover 56percent of failed drivesnot statedFAST 2007
Backblaze, October 2016 postSMART 5, 187, 188, 197, 19823.3percent of failed drives4.2 percent of operational drivesBackblaze 2016

Where errors cluster

Concentration and repeat-error figures from the same studies.
MeasureMediumValueUnitSource
Median errors per disk with at least one errorHDD3errors per diskSIGMETRICS 2007
Error disks with more than 1000 errorsHDD0.2percent of error disksSIGMETRICS 2007
Share of uncorrectable errors in the 10% of SSDs with most errorsSSD (Facebook)over 80percent of errorsSIGMETRICS 2015
Chance of another uncorrectable error next month, after oneSSD (Google flash)nearly 30percent of drive monthsFAST 2016
Same chance in a random monthSSD (Google flash)about 2percent of drive monthsFAST 2016

Reading the numbers

The direct answer to what fails most often is unreadable data, not a drive that stops. On hard disks, 3.45% of 1.53 million disks (53,820) developed at least one latent sector error over 32 months, which is an error found only when the sector is read; nearline disks were affected at 8.5% and enterprise disks at 1.9%. Google's flash paper restates the disk figure as 3.5% over 32 months, so the number is repeated by an independent group, but it is the same underlying dataset and not a second measurement. On SSDs, the two public field studies measure read failures on different populations and agree on direction, not on size: in Google's drives (Table 2 of the flash paper), 20.3–90.5% had at least one uncorrectable error and 19.1–62.7% at least one final read error (a read that still fails after retries) in the first four years (three for the eMLC models), and 4.2–34.1% of Facebook's SSDs per platform had an uncorrectable error.

Errors cluster rather than spread evenly. In the disk study, the median error disk had 3 errors, the mode was 1 error (30% of error disks), and 0.2% of error disks had more than 1000. In the Facebook SSD data, the 10% of SSDs with the most errors held over 80% of all uncorrectable errors on every platform. In Google's flash data, a drive with an uncorrectable error had nearly a 30% chance of another in the next month, against about 2% in a random month, and in the MLC models half of the drives that reached two bad blocks went on to about 200 or more. That is the practical reason a first error matters more than a count.

On the link, the evidence is thinner. The Google disk study reports CRC errors on about 2% of its disks and says they point to cables and connectors more than to the drive. The Linux libata documentation agrees on the classification: an ICRC bit means corruption during transfer and is handled as an ATA bus error, not a device error, and a SCSI sense of HARDWARE ERROR with ASC/ASCQ 47h/00h (parity error) is also treated as a bus error. By contrast the UNC bit is a media error reported only after the drive's retries failed. Neither source gives a fleet-wide link failure rate, so none is stated here.

The counters miss failures. Google found over 56% of failed drives had no count on scan errors, reallocation, offline reallocation or probational count. Backblaze reported that 76.7% of failed drives had at least one of five SMART attributes (5, 187, 188, 197, 198) above zero, leaving 23.3% with none, while 4.2% of operational drives also had a non-zero raw value. Those two numbers use different attribute sets and cannot be averaged. For SSDs, Google found the raw bit error rate is not a good predictor of uncorrectable errors, so a rising corrected-error count alone does not say a read failure is coming.

Related: the SMART failure-signal noteand the hard disk drive card.

Method

Every figure is copied from a source that was opened on the access date shown in the source list, not refit. Each fault type is backed by two independent publications where one exists: latent sector errors (Bairavasundaram et al., with the figure restated in Schroeder et al.), SSD uncorrectable errors (Schroeder et al. at Google and Meza et al. at Facebook), SMART blind spots (Pinheiro et al. at Google and Backblaze), and the link (Pinheiro et al. and the Linux kernel libata documentation). The tables keep each study's own unit and window so rows are not merged into one rate.

Limits

The studies use different populations, years, failure definitions and windows, so rows must not be ranked against each other. The hard-disk latent sector error study covers 32 months of enterprise and nearline disks from production storage systems. The Google flash study uses custom PCIe drives (ten models, the first four years of in-field errors per drive, three for the two eMLC models; the paper's own text and summary quote 20–63% for uncorrectable errors, but its Table 2 gives 19.1–62.7% for final read errors and 20.3–90.5% for uncorrectable errors, so this page uses the table), not retail SATA or NVMe drives. The Facebook figures are platform-level incidence at the time of the study. The Backblaze figure comes from one operator's October 2016 post covering 67,814 drives and five SMART attributes; its per-attribute chart is an image and is not reproduced. The kernel documentation describes how Linux classifies ATA errors; it gives no failure rates. This page does not cover NVMe health-log fields, SCSI sense data in practice, or tape: no source for those was opened for this note, so they are not claimed. A failure is not the same as an error: a drive can report errors for years, and a replacement is a policy decision.

Sources

  1. 01An Analysis of Latent Sector Errors in Disk Drives, Bairavasundaram, Goodson, Pasupathy and Schindler, SIGMETRICS 2007 · accessed 2026-10-09
  2. 02Flash Reliability in Production: The Expected and the Unexpected, Schroeder, Lagisetty and Merchant, FAST 2016 · accessed 2026-10-09
  3. 03A Large-Scale Study of Flash Memory Failures in the Field, Meza, Wu, Kumar and Mutlu, SIGMETRICS 2015 · accessed 2026-10-09
  4. 04Failure Trends in a Large Disk Drive Population, Pinheiro, Weber and Barroso, FAST 2007 (USENIX HTML) · accessed 2026-10-09
  5. 05What SMART Stats Tell Us About Hard Drives, Backblaze, October 6, 2016 · accessed 2026-10-09
  6. 06libATA Developer's Guide, ATA error handling (Linux kernel documentation) · accessed 2026-10-09