What fails most often: unreadable data on disks, uncorrectable reads on SSDs, and bus errors on the link
Published 2026-10-09
The most frequent field fault is not a dead drive but data that cannot be read: 3.45% of 1.53 million hard disks developed at least one latent sector error in 32 months, and 20.3–90.5% of Google's flash drives (ten models, first 4 years; 3 for eMLC) had at least one uncorrectable error and 19.1–62.7% a final read error, while 4.2–34.1% of Facebook's SSD platforms reported at least one uncorrectable error. Link faults show up separately as CRC counts (about 2% of Google's disks), and the public counters miss a large share of failed drives: over 56% of failed drives in the Google study and 23.3% in Backblaze's 2016 log showed no signal on the SMART attributes each used.
- Hard disks (latent errors)
- 1.53 million enterprise and nearline disks, 32 months, production storage systems
- SSDs, Google
- Ten drive models (MLC, SLC, eMLC), many millions of drive days, first four years in the field (three for eMLC)
- SSDs, Facebook
- Majority of flash-based SSDs in Facebook data centers over nearly four years; platform-level incidence
- SMART blind spots
- Google: over 100,000 consumer-grade ATA disks; Backblaze: 67,814 drives, five SMART attributes
- Link
- CRC counts in the Google disk study; error classification in the Linux libata documentation
How often each fault appears
| Fault (as the source names it) | Medium | Value | Unit | Population and window | Source |
|---|---|---|---|---|---|
| Latent sector error, at least one | HDD | 3.45 | percent of disks | 1.53 million disks, 32 months | SIGMETRICS 2007 |
| Latent sector error, nearline disks | HDD | 8.5 | percent of disks | nearline subset, 32 months | SIGMETRICS 2007 |
| Latent sector error, enterprise disks | HDD | 1.9 | percent of disks | enterprise subset, 32 months | SIGMETRICS 2007 |
| final read error, at least one | SSD (Google flash) | 19.1–62.7 | percent of drives | ten models, first 4 years (eMLC: 3); Table 2; paper text states 20–63 | FAST 2016 |
| uncorrectable error, at least one | SSD (Google flash) | 20.3–90.5 | percent of drives | ten models, first 4 years (eMLC: 3); Table 2 | FAST 2016 |
| Uncorrectable error, at least one | SSD (Facebook) | 4.2–34.1 | percent of SSDs | per platform, about 4 years of data | SIGMETRICS 2015 |
| Bad block developed in the field | SSD (Google flash) | 30–80 | percent of drives | ten models, first 4 years (eMLC: 3) | FAST 2016 |
| Bad chip | SSD (Google flash) | 2–7 | percent of drives | ten models, first 4 years (eMLC: 3) | FAST 2016 |
| CRC errors | HDD (link) | about 2 | percent of disks | over 100,000 ATA disks, 9 months of SMART data | FAST 2007 |
What the public counters did not see
| Study | Attributes checked | Failed drives with no signal | Unit | Operational drives with a signal | Source |
|---|---|---|---|---|---|
| Google, 2005–2006 window | scan errors, reallocation, offline reallocation, probational count | over 56 | percent of failed drives | not stated | FAST 2007 |
| Backblaze, October 2016 post | SMART 5, 187, 188, 197, 198 | 23.3 | percent of failed drives | 4.2 percent of operational drives | Backblaze 2016 |
Where errors cluster
| Measure | Medium | Value | Unit | Source |
|---|---|---|---|---|
| Median errors per disk with at least one error | HDD | 3 | errors per disk | SIGMETRICS 2007 |
| Error disks with more than 1000 errors | HDD | 0.2 | percent of error disks | SIGMETRICS 2007 |
| Share of uncorrectable errors in the 10% of SSDs with most errors | SSD (Facebook) | over 80 | percent of errors | SIGMETRICS 2015 |
| Chance of another uncorrectable error next month, after one | SSD (Google flash) | nearly 30 | percent of drive months | FAST 2016 |
| Same chance in a random month | SSD (Google flash) | about 2 | percent of drive months | FAST 2016 |
Reading the numbers
The direct answer to what fails most often is unreadable data, not a drive that stops. On hard disks, 3.45% of 1.53 million disks (53,820) developed at least one latent sector error over 32 months, which is an error found only when the sector is read; nearline disks were affected at 8.5% and enterprise disks at 1.9%. Google's flash paper restates the disk figure as 3.5% over 32 months, so the number is repeated by an independent group, but it is the same underlying dataset and not a second measurement. On SSDs, the two public field studies measure read failures on different populations and agree on direction, not on size: in Google's drives (Table 2 of the flash paper), 20.3–90.5% had at least one uncorrectable error and 19.1–62.7% at least one final read error (a read that still fails after retries) in the first four years (three for the eMLC models), and 4.2–34.1% of Facebook's SSDs per platform had an uncorrectable error.
Errors cluster rather than spread evenly. In the disk study, the median error disk had 3 errors, the mode was 1 error (30% of error disks), and 0.2% of error disks had more than 1000. In the Facebook SSD data, the 10% of SSDs with the most errors held over 80% of all uncorrectable errors on every platform. In Google's flash data, a drive with an uncorrectable error had nearly a 30% chance of another in the next month, against about 2% in a random month, and in the MLC models half of the drives that reached two bad blocks went on to about 200 or more. That is the practical reason a first error matters more than a count.
On the link, the evidence is thinner. The Google disk study reports CRC errors on about 2% of its disks and says they point to cables and connectors more than to the drive. The Linux libata documentation agrees on the classification: an ICRC bit means corruption during transfer and is handled as an ATA bus error, not a device error, and a SCSI sense of HARDWARE ERROR with ASC/ASCQ 47h/00h (parity error) is also treated as a bus error. By contrast the UNC bit is a media error reported only after the drive's retries failed. Neither source gives a fleet-wide link failure rate, so none is stated here.
The counters miss failures. Google found over 56% of failed drives had no count on scan errors, reallocation, offline reallocation or probational count. Backblaze reported that 76.7% of failed drives had at least one of five SMART attributes (5, 187, 188, 197, 198) above zero, leaving 23.3% with none, while 4.2% of operational drives also had a non-zero raw value. Those two numbers use different attribute sets and cannot be averaged. For SSDs, Google found the raw bit error rate is not a good predictor of uncorrectable errors, so a rising corrected-error count alone does not say a read failure is coming.
Related: the SMART failure-signal noteand the hard disk drive card.
Method
Every figure is copied from a source that was opened on the access date shown in the source list, not refit. Each fault type is backed by two independent publications where one exists: latent sector errors (Bairavasundaram et al., with the figure restated in Schroeder et al.), SSD uncorrectable errors (Schroeder et al. at Google and Meza et al. at Facebook), SMART blind spots (Pinheiro et al. at Google and Backblaze), and the link (Pinheiro et al. and the Linux kernel libata documentation). The tables keep each study's own unit and window so rows are not merged into one rate.
Limits
The studies use different populations, years, failure definitions and windows, so rows must not be ranked against each other. The hard-disk latent sector error study covers 32 months of enterprise and nearline disks from production storage systems. The Google flash study uses custom PCIe drives (ten models, the first four years of in-field errors per drive, three for the two eMLC models; the paper's own text and summary quote 20–63% for uncorrectable errors, but its Table 2 gives 19.1–62.7% for final read errors and 20.3–90.5% for uncorrectable errors, so this page uses the table), not retail SATA or NVMe drives. The Facebook figures are platform-level incidence at the time of the study. The Backblaze figure comes from one operator's October 2016 post covering 67,814 drives and five SMART attributes; its per-attribute chart is an image and is not reproduced. The kernel documentation describes how Linux classifies ATA errors; it gives no failure rates. This page does not cover NVMe health-log fields, SCSI sense data in practice, or tape: no source for those was opened for this note, so they are not claimed. A failure is not the same as an error: a drive can report errors for years, and a replacement is a policy decision.
Sources
- 01An Analysis of Latent Sector Errors in Disk Drives, Bairavasundaram, Goodson, Pasupathy and Schindler, SIGMETRICS 2007 · accessed 2026-10-09
- 02Flash Reliability in Production: The Expected and the Unexpected, Schroeder, Lagisetty and Merchant, FAST 2016 · accessed 2026-10-09
- 03A Large-Scale Study of Flash Memory Failures in the Field, Meza, Wu, Kumar and Mutlu, SIGMETRICS 2015 · accessed 2026-10-09
- 04Failure Trends in a Large Disk Drive Population, Pinheiro, Weber and Barroso, FAST 2007 (USENIX HTML) · accessed 2026-10-09
- 05What SMART Stats Tell Us About Hard Drives, Backblaze, October 6, 2016 · accessed 2026-10-09
- 06libATA Developer's Guide, ATA error handling (Linux kernel documentation) · accessed 2026-10-09