Skip to content

Paper note

Google measured a disk MTTF of 10 to 50 years but a node MTTF of 4.3 months; 37% of node failures came in bursts

Published 2026-10-10

In Ford et al. (OSDI 2010), one year of data from tens of Google storage cells of 1000 to 7000 nodes each, the mean time to failure was 10 to 50 years for a disk, 10.2 years for a rack, and 4.3 months for a node, so node failures, not disk failures, drove availability. 37% of failures were part of a burst of at least 2 nodes. In the authors' model, a 10% cut in the disk failure rate raised stripe availability by less than 1.5%, while a 10% cut in the node failure rate raised data availability by 18%. Maneas et al. (FAST 2020, NetApp SSDs) and Schroeder and Gibson (FAST 2007, HPC1 disks) also found strong clustering of drive replacements in time, so a failure rate alone does not describe risk.

Population
Tens of Google storage cells, 1000 to 7000 nodes each, one operator
Window
One year of operation; events of 15 minutes or longer
Failure definition
Node unavailability from software or hardware causes; disk, node, and rack MTTF reported separately
Burst definition
Node failures each within 120 seconds of the next
Cross-checks
NetApp SSD RAID groups, FAST 2020 (about 1.4 million drives); HPC1 disk replacements, FAST 2007

Numbers

Availability and correlated-failure figures as printed in Ford et al., OSDI 2010 (Google storage cells, one year), rows 1 to 10, and in two cross-check sources, rows 11 to 14. MTTF is mean time to failure. Qualifiers such as 'less than' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
Storage nodes per cell studied1000 to 7000nodes per cellTens of Google storage cells, one year
Node unavailability events lasting longer than 15 minutes10percent of eventsSame cells
Disk mean time to failure10 to 50yearsGoogle cells; counts software and hardware causes
Node mean time to failure4.3monthsSame cells
Rack mean time to failure10.2yearsSame cells
Failures that are part of a burst of at least 2 nodes37percent of failuresBursts defined with a 120 second window
Random failures wrongly merged into a burst8.0percent of non-correlated failuresSame window; 0.068% for a burst of at least 10 nodes
Older data blocks that fail a checksum on scrubbing1 in 10^6 to 10^7fraction of older blocksGFS background scrubbing
Gain in stripe availability from a 10% cut in the disk failure rate1.5percent (upper bound)Model with R = 3 replication
Gain in data availability from a 10% cut in the node failure rate18percentSame model
Chance of a drive replacement in a random week0.0504percentNetApp RAID groups, about 1.4 million SSDs
Chance of a replacement within a week of a previous one in the same RAID group9.39percentSame fleet
Correlation of disk replacements in consecutive weeks0.72correlation coefficientHPC1 system, whole lifetime
Spread in expected weekly replacements after a quiet versus a busy week9timesHPC1 system

This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 10 are from Ford et al. (OSDI 2010). Rows 11 and 12 are from Maneas et al. (FAST 2020) and rows 13 and 14 from Schroeder and Gibson (FAST 2007); they are cross-checks with different measures.

Method

The figures are copied from the PDF of Ford, Labelle, Popovici, Stokely, Truong, Barroso, Grimes, and Quinlan, USENIX OSDI 2010, not refit. The data are node unavailability events from tens of Google storage cells, each with 1000 to 7000 nodes, over one year; the paper counts an event only if the node is unavailable for 15 minutes or longer. Component mean time to failure includes software errors and hardware failures. A failure burst is a run of node failures each within 120 seconds of the next. The availability results come from the authors' Markov model. Two other sources check the surrounding claim that failures cluster in time: Maneas et al., FAST 2020 (about 1.4 million NetApp SSDs in RAID groups) and Schroeder and Gibson, FAST 2007 (the HPC1 disk replacement log). Neither tests the Google model results.

Limits

The numbers are one operator's own cells, hardware, and software, and they describe nodes and racks as well as disks. Disk MTTF is printed as a range, 10 to 50 years, and that is not the same as a vendor rating or an annual replacement rate. The 10% sensitivity results are from the authors' model and assume 3-way replication (R = 3); they are not an experiment on a live system. The 37% share depends on the 120 second window; the authors estimate that 8.0% of non-correlated failures could be merged by chance. Maneas et al. count drive replacements in RAID groups of SSDs, and Schroeder and Gibson count replacements in one HPC1 system, so their measures are not the same as node unavailability and the numbers are not comparable. The paper does not model fail-slow drives, and its only silent corruption figure is one scrubbing rate.

What the paper found

The answer. In Ford et al., the unit that fails most often is the node, not the disk. The mean time to failure in Google's cells was 10 to 50 years for a disk, 10.2 years for a rack of nodes, and 4.3 months for a node, and the node count includes software restarts and other transient causes. Because nodes fail so much more often, the authors find they matter far more for data availability, even though disk failures are the ones that are permanent. In their model with 3-way replication, a 10% cut in the disk failure rate raised stripe availability by less than 1.5%, a 10% cut in the latent disk error rate had a negligible effect, and a 10% cut in the node failure rate raised availability by 18%.

Failures arrive in bursts. When the authors group node failures that occur within 120 seconds of each other, 37% of failures are part of a burst of at least 2 nodes, and they estimate that close to 37% are truly correlated, since only 8.0% of non-correlated failures could be merged by chance. Most large bursts are tied to rack-level or multi-rack events such as shared switches, power, or planned kernel upgrades. Ignoring such correlation makes a model too optimistic: with correlated failures included, even a 90% reduction in recovery time gave only a 6% reduction in unavailability.

Disk-level numbers are in line with other studies. The authors say their disk figures are comparable with published disk and storage subsystem failure rates, and background scrubbing found checksum mismatches in 1 in 10^6 to 10^7 older data blocks, concentrated on a small number of disks. Disk errors are therefore a small contributor to unavailability but a real one for data integrity, which the system handles with scrubbing and verification on reads.

Other fleets show the same clustering. Maneas et al. found a 0.0504% chance of a drive replacement in a random week but 9.39% in the week after a previous replacement in the same RAID group, more than 180 times higher, and 52% of consecutive replacements fell within a week. Schroeder and Gibson found a correlation coefficient of 0.72 between disk replacement counts in consecutive weeks of the HPC1 system, and the expected number of replacements in a week varied by a factor of 9 depending on how busy the previous week was. The three fleets, devices, and measures differ, so the sizes are not comparable, but all three reject the assumption of independent failures.

Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.

Sources

  1. 01Availability in Globally Distributed Storage Systems, Ford, Labelle, Popovici, Stokely, Truong, Barroso, Grimes, and Quinlan, USENIX OSDI 2010 (USENIX PDF) · accessed 2026-10-10
  2. 02A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, section 6 on correlated replacements) · accessed 2026-10-10
  3. 03Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX PDF, section 5.2 on correlations) · accessed 2026-10-10