Google measured a disk MTTF of 10 to 50 years but a node MTTF of 4.3 months; 37% of node failures came in bursts
Published 2026-10-10
In Ford et al. (OSDI 2010), one year of data from tens of Google storage cells of 1000 to 7000 nodes each, the mean time to failure was 10 to 50 years for a disk, 10.2 years for a rack, and 4.3 months for a node, so node failures, not disk failures, drove availability. 37% of failures were part of a burst of at least 2 nodes. In the authors' model, a 10% cut in the disk failure rate raised stripe availability by less than 1.5%, while a 10% cut in the node failure rate raised data availability by 18%. Maneas et al. (FAST 2020, NetApp SSDs) and Schroeder and Gibson (FAST 2007, HPC1 disks) also found strong clustering of drive replacements in time, so a failure rate alone does not describe risk.
- Population
- Tens of Google storage cells, 1000 to 7000 nodes each, one operator
- Window
- One year of operation; events of 15 minutes or longer
- Failure definition
- Node unavailability from software or hardware causes; disk, node, and rack MTTF reported separately
- Burst definition
- Node failures each within 120 seconds of the next
- Cross-checks
- NetApp SSD RAID groups, FAST 2020 (about 1.4 million drives); HPC1 disk replacements, FAST 2007
Numbers
This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 10 are from Ford et al. (OSDI 2010). Rows 11 and 12 are from Maneas et al. (FAST 2020) and rows 13 and 14 from Schroeder and Gibson (FAST 2007); they are cross-checks with different measures.
Method
The figures are copied from the PDF of Ford, Labelle, Popovici, Stokely, Truong, Barroso, Grimes, and Quinlan, USENIX OSDI 2010, not refit. The data are node unavailability events from tens of Google storage cells, each with 1000 to 7000 nodes, over one year; the paper counts an event only if the node is unavailable for 15 minutes or longer. Component mean time to failure includes software errors and hardware failures. A failure burst is a run of node failures each within 120 seconds of the next. The availability results come from the authors' Markov model. Two other sources check the surrounding claim that failures cluster in time: Maneas et al., FAST 2020 (about 1.4 million NetApp SSDs in RAID groups) and Schroeder and Gibson, FAST 2007 (the HPC1 disk replacement log). Neither tests the Google model results.
Limits
The numbers are one operator's own cells, hardware, and software, and they describe nodes and racks as well as disks. Disk MTTF is printed as a range, 10 to 50 years, and that is not the same as a vendor rating or an annual replacement rate. The 10% sensitivity results are from the authors' model and assume 3-way replication (R = 3); they are not an experiment on a live system. The 37% share depends on the 120 second window; the authors estimate that 8.0% of non-correlated failures could be merged by chance. Maneas et al. count drive replacements in RAID groups of SSDs, and Schroeder and Gibson count replacements in one HPC1 system, so their measures are not the same as node unavailability and the numbers are not comparable. The paper does not model fail-slow drives, and its only silent corruption figure is one scrubbing rate.
What the paper found
The answer. In Ford et al., the unit that fails most often is the node, not the disk. The mean time to failure in Google's cells was 10 to 50 years for a disk, 10.2 years for a rack of nodes, and 4.3 months for a node, and the node count includes software restarts and other transient causes. Because nodes fail so much more often, the authors find they matter far more for data availability, even though disk failures are the ones that are permanent. In their model with 3-way replication, a 10% cut in the disk failure rate raised stripe availability by less than 1.5%, a 10% cut in the latent disk error rate had a negligible effect, and a 10% cut in the node failure rate raised availability by 18%.
Failures arrive in bursts. When the authors group node failures that occur within 120 seconds of each other, 37% of failures are part of a burst of at least 2 nodes, and they estimate that close to 37% are truly correlated, since only 8.0% of non-correlated failures could be merged by chance. Most large bursts are tied to rack-level or multi-rack events such as shared switches, power, or planned kernel upgrades. Ignoring such correlation makes a model too optimistic: with correlated failures included, even a 90% reduction in recovery time gave only a 6% reduction in unavailability.
Disk-level numbers are in line with other studies. The authors say their disk figures are comparable with published disk and storage subsystem failure rates, and background scrubbing found checksum mismatches in 1 in 10^6 to 10^7 older data blocks, concentrated on a small number of disks. Disk errors are therefore a small contributor to unavailability but a real one for data integrity, which the system handles with scrubbing and verification on reads.
Other fleets show the same clustering. Maneas et al. found a 0.0504% chance of a drive replacement in a random week but 9.39% in the week after a previous replacement in the same RAID group, more than 180 times higher, and 52% of consecutive replacements fell within a week. Schroeder and Gibson found a correlation coefficient of 0.72 between disk replacement counts in consecutive weeks of the HPC1 system, and the expected number of replacements in a week varied by a factor of 9 depending on how busy the previous week was. The three fleets, devices, and measures differ, so the sizes are not comparable, but all three reject the assumption of independent failures.
Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.
Sources
- 01Availability in Globally Distributed Storage Systems, Ford, Labelle, Popovici, Stokely, Truong, Barroso, Grimes, and Quinlan, USENIX OSDI 2010 (USENIX PDF) · accessed 2026-10-10
- 02A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, section 6 on correlated replacements) · accessed 2026-10-10
- 03Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX PDF, section 5.2 on correlations) · accessed 2026-10-10