Field disk replacements averaged 3.01% a year, 3.4 times the 0.88% a 1,000,000-hour MTTF implies
Published 2026-10-09 · updated 2026-10-10
In Schroeder and Gibson (FAST 2007), disk replacement logs from large production systems (more than 100,000 drives, up to five years of use) gave a weighted average annual replacement rate of 3.01% for drives under five years old, 3.4 times the 0.88% annual failure rate that a datasheet MTTF of 1,000,000 hours implies. Replacement rates also kept rising with age instead of settling after year one, and the time between replacements was far from the exponential model that most reliability arithmetic assumes. A replacement is not a confirmed failure, so read the 3.4 times as a gap between datasheet arithmetic and field replacement practice. Google's FAST 2007 paper reports a 1.7% to over 8.6% range, and the Backblaze Q2 2026 report gives a lifetime fleet rate of 1.41%.
- Population
- More than 100,000 drives in seven data sets from at least four vendors: SCSI, FC, and SATA
- Window
- Data sets span one month to five years, with collection periods from 2001 to 2006
- Sites
- Three HPC clusters, one multi-site HPC warranty log, and three internet-provider data sets
- Measure
- Annual replacement rate (ARR): drives replaced per year, as a percent of the drives in use. Not a confirmed failure rate
- Datasheet baseline
- MTTF of 1,000,000 to 1,500,000 hours, which the paper converts to a nominal AFR of 0.88% to 0.58%
Numbers
| Measure | Value | Unit | Population or scope | Source |
|---|---|---|---|---|
| Datasheet MTTF to nominal AFR | 1,000,000 to 1,500,000 hours is 0.88% to 0.58% | hours; percent of drives per year | Vendor datasheet arithmetic: power-on hours per year divided by MTTF | FAST 2007 |
| Weighted average ARR, drives under five years | 3.01 | percent of drives per year | All seven data sets, weighted by drive count | FAST 2007 |
| Weighted average ARR without COM3 | 2.86 | percent of drives per year | Six data sets; 3.3 times 0.88% | FAST 2007 |
| Range of ARR by data set and drive type | 0.5 to 13.5 | percent of drives per year | Drives within the five-year nominal life; one group at 24% is excluded | FAST 2007 |
| HPC1 ARR over five years | 3.4 | percent of drives per year | HPC1, a 765-node cluster, August 2001 to May 2006 | FAST 2007 |
| Replacement rate versus datasheet, years 4 and 5 | 7 to 10 | times the rate the datasheet MTTF implies | HPC1, within nominal life | FAST 2007 |
| Replacement rate versus datasheet, systems 5 to 8 years old | up to 30 | times the rate the datasheet MTTF implies | Older systems in the data sets | FAST 2007 |
| Correlation of replacement counts, consecutive weeks and months | 0.72 and 0.79 | correlation coefficient, unitless | HPC1 only | FAST 2007 |
| Time between replacements, Weibull shape and variability | 0.71 shape; 2.4 squared coefficient of variation | unitless | HPC1 only; an exponential model would give 1 for both | FAST 2007 |
| Google AFR, first year to three-year-old group | 1.7 to over 8.6 | percent of drives per year | Over 100,000 ATA drives, 5400 to 7200 rpm, repairs database of about five years | Google, FAST 2007 |
| Backblaze quarterly AFR, Q2 2026 | 1.73 | percent of drives per year | 354,415 drives, 31,553,350 drive days, 1,498 failures, April to June 2026 | Backblaze Q2 2026 |
| Backblaze lifetime AFR to June 30, 2026 | 1.41 | percent of drives per year | 353,041 drives, 558,148,607 drive days, 21,609 failures, models with more than 500 drives | Backblaze Q2 2026 |
This table is a short extract of printed figures, not a copy of the papers and not the underlying logs. The canonical text for rows 1 to 9 is the USENIX HTML proceedings of Schroeder and Gibson. Rows 10 to 12 are different studies and are listed for scale only. The 'times' ratios are the paper's own.
Method
The figures are copied from the USENIX HTML proceedings of Schroeder and Gibson, FAST 2007, not refit. The authors collected hardware replacement logs from seven data sets: three high-performance computing clusters (HPC1 to HPC3), a warranty log covering dozens of HPC sites (HPC4), and three data sets from an internet provider (COM1 to COM3). They report the annual replacement rate (ARR), the share of drives replaced per year, because a replacement is not necessarily a failure. Datasheet AFR is the vendor's number: power-on hours per year divided by MTTF, so 1,000,000 hours is 0.88% and 1,500,000 hours is 0.58%. For the time between replacements they used the HPC1 data, the only set with timestamps of when a problem was detected, and fitted exponential, Weibull, gamma, and lognormal distributions by maximum likelihood, judged by visual fit, negative log-likelihood, and chi-square tests at the 0.05 level. Two other sources cross-check the level of the rates: the Google FAST 2007 paper (a different fleet) and the Backblaze Q2 2026 report (a different era and drive mix). Neither one tests the 3.4 times ratio. The authors' own CMU project page (opened 2026-10-10) restates the headline results, including datasheet rates exceeded by a factor of 2 to 10 for drives under five years old and a factor of 30 for five to eight year old drives. It is the same research group, so it confirms that the paper text is stated consistently but is not an independent check.
Limits
A replacement is a customer decision, not a manufacturer-confirmed failure. The paper cites a vendor's personal communication that 43% of returned disks show no problem, and notes that even scaling the ARRs down to 57% leaves most estimates more than twice the datasheet AFR. Customers also use different tests and thresholds, and the paper says a customer may declare a disk faulty while its maker sees it as healthy. The ARR for each data set cannot be recomputed from Table 1 because disk counts changed during collection and some logs hold events other than replacements. The data providers chose which systems to share. The 100,000-drive sample is small next to the 300 million drives the paper says were built in 2006, so bad batches are mostly outside the sample. The one bad batch the paper names (11,000 SATA drives in HPC3, replaced in October 2006 after high media error rates, attributed to lubricant breakdown that raised head flying height) is not recorded as failures. Infant failures caught in manufacturing or installation tests are not in the logs. The time-between-replacement results come from one system, HPC1, so they are not a fleet-wide result. Differences between interfaces say little about quality, because operating conditions are mixed in. Several results are shown only as charts, so this page uses only numbers printed in the text or tables of the paper. The drives are SCSI, FC, and SATA models in use from about 1998 to 2006. The Backblaze rates come from a different era with a different drive mix, and the page I opened does not define a failure, so those rows are context, not a replication. The CMU project page is a summary by the same authors, not an independent source. The independent checks on level are the Google paper and the Backblaze report only.
What the paper found
The headline gap. The weighted average ARR across the data sets, for drives within their five-year nominal life, was 3.01%. The paper calls this 3.4 times the 0.88% that a 1,000,000-hour MTTF implies. Without the COM3 data, which had the highest rates, the average was still 2.86%, or 3.3 times 0.88%. The observed ARRs by data set and drive type ranged from 0.5% to 13.5%, up to a factor of 15 above the datasheet AFR. HPC1, which covers almost exactly the five-year nominal life, had an ARR of 3.4%. One group of COM3 drives deployed in 1998 had an ARR of 24%, but those were more than seven years old and outside the nominal life.
The gap grows with age. Replacement rates were higher than the datasheet suggested in every year except year 1, and the HPC1 file system nodes had no replacements in the first 12 months. In HPC1's second year the rates were 20% above expectation for the file system nodes and a factor of two for the compute nodes. In years 4 and 5 they were 7 to 10 times what the datasheet MTTF implies. For systems 5 to 8 years old, datasheet MTTFs understated replacement rates by as much as a factor of 30. The paper finds no steady state after year 1 and no strong infant mortality. It sees a steady rise that it calls an early onset of wear-out, and it recommends that reliability standards include wear-out. HPC1 replacement rates nearly doubled from year 1 to year 2, and from year 2 to year 3.
The independence model does not fit. A Poisson process, with independent failures and exponentially distributed gaps, is the usual assumption behind RAID loss estimates. In HPC1 the correlation coefficient between replacement counts in consecutive weeks was 0.72, and between consecutive months 0.79. The expected number of replacements in a week varied by a factor of 9 depending on whether the preceding week was in the lowest or the highest third. The time between replacements was fit best by a Weibull distribution with a shape of 0.71 over all HPC1 nodes (0.73 for compute nodes, 0.76 for file system nodes), and the exponential model was rejected at the 0.05 level. The squared coefficient of variation was 2.4, against 1 for an exponential distribution. A shape under 1 means a decreasing hazard rate: right after a replacement the expected wait for the next one was about 4 days, but after 10 days without a replacement it had grown to 10 days, and after 20 days to 15 days. For a rebuild window of a few hours, the paper reports that two drives failing within one hour was four times more likely in the real data than under the exponential model, and within 10 hours two times more likely.
Where a drive sits among the parts that get replaced. In HPC1 the hard drive was 30.6% of all hardware replacements, in COM1 18.1%, and in COM2 49.1%. Among HPC1 node outages attributed to hardware, drives were 16% of outage records, behind CPU at 44% and memory at 29%, but about 90% of the drive problems led to a drive replacement, a costlier repair than the simple fix that cleared most CPU and memory events.
Two other sources agree on the level, not on the ratio. The Google paper (Pinheiro, Weber, and Barroso, FAST 2007) reports annualized failure rates from 1.7% in the first year to over 8.6% in the three-year-old population for more than 100,000 consumer-grade ATA drives, and says vendors often quote yearly rates below 2%. The Backblaze Q2 2026 report gives 1.73% for the quarter (1,498 failures over 31,553,350 drive days, 354,415 drives) and 1.41% over the lifetime of its fleet (21,609 failures over 558,148,607 drive days, 353,041 drives). As my own arithmetic, 1.41% is about 1.6 times 0.88%. That is a different era and drive mix, counted under Backblaze's own rules, so it is not a replication of the 3.4 times figure. The Google paper warns that its age groups mix drive models, so its age pattern is not a pure aging effect; the age finding here therefore rests on one study.
The medium behind these rates is thehard disk drivecard. A related ratio table on SMART warning signals is in theGoogle FAST 2007 analysis.
Sources
- 01Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX HTML proceedings, pages 1 to 16, Best Paper) · accessed 2026-10-09
- 02Disk Failures in the Real World, USENIX FAST 2007 abstract page · accessed 2026-10-09
- 03Analyzing Failure Data, CMU Petascale Data Storage Institute project page by the paper's authors (summary of the FAST 2007 results) · accessed 2026-10-10
- 04Failure Trends in a Large Disk Drive Population, Pinheiro, Weber, and Barroso, FAST 2007 (USENIX HTML proceedings) · accessed 2026-10-09
- 05Backblaze Drive Stats for Q2 2026, published September 29, 2026 · accessed 2026-10-09