Skip to content

Stat analysis

Over 56% of failed drives had no count on the four strong SMART signals

Published 2026-10-09

Over 56% of failed drives in the Google production population of more than 100,000 consumer-grade serial and parallel ATA drives (December 2005–August 2006) had no count in any of the four strong SMART signals (scan errors, reallocation count, offline reallocation, and probational count), so models based only on those signals can never predict more than half of the failed drives.

Population
More than 100,000 consumer-grade serial and parallel ATA drives in Google production service
Speed and capacity
5400–7200 rpm and 80–400 GB
Commissioned
In or after 2001, at least nine models; manufacturer, model, and vintage mixes withheld
SMART window
December 2005 through August 2006, nine months
Failure
Replaced during a repair, which can lag the observed event by a few days; not a manufacturer confirmation that the drive was defective

Comparison

SMART signals and failed-drive coverage for more than 100,000 consumer-grade serial and parallel ATA drives in Google production service (5400–7200 rpm, 80–400 GB, commissioned in or after 2001, at least nine models; manufacturer, model, and vintage withheld). SMART counts were collected from December 2005 through August 2006. Failure means the drive was replaced during a repair. The 60-day column is the paper’s ratio versus a zero count after the first event. A critical threshold is reported only when that ratio is at least 10 times and confidence is above 95%. “not stated” means the paper does not give the figure. “does not apply” means the column is a different measure. Qualifiers such as “fewer than,” “about,” and “over” are the paper’s wording.
SMART checkNon-zero count (percent of drives)Failed drives at zero (percent of failed drives)60-day failure likelihood vs zero count (times)Annualized failure rate vs zero count (times)Critical threshold (count)Source
Scan errorsfewer than 2not stated39101FAST 2007
Reallocation countabout 9not statedover 143–61FAST 2007
Offline reallocationabout 4not statedover 21not statednot statedFAST 2007
Probational countabout 2not stated16not stated1FAST 2007
Scan errors, reallocation count, offline reallocation, and probational count all at zerodoes not applyover 56does not applydoes not applydoes not applyFAST 2007
Every other SMART parameter except temperature at zerodoes not applyover 36does not applydoes not applydoes not applyFAST 2007

This table is a short extract of reported ratios, not a copy of the paper and not the underlying records. The canonical text is the USENIX HTML proceedings.

Method

The ratios are copied from Pinheiro, Weber, and Barroso, FAST 2007, not refit on another fleet. For each SMART parameter they looked for a count that raised the probability of failure in the next 60 days by at least 10 times versus drives with a zero count for that parameter, and they reported a critical threshold only above 95% confidence. Annualized failure rate is a separate comparison of the non-zero group with the zero group. The coverage rows are the share of failed drives whose counts stayed at zero: an upper bound on what those signals can flag, not a false-positive rate and not the accuracy of a deployed model. The authors state that these parameters were not used in repair diagnostics when the data were collected, so a repair was not defined as a threshold on these counts. Replacement time can still lag the observed event by a few days.

Limits

The SMART sample is the population in the scope list, collected from December 2005 through August 2006, a nine-month window. Drives were powered on for essentially all of their recorded life in rack-mounted servers. Burn-in fallout is excluded. Negative counts and impossible values were dropped, which reduced the sample by less than 0.1%. The paper does not publish the drive records or a data license, and it withholds manufacturer, model, and vintage mixes as proprietary. Absolute 60-day failure probabilities are not stated, only ratios against zero-count drives. The age-group annualized failure rates, 1.7% in the first year of operation and over 8.6% in the three-year-old population, come from a separate repairs database of about five years. Those age groups mix different models, so they are not rates from the SMART window. The USENIX HTML is labeled pages 17–28; the Google Research record says pages 17–29. Not measured: per-model counts, vibration, a crisp temperature-error threshold, the false-positive rate of any predictor, or any fleet other than this Google population. An arbitrary check in the paper — more than half of observed time above 40°C — still left about 36% of all drives with no failure signal. That sentence says “all drives,” not “failed drives,” so it is not in the failed-drive column. Drive manufacturers often quote yearly failure rates below 2%. That is a vendor claim the paper records, not a measurement of these drives.

What the signal catches, and what it misses

Scan errors are counts from a background pass over the platter. The paper treats a large count as a sign of surface defects. Fewer than 2% of the drives showed any, nearly uniformly across models. That group’s annualized failure rate was 10 times the rate of drives with none. After the first scan error, a drive was 39 times more likely to fail within 60 days than a drive with no scan errors, and the critical threshold is one. A little over 70% of drives still survived the first 8 months after that first error. The 39-times figure is a ratio against quiet drives, not the percentage of drives that fail, and the signal is rare.

A reallocation count is how many times the drive remapped a sector it believed was damaged, typically after recurring soft errors or a hard error, onto a spare. About 9% of the population had a count above zero. The average impact on annualized failure rate was a factor of 3–6, and after the first reallocation a drive was over 14 times more likely to fail within 60 days than a drive with none. The critical threshold is one. About 85% of drives survived past 8 months after the first reallocation.

Offline reallocation is defined as the subset found during background scrubbing. It is meant to exclude sectors remapped because of errors on actual I/O. About 4% of drives had a non-zero count, concentrated on a subset of models, and after the first offline reallocation the 60-day failure chance was over 21 times the chance for drives with none. Some models reported more offline reallocations than total reallocations, so the counter is not a clean subset; the paper says to interpret it within specific models and does not name a critical threshold in that section. Probational counts are suspect sectors that may later be reallocated or may keep working. About 2% of drives had a non-zero probational count, lower than the reallocation shares, which the paper reads as sectors leaving probation after further observation. After the first event a drive was 16 times more likely to fail within 60 days than a drive at zero, and the critical threshold is one.

The fault those four signals miss is a failed drive that never incremented them. Over 56% of failed drives had no count in scan errors, reallocation count, offline reallocation, or probational count. Adding every other SMART parameter except temperature still left over 36% of failed drives at zero on all counts. The USENIX text concludes that SMART data alone are unlikely to predict failures of individual drives; section 5 says such models are likely to be severely limited because a large fraction of failed drives showed no SMART error signals. The Google Research abstract states the same conclusion. Seek errors do not close the gap: more than 72% of drives had them, they were widespread in one manufacturer only, and the other manufacturers showed no correlation with failure. CRC errors appeared on about 2% of drives and are described as less indicative of the drive than of cables and connectors.

The medium these counts come from is thehard disk drivecard.

Sources

  1. 01Failure Trends in a Large Disk Drive Population, Pinheiro, Weber, and Barroso, FAST 2007 (USENIX HTML proceedings) · accessed 2026-10-09
  2. 02Failure Trends in a Large Disk Drive Population (USENIX FAST 2007 event HTML) · accessed 2026-10-09
  3. 03Failure Trends in a Large Disk Drive Population (Google Research record) · accessed 2026-10-09