Skip to content

Paper note

Sector errors predicted a week ahead let a scrubber find them 1.7 to 1.8 times faster on hard disks, using accelerated scrubs 2% of the time

Published 2026-10-10

In Mahdisoltani et al. (USENIX ATC 2017), random-forest classifiers trained on drive monitoring data predicted a week ahead whether a disk would develop sector errors. For two hard disk models from Backblaze's public data, limiting false alarms to 10% of error-free weeks caught 90% and 95% of the weeks with errors, and limiting false alarms to 2% still caught 70% to 90%. Three Google SSD models were harder: 50% to 70% at 10% false alarms. When a simulated scrubber doubled its speed after a prediction and did so for at most 2% of the time, it found errors 1.7 to 1.8 times faster on hard disks and 1.4 to 1.5 times faster on two of the SSD models. The scores are weekly interval predictions on a held-out quarter of the data, and the text I read does not say the split was by time. Bairavasundaram et al. (2007) and Pinheiro et al. (2007) confirm that sector errors are common and that scrubbing finds most of them.

Hard disk data
Backblaze public SMART data, 7 models of 2,719 to 36,368 drives; more than a billion device hours in the full data set
SSD data
About 30,000 Google MLC SSDs in 3 models; uncorrectable errors and bad blocks per drive-week
Prediction target
Whether a drive shows a sector error in the next one-week interval
Scoring
False positive rate and false negative rate, 75% training and 25% testing
Cross-checks
NetApp, 1.53 million disks over 32 months (SIGMETRICS 2007); Google hard disks (FAST 2007)

Numbers

Sector-error prediction as printed in Mahdisoltani et al., USENIX ATC 2017 (Backblaze hard disks, Google SSDs), rows 1 to 12, and in two cross-check sources, rows 13 to 15. Rates for prediction are per one-week interval. FPR is the share of error-free intervals wrongly flagged. Qualifiers such as 'nearly' and 'about' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
Hard disks with at least one reallocated sector (SMART 5), lowest and highest model0.24 and 25.15percent of drivesHGST HMS5C4040BLE640 (9,426 drives) and Seagate ST3000DM001 (4,707 drives)ATC 2017
Hard disks with at least one reallocated sector, Hitachi HDS722020ALA33011.84percent of drives4,774 drivesATC 2017
SSDs with uncorrectable errors, three Google MLC models37.07 to 65.56percent of drivesAbout 10,000 drives per modelATC 2017
SSDs with bad blocks, three Google MLC models50.35 to 82.75percent of drivesAbout 10,000 drives per modelATC 2017
Error weeks caught at a 10% false positive rate, Hitachi and Seagate90 and 95percent of error intervalsRandom forest, prediction one week aheadATC 2017
Error weeks caught at a 2% false positive rate, hard disks70 to 90percent of error intervalsSame modelsATC 2017
Error weeks caught at a 10% false positive rate, SSDs50 to 70percent of error intervalsThree Google MLC models, uncorrectable errorsATC 2017
Training data size that still worked10 to 90percent of the usual training setHitachi disk; quality hardly affectedATC 2017
Faster error detection by an accelerated scrubber, hard disks1.7 to 1.8times faster than a fixed-rate scrubberScrub speed doubled after a prediction; accelerated mode at most 2% of the timeATC 2017
Faster error detection by an accelerated scrubber, MLC-A and MLC-D1.4 to 1.5times faster than a fixed-rate scrubberSame settingATC 2017
Time in accelerated scrub modeless than 2percent of the total timeAuthors' summary of the simulationATC 2017
Share of tested hard disk drive days affected by SMART 5, Seagate ST3000DM0011.77percent of drive days4,707 drivesATC 2017
Disks that developed latent sector errors over 32 months3.45percent of disksNetApp, 1.53 million nearline and enterprise disksSIGMETRICS 2007
Latent sector errors found by scrubbingmore than 60percent of errorsSame populationSIGMETRICS 2007
Google hard disks with a reallocation count above zeroabout 9percent of drivesMore than 100,000 ATA disks, December 2005 to August 2006Google, FAST 2007

This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 12 are from Mahdisoltani et al. (ATC 2017). Rows 13 to 15 are cross-checks from other fleets and are different measures.

Method

The figures are copied from the USENIX PDF of Mahdisoltani, Stefanovici, and Schroeder, USENIX ATC 2017, not refit. The hard disk data are Backblaze's public daily SMART snapshots for 7 models of 2,719 to 36,368 drives each; the SSD data are a random sample of about 30,000 drives from three MLC models used in a Google field study (MLC-A, MLC-B, and MLC-D, 10,115 to 10,258 drives each). The authors cut each drive's history into one-week intervals and predict whether the next week will see an error (SMART 5 reallocated sectors for hard disks, uncorrectable errors for SSDs). They compare several machine learning methods, keep 20% to 60% of the training instances as error instances by undersampling, and split the data 75% training and 25% testing. They report the false positive rate (error-free intervals wrongly flagged) and the false negative rate (error intervals missed), and then simulate a scrubber that doubles its speed when an error is predicted. Two other sources check the surrounding numbers: Bairavasundaram et al., SIGMETRICS 2007 (NetApp, 1.53 million disks) and Pinheiro et al., FAST 2007 (Google hard disks). Neither tests prediction.

Limits

The hard disk results come from two models picked because they are common and have the highest error rates, and they are not an average over all drives. The prediction target is a weekly interval, not a disk, and what is predicted is whether a sector error happens in the week, not whether the disk fails; the authors say the false-positive rates would be too high for triggering drive replacement and are acceptable only for light actions such as a scrub. The paper says that 75% of data is used for training and 25% for testing, and the text I read does not say whether the split is random or by date. Another paper, Xu et al. (ATC 2018, covered in a separate note), found that random splits gave 91.64% against 36.50% when tested forward in time, so these accuracy figures may be optimistic for a deployed predictor. The scrubbing result is a simulation with a scrubber that alternates between two speeds and predictions made once a week; no deployment is reported. The SSD data are a sample from drives of Google's own field study (Schroeder et al., FAST 2016), so the SSD rows are not an independent check of that study. Backblaze reports SMART values that vary by model and manufacturer, so not every model reports every attribute, and the hard disk models are older consumer drives.

What the paper found

The answer. Sector errors can be predicted a week ahead well enough to steer a scrubber. For the Hitachi HDS722020ALA330 and the Seagate ST3000DM001, random forests caught 90% and 95% of error weeks when 10% of error-free weeks were flagged by mistake, and 70% to 90% when only 2% were. The authors say the same prediction does not just restate SMART 197 (unstable sectors): removing it as an input gave the same results. Random forests matched or beat the other methods, including neural networks and support vector machines that needed much more tuning. Results for SMART 187 and SMART 197 were similar but slightly lower.

SSDs and little data. For the three Google SSD models, random forests caught 50% to 70% of uncorrectable-error weeks at a 10% false-positive rate, lower than for hard disks. Training on only 10% to 90% of the usual training data hardly changed the quality for the Hitachi disk. A predictor trained on other models still worked, with some loss: forests trained on MLC-B and MLC-D predicted MLC-A, and a forest trained on the Hitachi disk predicted errors for the Seagate disk.

The scrubbing use. A scrubber that reads data in the background finds latent errors, and a slower scrub means a longer window of vulnerability to data loss. In the authors' simulation, doubling the scrub speed after a prediction and keeping that mode to at most 2% of the time shortened the mean time to detect an error by 1.7 to 1.8 times for hard disks and 1.4 to 1.5 times for two of the SSD models. Their summary figure is nearly a factor of 2 with accelerated scrubbing less than 2% of the time.

What other fleets say. The paper's hard disk shares (0.24% to 25.15% of drives with a reallocated sector) are higher than the 3.45% of 1.53 million NetApp disks that developed latent sector errors over 32 months in Bairavasundaram et al., but that study's nearline models ranged from 5% to 20% at 24 months, and Mahdisoltani et al. call their numbers in line with those. Bairavasundaram et al. also found that more than 60% of the errors were discovered by scrubbing, which is the premise of the scrub-rate idea. Pinheiro et al. found about 9% of Google's hard disks with a nonzero reallocation count.

Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.

Sources

  1. 01Improving Storage System Reliability with Proactive Error Prediction, Mahdisoltani, Stefanovici, and Schroeder, USENIX ATC 2017 (USENIX PDF) · accessed 2026-10-10
  2. 02An Analysis of Latent Sector Errors in Disk Drives, Bairavasundaram, Goodson, Pasupathy, and Schindler, SIGMETRICS 2007 (University of Wisconsin ADSL PDF) · accessed 2026-10-10
  3. 03Failure Trends in a Large Disk Drive Population, Pinheiro, Weber, and Barroso, FAST 2007 (USENIX HTML proceedings) · accessed 2026-10-09