Sector errors predicted a week ahead let a scrubber find them 1.7 to 1.8 times faster on hard disks, using accelerated scrubs 2% of the time
Published 2026-10-10
In Mahdisoltani et al. (USENIX ATC 2017), random-forest classifiers trained on drive monitoring data predicted a week ahead whether a disk would develop sector errors. For two hard disk models from Backblaze's public data, limiting false alarms to 10% of error-free weeks caught 90% and 95% of the weeks with errors, and limiting false alarms to 2% still caught 70% to 90%. Three Google SSD models were harder: 50% to 70% at 10% false alarms. When a simulated scrubber doubled its speed after a prediction and did so for at most 2% of the time, it found errors 1.7 to 1.8 times faster on hard disks and 1.4 to 1.5 times faster on two of the SSD models. The scores are weekly interval predictions on a held-out quarter of the data, and the text I read does not say the split was by time. Bairavasundaram et al. (2007) and Pinheiro et al. (2007) confirm that sector errors are common and that scrubbing finds most of them.
- Hard disk data
- Backblaze public SMART data, 7 models of 2,719 to 36,368 drives; more than a billion device hours in the full data set
- SSD data
- About 30,000 Google MLC SSDs in 3 models; uncorrectable errors and bad blocks per drive-week
- Prediction target
- Whether a drive shows a sector error in the next one-week interval
- Scoring
- False positive rate and false negative rate, 75% training and 25% testing
- Cross-checks
- NetApp, 1.53 million disks over 32 months (SIGMETRICS 2007); Google hard disks (FAST 2007)
Numbers
| Measure | Value | Unit | Population or scope | Source |
|---|---|---|---|---|
| Hard disks with at least one reallocated sector (SMART 5), lowest and highest model | 0.24 and 25.15 | percent of drives | HGST HMS5C4040BLE640 (9,426 drives) and Seagate ST3000DM001 (4,707 drives) | ATC 2017 |
| Hard disks with at least one reallocated sector, Hitachi HDS722020ALA330 | 11.84 | percent of drives | 4,774 drives | ATC 2017 |
| SSDs with uncorrectable errors, three Google MLC models | 37.07 to 65.56 | percent of drives | About 10,000 drives per model | ATC 2017 |
| SSDs with bad blocks, three Google MLC models | 50.35 to 82.75 | percent of drives | About 10,000 drives per model | ATC 2017 |
| Error weeks caught at a 10% false positive rate, Hitachi and Seagate | 90 and 95 | percent of error intervals | Random forest, prediction one week ahead | ATC 2017 |
| Error weeks caught at a 2% false positive rate, hard disks | 70 to 90 | percent of error intervals | Same models | ATC 2017 |
| Error weeks caught at a 10% false positive rate, SSDs | 50 to 70 | percent of error intervals | Three Google MLC models, uncorrectable errors | ATC 2017 |
| Training data size that still worked | 10 to 90 | percent of the usual training set | Hitachi disk; quality hardly affected | ATC 2017 |
| Faster error detection by an accelerated scrubber, hard disks | 1.7 to 1.8 | times faster than a fixed-rate scrubber | Scrub speed doubled after a prediction; accelerated mode at most 2% of the time | ATC 2017 |
| Faster error detection by an accelerated scrubber, MLC-A and MLC-D | 1.4 to 1.5 | times faster than a fixed-rate scrubber | Same setting | ATC 2017 |
| Time in accelerated scrub mode | less than 2 | percent of the total time | Authors' summary of the simulation | ATC 2017 |
| Share of tested hard disk drive days affected by SMART 5, Seagate ST3000DM001 | 1.77 | percent of drive days | 4,707 drives | ATC 2017 |
| Disks that developed latent sector errors over 32 months | 3.45 | percent of disks | NetApp, 1.53 million nearline and enterprise disks | SIGMETRICS 2007 |
| Latent sector errors found by scrubbing | more than 60 | percent of errors | Same population | SIGMETRICS 2007 |
| Google hard disks with a reallocation count above zero | about 9 | percent of drives | More than 100,000 ATA disks, December 2005 to August 2006 | Google, FAST 2007 |
This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 12 are from Mahdisoltani et al. (ATC 2017). Rows 13 to 15 are cross-checks from other fleets and are different measures.
Method
The figures are copied from the USENIX PDF of Mahdisoltani, Stefanovici, and Schroeder, USENIX ATC 2017, not refit. The hard disk data are Backblaze's public daily SMART snapshots for 7 models of 2,719 to 36,368 drives each; the SSD data are a random sample of about 30,000 drives from three MLC models used in a Google field study (MLC-A, MLC-B, and MLC-D, 10,115 to 10,258 drives each). The authors cut each drive's history into one-week intervals and predict whether the next week will see an error (SMART 5 reallocated sectors for hard disks, uncorrectable errors for SSDs). They compare several machine learning methods, keep 20% to 60% of the training instances as error instances by undersampling, and split the data 75% training and 25% testing. They report the false positive rate (error-free intervals wrongly flagged) and the false negative rate (error intervals missed), and then simulate a scrubber that doubles its speed when an error is predicted. Two other sources check the surrounding numbers: Bairavasundaram et al., SIGMETRICS 2007 (NetApp, 1.53 million disks) and Pinheiro et al., FAST 2007 (Google hard disks). Neither tests prediction.
Limits
The hard disk results come from two models picked because they are common and have the highest error rates, and they are not an average over all drives. The prediction target is a weekly interval, not a disk, and what is predicted is whether a sector error happens in the week, not whether the disk fails; the authors say the false-positive rates would be too high for triggering drive replacement and are acceptable only for light actions such as a scrub. The paper says that 75% of data is used for training and 25% for testing, and the text I read does not say whether the split is random or by date. Another paper, Xu et al. (ATC 2018, covered in a separate note), found that random splits gave 91.64% against 36.50% when tested forward in time, so these accuracy figures may be optimistic for a deployed predictor. The scrubbing result is a simulation with a scrubber that alternates between two speeds and predictions made once a week; no deployment is reported. The SSD data are a sample from drives of Google's own field study (Schroeder et al., FAST 2016), so the SSD rows are not an independent check of that study. Backblaze reports SMART values that vary by model and manufacturer, so not every model reports every attribute, and the hard disk models are older consumer drives.
What the paper found
The answer. Sector errors can be predicted a week ahead well enough to steer a scrubber. For the Hitachi HDS722020ALA330 and the Seagate ST3000DM001, random forests caught 90% and 95% of error weeks when 10% of error-free weeks were flagged by mistake, and 70% to 90% when only 2% were. The authors say the same prediction does not just restate SMART 197 (unstable sectors): removing it as an input gave the same results. Random forests matched or beat the other methods, including neural networks and support vector machines that needed much more tuning. Results for SMART 187 and SMART 197 were similar but slightly lower.
SSDs and little data. For the three Google SSD models, random forests caught 50% to 70% of uncorrectable-error weeks at a 10% false-positive rate, lower than for hard disks. Training on only 10% to 90% of the usual training data hardly changed the quality for the Hitachi disk. A predictor trained on other models still worked, with some loss: forests trained on MLC-B and MLC-D predicted MLC-A, and a forest trained on the Hitachi disk predicted errors for the Seagate disk.
The scrubbing use. A scrubber that reads data in the background finds latent errors, and a slower scrub means a longer window of vulnerability to data loss. In the authors' simulation, doubling the scrub speed after a prediction and keeping that mode to at most 2% of the time shortened the mean time to detect an error by 1.7 to 1.8 times for hard disks and 1.4 to 1.5 times for two of the SSD models. Their summary figure is nearly a factor of 2 with accelerated scrubbing less than 2% of the time.
What other fleets say. The paper's hard disk shares (0.24% to 25.15% of drives with a reallocated sector) are higher than the 3.45% of 1.53 million NetApp disks that developed latent sector errors over 32 months in Bairavasundaram et al., but that study's nearline models ranged from 5% to 20% at 24 months, and Mahdisoltani et al. call their numbers in line with those. Bairavasundaram et al. also found that more than 60% of the errors were discovered by scrubbing, which is the premise of the scrub-rate idea. Pinheiro et al. found about 9% of Google's hard disks with a nonzero reallocation count.
Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.
Sources
- 01Improving Storage System Reliability with Proactive Error Prediction, Mahdisoltani, Stefanovici, and Schroeder, USENIX ATC 2017 (USENIX PDF) · accessed 2026-10-10
- 02An Analysis of Latent Sector Errors in Disk Drives, Bairavasundaram, Goodson, Pasupathy, and Schindler, SIGMETRICS 2007 (University of Wisconsin ADSL PDF) · accessed 2026-10-10
- 03Failure Trends in a Large Disk Drive Population, Pinheiro, Weber, and Barroso, FAST 2007 (USENIX HTML proceedings) · accessed 2026-10-09