A 2016 SMART predictor reached 98% recall on one Seagate model but 81% on a Hitachi model, scored with random splits on Backblaze data
Published 2026-10-11
In Botezatu et al. (KDD 2016, IBM Research), a SMART-based predictor trained on Backblaze data from 50,984 disks over 17 months reached 98% recall for replaced disks of one Seagate model and 81% for one Hitachi model, against 53% and 44% for a simple decision tree, and found 92% of replaced Seagate disks 10 days ahead. The scores come from 100 random 80/20 splits of a downsampled set, not from a split by time, and only 2.5% to 3% of disks were replaced. Xu et al. (ATC 2018, Microsoft Azure) cite this kind of cross-validation and show a true positive rate of 91.64% falling to 36.50% when tested forward in time, and Backblaze's own 2016 post found 23.3% of failed drives showed no warning from five SMART counts, so the headline numbers are an upper bound for a deployed predictor.
- Population
- 50,984 Backblaze disks; one Seagate and one Hitachi family modeled, with two sibling models for transfer learning
- Window
- 17 months of daily SMART data, April 2013 to June 2015 collection, early months dropped
- Label
- A disk marked failed on the day before its replacement or removal
- Evaluation
- 100 random splits of 80% training and 20% test; healthy class downsampled
- Cross-checks
- Azure disk errors, ATC 2018; Backblaze SMART post, 2016; a data center operator's disks, FAST 2020
Numbers
This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 10 are from Botezatu et al. (KDD 2016). Rows 11 and 12 are from Xu et al. (ATC 2018), rows 13 and 14 from Backblaze (2016), and row 15 from Lu et al. (FAST 2020); they are cross-checks with different measures.
Method
The figures are copied from the KDD PDF of Botezatu, Giurgiu, Bogojeska, and Wiesmann, ACM SIGKDD 2016, not refit. The data are Backblaze's public daily SMART logs for 50,984 disks, collected from April 2013 to June 2015 with the first months dropped, leaving 17 months. The authors keep one Seagate and one Hitachi family, pick SMART attributes by detecting changepoints, summarize each disk's recent history with exponential smoothing, downsample the healthy disks with k-means to 1,000 (Seagate) or 500 (Hitachi), and fit a regularized greedy forest. They score it with precision, recall, and F-score over 100 random splits of 80% training and 20% test. Three other sources check the surrounding claims: Xu et al., ATC 2018 (Microsoft Azure disk errors, time-forward testing), Backblaze's 2016 post on SMART counts in failed and operational drives, and Lu et al., FAST 2020 (disks of a leading data center operator). None re-runs the Botezatu model.
Limits
The numbers are for two disk families from one public fleet in 2013 to 2015, and the authors say a separate model is needed for each manufacturer. The scores come from random 80/20 splits of data that the authors downsampled to balance healthy and replaced disks; the paper does not say the test set keeps the true 2.5% to 3% share of replaced disks, so the precision values would be lower on a real fleet, and the paper's own false-alarm claim is not tested at fleet scale. Random splits can place a disk's earlier and later records on both sides, and Xu et al. show on Azure data that this inflates results (91.64% against 36.50% true positive rate), though their data and model differ. Backblaze's post is the same fleet's own analysis, so it adds a view of SMART counts but not an independent dataset. Lu et al. also use a random 5-fold partition. A replacement is not always a fail-stop failure. The 14 to 19% Hitachi gap is the authors' own summary and the paper ties it to fewer disks and fewer usable SMART values.
What the paper found
The answer. In Botezatu et al., SMART-based prediction of disk replacement worked well on one Seagate model and less well on one Hitachi model, and the scores are an upper bound. On Backblaze data for 50,984 disks, the predictor reached 98% recall for replaced Seagate ST4000DM000 disks (error of 1% to 2% over 100 runs) and 81% for the Hitachi HDS722020ALA330, against 53% and 44% for a simple decision tree on a small attribute subset. The Hitachi scores were 14% to 19% lower, which the authors tie to fewer disks and fewer usable SMART values. Only 2.5% to 3% of these disks were replaced, and the authors downsampled the healthy class to 1,000 (Seagate) and 500 (Hitachi) before training.
Lead time. Using snapshots taken before the replacement, the model flagged 97% of replaced Seagate disks 3 days ahead, 92% 10 days ahead, and 73% 30 days ahead, and 84% of replaced Hitachi disks 3 days ahead and 75% 30 days ahead. The authors read the length of useful history from the SMART attribute: about 12 days for reallocated sectors, 10 days for pending sectors, 4 days for the read error rate, and 25 days for seek and transfer error rates.
Why the scores are optimistic. The authors score the model with 100 random 80/20 splits. Xu et al. (ATC 2018) name Botezatu et al. among the studies that use cross-validation and show on Azure data that a true positive rate of 91.64% at a 0.1% false positive rate with cross-validation falls to 36.50% when the model is trained on the past and tested on the future, because random splits let future and environment-specific information leak into training. Lu et al. (FAST 2020) also use a random 5-fold partition for their 0.95 MCC on a data center operator's disks. The data and models differ, so the sizes are not comparable, but the direction is the same: a time-forward test is the fair one.
What SMART counts alone miss. In the same Backblaze fleet's own 2016 post, 76.7% of failed drives had at least one of five SMART counts above zero, against 4.2% of operational drives, so 23.3% of failed drives showed no warning from those counts. That is consistent with Botezatu et al.'s finding that the useful signal comes from combining several attributes over time, and with the limits of any model built on SMART alone.
Background on drive types is in the drive technology index. Warning signs before a failure are also covered in the SMART failure signals analysis.
Sources
- 01Predicting Disk Replacement towards Reliable Data Centers, Botezatu, Giurgiu, Bogojeska, and Wiesmann, ACM SIGKDD 2016, pages 39 to 48 (KDD PDF) · accessed 2026-10-11
- 02Improving Service Availability of Cloud Systems by Predicting Disk Error, Xu et al., USENIX ATC 2018 (USENIX PDF, section 3 on online prediction) · accessed 2026-10-11
- 03What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 · accessed 2026-10-11
- 04Making Disk Failure Predictions SMARTer!, Lu et al., USENIX FAST 2020 (USENIX PDF, section 4) · accessed 2026-10-11