Skip to content

Paper note

A 2016 SMART predictor reached 98% recall on one Seagate model but 81% on a Hitachi model, scored with random splits on Backblaze data

Published 2026-10-11

In Botezatu et al. (KDD 2016, IBM Research), a SMART-based predictor trained on Backblaze data from 50,984 disks over 17 months reached 98% recall for replaced disks of one Seagate model and 81% for one Hitachi model, against 53% and 44% for a simple decision tree, and found 92% of replaced Seagate disks 10 days ahead. The scores come from 100 random 80/20 splits of a downsampled set, not from a split by time, and only 2.5% to 3% of disks were replaced. Xu et al. (ATC 2018, Microsoft Azure) cite this kind of cross-validation and show a true positive rate of 91.64% falling to 36.50% when tested forward in time, and Backblaze's own 2016 post found 23.3% of failed drives showed no warning from five SMART counts, so the headline numbers are an upper bound for a deployed predictor.

Population
50,984 Backblaze disks; one Seagate and one Hitachi family modeled, with two sibling models for transfer learning
Window
17 months of daily SMART data, April 2013 to June 2015 collection, early months dropped
Label
A disk marked failed on the day before its replacement or removal
Evaluation
100 random splits of 80% training and 20% test; healthy class downsampled
Cross-checks
Azure disk errors, ATC 2018; Backblaze SMART post, 2016; a data center operator's disks, FAST 2020

Numbers

Disk replacement prediction figures as printed in Botezatu et al., KDD 2016 (Backblaze data, 50,984 disks), rows 1 to 10, and in three cross-check sources, rows 11 to 15. Qualifiers such as 'to' mark printed ranges.
MeasureValueUnitPopulation or scopeSource
Disks in the Backblaze dataset used50984disksBackblaze public daily SMART logs
Months of data kept17monthsApril 2013 to June 2015 collection; first months dropped
Replaced share of the two main models2.5 to 3percent of disksSeagate ST4000DM000 (SgtA) and Hitachi HDS722020ALA330 (HitA)
Error on the best Seagate model1 to 2percent over 100 runsSgtA, regularized greedy forest
Recall for replaced disks, Seagate model98percentSgtA, 100 random 80/20 splits
Recall for replaced disks, Hitachi model81percentHitA, 100 random 80/20 splits
Recall of a simple decision tree on a small set of SMART attributes53percentSgtA; 44% for HitA
Replaced Seagate disks predicted 10 days ahead92percentSgtA; 97% at 3 days
Replaced Seagate disks predicted 30 days ahead73percentSgtA; HitA 75% at 30 days
Evaluation repeats with random training and test splits100splits of 80% training and 20% testSame models
True positive rate with random cross-validation at a 0.1% false positive rate91.64percentMicrosoft Azure, Dataset 1
True positive rate with online prediction at the same false positive rate36.5percentSame dataset and model, training before testing in time
Failed drives with at least one of five SMART counts above zero76.7percent of failed drivesBackblaze, 67,814 drives, 2016
Operational drives with at least one of the five SMART counts above zero4.2percent of operational drivesSame fleet
Matthews correlation coefficient with 5-fold cross-validation0.95correlation coefficientdisks of a leading data center operator, 10-day lead time

This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 10 are from Botezatu et al. (KDD 2016). Rows 11 and 12 are from Xu et al. (ATC 2018), rows 13 and 14 from Backblaze (2016), and row 15 from Lu et al. (FAST 2020); they are cross-checks with different measures.

Method

The figures are copied from the KDD PDF of Botezatu, Giurgiu, Bogojeska, and Wiesmann, ACM SIGKDD 2016, not refit. The data are Backblaze's public daily SMART logs for 50,984 disks, collected from April 2013 to June 2015 with the first months dropped, leaving 17 months. The authors keep one Seagate and one Hitachi family, pick SMART attributes by detecting changepoints, summarize each disk's recent history with exponential smoothing, downsample the healthy disks with k-means to 1,000 (Seagate) or 500 (Hitachi), and fit a regularized greedy forest. They score it with precision, recall, and F-score over 100 random splits of 80% training and 20% test. Three other sources check the surrounding claims: Xu et al., ATC 2018 (Microsoft Azure disk errors, time-forward testing), Backblaze's 2016 post on SMART counts in failed and operational drives, and Lu et al., FAST 2020 (disks of a leading data center operator). None re-runs the Botezatu model.

Limits

The numbers are for two disk families from one public fleet in 2013 to 2015, and the authors say a separate model is needed for each manufacturer. The scores come from random 80/20 splits of data that the authors downsampled to balance healthy and replaced disks; the paper does not say the test set keeps the true 2.5% to 3% share of replaced disks, so the precision values would be lower on a real fleet, and the paper's own false-alarm claim is not tested at fleet scale. Random splits can place a disk's earlier and later records on both sides, and Xu et al. show on Azure data that this inflates results (91.64% against 36.50% true positive rate), though their data and model differ. Backblaze's post is the same fleet's own analysis, so it adds a view of SMART counts but not an independent dataset. Lu et al. also use a random 5-fold partition. A replacement is not always a fail-stop failure. The 14 to 19% Hitachi gap is the authors' own summary and the paper ties it to fewer disks and fewer usable SMART values.

What the paper found

The answer. In Botezatu et al., SMART-based prediction of disk replacement worked well on one Seagate model and less well on one Hitachi model, and the scores are an upper bound. On Backblaze data for 50,984 disks, the predictor reached 98% recall for replaced Seagate ST4000DM000 disks (error of 1% to 2% over 100 runs) and 81% for the Hitachi HDS722020ALA330, against 53% and 44% for a simple decision tree on a small attribute subset. The Hitachi scores were 14% to 19% lower, which the authors tie to fewer disks and fewer usable SMART values. Only 2.5% to 3% of these disks were replaced, and the authors downsampled the healthy class to 1,000 (Seagate) and 500 (Hitachi) before training.

Lead time. Using snapshots taken before the replacement, the model flagged 97% of replaced Seagate disks 3 days ahead, 92% 10 days ahead, and 73% 30 days ahead, and 84% of replaced Hitachi disks 3 days ahead and 75% 30 days ahead. The authors read the length of useful history from the SMART attribute: about 12 days for reallocated sectors, 10 days for pending sectors, 4 days for the read error rate, and 25 days for seek and transfer error rates.

Why the scores are optimistic. The authors score the model with 100 random 80/20 splits. Xu et al. (ATC 2018) name Botezatu et al. among the studies that use cross-validation and show on Azure data that a true positive rate of 91.64% at a 0.1% false positive rate with cross-validation falls to 36.50% when the model is trained on the past and tested on the future, because random splits let future and environment-specific information leak into training. Lu et al. (FAST 2020) also use a random 5-fold partition for their 0.95 MCC on a data center operator's disks. The data and models differ, so the sizes are not comparable, but the direction is the same: a time-forward test is the fair one.

What SMART counts alone miss. In the same Backblaze fleet's own 2016 post, 76.7% of failed drives had at least one of five SMART counts above zero, against 4.2% of operational drives, so 23.3% of failed drives showed no warning from those counts. That is consistent with Botezatu et al.'s finding that the useful signal comes from combining several attributes over time, and with the limits of any model built on SMART alone.

Background on drive types is in the drive technology index. Warning signs before a failure are also covered in the SMART failure signals analysis.

Sources

  1. 01Predicting Disk Replacement towards Reliable Data Centers, Botezatu, Giurgiu, Bogojeska, and Wiesmann, ACM SIGKDD 2016, pages 39 to 48 (KDD PDF) · accessed 2026-10-11
  2. 02Improving Service Availability of Cloud Systems by Predicting Disk Error, Xu et al., USENIX ATC 2018 (USENIX PDF, section 3 on online prediction) · accessed 2026-10-11
  3. 03What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 · accessed 2026-10-11
  4. 04Making Disk Failure Predictions SMARTer!, Lu et al., USENIX FAST 2020 (USENIX PDF, section 4) · accessed 2026-10-11