Skip to content

Paper note

Adding performance and location data lifted 10-day disk failure prediction to 0.95 MCC across 380,000 disks

Published 2026-10-10

In Lu et al. (FAST 2020), models fed only SMART attributes missed many failing disks, and adding performance metrics and rack location raised the best model, a CNN-LSTM, to 0.95 MCC and 0.95 F-measure for a 10-day prediction horizon on 380,000 hard disks from 64 sites over about two months, against 0.77 MCC for the next best method, a random forest. Location alone helped by less than 10% in MCC, and only when performance data were present. The result comes from one operator's data center and one short window, so it is a measure of what is possible there, not a general error rate. Backblaze (2016) and Google (FAST 2007) both report that a large share of failed drives show no SMART warning, which is the same gap from other fleets.

Population
380,000 hard disks, 10,000 server racks, 64 data center sites of one operator
Window
About 70 days; SMART collected once a day; 10-day prediction horizon
Failure definition
The operator decides a disk must be replaced or repaired after a failed read or write and a bad restart
Features
14 SMART attributes (more than 97% of disks report them), performance metrics, and rack location
Cross-checks
Backblaze (October 2016, 67,814 drives) and Google (FAST 2007, more than 100,000 ATA disks)

Numbers

Disk failure prediction as printed in Lu et al., FAST 2020 (380,000 hard disks, 64 sites, about 70 days, 10-day horizon), rows 1 to 10, and in two cross-check sources about SMART warnings, rows 11 to 13. MCC is the Matthews correlation coefficient, from -1 to 1. Qualifiers such as 'up to' and 'roughly' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
Hard disks in the study380,000disks64 sites, 10,000 server racks, one operator housing more than two million disksFAST 2020
Observation windowabout 70daysSMART collected once a dayFAST 2020
Prediction horizon10daysChosen so operators have time to actFAST 2020
Best model, MCC with SMART, performance, and location data0.95MCC scoreCNN-LSTM, SPL feature groupFAST 2020
Best model, F-measure0.95F-measurePrinted as 'up to' and 'on average' for a 10-day horizonFAST 2020
Next best method, MCC, same data0.77MCC scoreRandom forest, SPL feature groupFAST 2020
Gain from adding location informationless than 10percent of MCCCNN-LSTM; helps mainly with performance features presentFAST 2020
MCC on an unseen site after training on 62 sitesabove 0.90MCC scoreCNN-LSTM, SPL group, 10-day horizon, site AFAST 2020
Drop for RF and GBDT on an unseen sitemore than 15percent, in some casesSame testFAST 2020
Training time of the CNN-LSTMup to 4hours per training runSame studyFAST 2020
Failed drives with one or more of five SMART counts above zero76.7percent of failed drives67,814 drives in 2016; 23.3% showed none; operational drives 4.2%Backblaze 2016
Operational drives with one or more of five SMART counts above zero4.2percent of operational drivesSame fleetBackblaze 2016
Failed drives with no count on four strong SMART signalsover 56percent of failed drivesMore than 100,000 ATA disks, December 2005 to August 2006Google, FAST 2007

This table is a short extract of printed figures, not a copy of the papers and not the data set. Rows 1 to 10 are from Lu et al. (FAST 2020). Rows 11 to 13 are from the two cross-check sources and measure how often failed drives had a SMART warning, which is a different quantity.

Method

The figures are copied from the USENIX PDF of Lu, Luo, Patel, Yao, Tiwari, and Shi, FAST 2020 (pages 151 to 167), not refit. The data cover 380,000 hard disks across 64 data center sites and 10,000 server racks for roughly 70 days, from one large data center operator that houses more than two million disks; only disks with SMART, performance, and location logs are included. A disk is counted as failed when the operator decides it must be replaced or repaired because a read or write failed and the disk cannot work properly after a restart. The authors compare Bayes, random forest (RF), gradient-boosted trees (GBDT), LSTM, and a CNN-LSTM they propose, on six feature groups drawn from SMART (S), performance (P), and location (L) data. They score a disk as failing or not within a 10-day horizon with F-measure and the Matthews correlation coefficient (MCC), because healthy disks far outnumber failed ones. Two other sources check the SMART gap only: Backblaze's October 2016 post and Pinheiro, Weber, and Barroso, FAST 2007. Neither tests machine learning models or the 0.95 figure.

Limits

The 0.95 scores are the paper's own, from one operator, about 70 days of data, and a 10-day horizon, with the model trained and tested on the same operator's disks; the number of failed disks is not stated in the text I read, and the headline scores are read from the paper's abstract, introduction, and evaluation prose, not from the charts. The abstract says 0.95 on average and the introduction says up to 0.95, and the scores differ by model and feature group. The paper prints '2.6 million device hours' for 380,000 disks over roughly 70 days, which does not match that many disks and days, and this note repeats the printed number only as a caution. Fewer failures were predicted correctly in racks with few failures, because there were too few failed samples to learn from. Portability across sites held when training on 62 sites and testing on an unseen one (above 0.90 MCC for CNN-LSTM), but RF and GBDT dropped by more than 15% in some cases, and training on one site and testing on another dropped MCC significantly. The paper says SMART attributes are of limited use only for 'hard-to-predict' failures and does not claim every failed disk shows a change in performance. The cross-check sources do not use machine learning: Backblaze describes five attributes in 2016, and Google's population is consumer-grade ATA disks from 2005 and 2006.

What the paper found

The answer. For a 10-day horizon the best model in Lu et al., a CNN-LSTM using SMART, performance, and location data together (the SPL group), scored 0.95 MCC and 0.95 F-measure, against 0.77 MCC for the next best method, a random forest, with the same data. Models using SMART alone had a very high false-negative rate, meaning failing disks were predicted healthy, and adding performance and location data cut it significantly. The authors argue that SMART values often change only a few hours before a failure, so a prediction several days ahead needs signals that move earlier.

What each kind of data added. Performance metrics were the larger addition. Location helped less: adding it improved quality across models, but by less than 10% in MCC for the CNN-LSTM, and it mattered only when performance features were present. The authors explain that disks in the same rack share temperature, humidity, and vibration. They also show that SMART values of a failed disk can look like those of healthy disks on the same server until the failure, and they show the same flat pattern in a public Baidu data set sampled hourly, where a failed disk's 12 normalized SMART attributes did not vary noticeably in the 477 hours before the failure.

Model choice and trade-offs. Without performance and location data, simpler tree models (RF and GBDT) did roughly as well as the neural networks and sometimes beat the LSTM, and the CNN-LSTM took up to four hours to train once. The authors also describe a trade-off between false positives (replacing healthy disks) and false negatives (missing a failure), which the operator should set by cost: GBDT gave a lower false-positive rate and a higher false-negative rate for the SPL group. Mispredictions clustered in racks with few failures, and early in the test period the models predicted too many disks as healthy for lack of training data.

What other fleets say about the SMART gap. Backblaze's 2016 post reports that 76.7% of failed drives, against 4.2% of operational drives, had one or more of five SMART counts above zero, so 23.3% of failed drives showed no warning in those counts. Google's FAST 2007 population found over 56% of failed drives with no count on any of four strong SMART signals. Those fleets and failure definitions differ from the operator in Lu et al., and neither tests a prediction model, but they support the paper's starting point that SMART alone is an incomplete warning.

Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.

Sources

  1. 01Making Disk Failure Predictions SMARTer!, Lu, Luo, Patel, Yao, Tiwari, and Shi, FAST 2020 (USENIX PDF, pages 151 to 167) · accessed 2026-10-10
  2. 02Making Disk Failure Predictions SMARTer!, USENIX FAST 2020 presentation page (abstract and citation) · accessed 2026-10-10
  3. 03What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 · accessed 2026-10-10
  4. 04Failure Trends in a Large Disk Drive Population, Pinheiro, Weber, and Barroso, FAST 2007 (USENIX HTML proceedings) · accessed 2026-10-09