Skip to content

Paper note

Azure disk-error prediction caught 36.5% of faulty disks at a 0.1% false-positive rate, and random cross-validation said 91.6%

Published 2026-10-10

In Xu et al. (USENIX ATC 2018), Microsoft's disk-error predictor, CDEF, found 36.50% of faulty disks on one test set when it was allowed to wrongly flag only 0.1% of healthy disks (29.67% to 41.09% across three test sets), with about 10,000 healthy disks per 3 faulty ones. Scored with a random split, the same task gave 91.64%, which is why the authors insist on testing forward in time. Adding system-level signals such as OS events to SMART raised the average catch rate at that false-positive rate from 27.6% to 30.3%, and to 35.8% with feature selection. In production at Azure the authors report about 63,000 fewer minutes of VM downtime a month. The data are one cloud system in one month, and labels come from engineers' root-cause analysis of service problems.

System
One Microsoft cloud system; disks in clusters that host virtual machines, not RAID storage clusters
Window
Training on October 2017; testing on November 2017 in three test sets
Target
Disk errors found through root-cause analysis of service issues, not only whole-disk failure
Class balance
About 10,000 healthy disks to 3 faulty disks per test set; about 300 of 1,000,000 disks become faulty per day
Cross-checks
Lu et al., FAST 2020 (380,000 disks, 64 sites); Backblaze, October 2016 (67,814 drives)

Numbers

Disk-error prediction as printed in Xu et al., USENIX ATC 2018 (one Microsoft cloud system, training October 2017, testing November 2017), rows 1 to 12, and in two cross-check sources, rows 13 to 15. TPR is the share of faulty disks flagged; FPR is the share of healthy disks flagged wrongly. Qualifiers such as 'about' and 'around' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
Healthy to faulty disks in each test setabout 10,000 to 3ratioThree test sets from November 2017ATC 2018
Disks that become faulty per dayabout 300 of 1,000,000disksThe studied cloud systemATC 2018
CDEF true positive rate at 0.1% false positives, Dataset 136.50percent of faulty disksCost-sensitive ranking modelATC 2018
CDEF true positive rate at 0.1% false positives, Datasets 2 and 341.09 and 29.67percent of faulty disksSame modelATC 2018
Random forest true positive rate at 0.1% false positives, Datasets 1 to 330.51, 34.11, and 18.81percent of faulty disksSMART features onlyATC 2018
SVM true positive rate at 0.1% false positives, Datasets 1 to 315.51, 21.71, and 7.20percent of faulty disksSMART features onlyATC 2018
Area under the ROC curve on Dataset 1, CDEF, random forest, SVM0.93, 0.85, and 0.53AUCA random guess is 0.5ATC 2018
Average true positive rate with SMART features alone27.6percent of faulty disksAt 0.1% false positives, three datasetsATC 2018
Average true positive rate with SMART plus system-level signals30.3percent of faulty disksSame settingATC 2018
Average true positive rate with feature selection added35.8percent of faulty disksSame setting; this is CDEFATC 2018
True positive rate on Dataset 1 with cross-validation instead of a time split91.64percent of faulty disksAt 0.1% false positives; the time split gave 36.50ATC 2018
Average VM downtime saved after deploymentaround 63,000minutes per monthAzure; 99.999% availability allows 26 seconds per monthATC 2018
Failed drives with one or more of five SMART counts above zero76.7percent of failed drivesBackblaze, 67,814 drives in 2016; 23.3% showed noneBackblaze 2016
Operational drives with one or more of five SMART counts above zero4.2percent of operational drivesSame fleetBackblaze 2016
Best model on an unseen site, trained on 62 sitesabove 0.90MCC scoreLu et al., 380,000 disks; CNN-LSTM, 10-day horizon, with SMART, performance, and location dataFAST 2020

This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 12 are from Xu et al. (ATC 2018). Rows 13 to 15 are cross-checks from other fleets and are different measures.

Method

The figures are copied from the USENIX PDF of Xu, Sui, Yao, Zhang, Lin, Dang, Li, Jiang, Zhang, Lou, Chintalapati, and Zhang, USENIX ATC 2018 (pages 481 to 494), not refit. The authors predict disk errors (sector errors, latency errors, timeouts, and partition errors that show up as VM downtime), not only complete disk failure. Faulty disks are those found through root-cause analysis of service issues by field engineers. They train on October 2017 data from a Microsoft cloud system and test on November 2017, split into three test sets, and they compare their cost-sensitive ranking model (CDEF) with random forest and SVM baselines on SMART data. They report the true positive rate (TPR) at a false positive rate (FPR) of 0.1% and the area under the ROC curve (AUC). Two other sources check the surrounding claims: Lu et al., FAST 2020 (another operator) on SMART-only prediction, and a Backblaze post of 6 October 2016 on how often failed drives show SMART warnings. Neither reproduces the Azure numbers.

Limits

The authors write that the data come from one cloud service system of one company, so the results may not generalize. Labels come from engineers' root-cause analyses, and drives sent for repair may return, so some labels are noisy, which the authors acknowledge. They report only FPR, TPR, and AUC, and say precision and recall are future work, so the number of false alarms in absolute terms is not stated. By this note's arithmetic, with about 10,000 healthy disks per 3 faulty ones, a 0.1% false-positive rate is about 10 healthy disks wrongly flagged while a 36.5% catch rate finds about 1 of the 3 faulty ones, so most flagged disks would be healthy; the paper does not report precision. The 91.64% cross-validation figure is a warning about method, not a claim about performance. The 63,000 minutes of VM downtime saved per month is the authors' own average after deployment, and the text I read does not describe a control group. The model flags a disk by ranking and a cutoff, and predicted disks are live-migrated and stress-tested before replacement, so the actions are cheap compared with replacing a disk. The cross-check sources are from other fleets with other failure definitions and do not test this model.

What the paper found

The answer. At a false-positive rate of 0.1%, the Azure predictor CDEF caught 36.50%, 41.09%, and 29.67% of faulty disks on the three test sets, against 30.51%, 34.11%, and 18.81% for a random forest and 15.51%, 21.71%, and 7.20% for an SVM, both trained on SMART data. On Dataset 1 the area under the ROC curve was 0.93 for CDEF, 0.85 for the random forest, and 0.53 for the SVM. The authors note that the SVM has low cost because it rarely raises false alarms, and that it does poorly at finding faulty disks.

What helped. The authors' main addition was system-level signals, such as events from the operating system, on top of SMART. On average the catch rate at a 0.1% false-positive rate was 27.6% with SMART features alone, 30.3% with SMART plus system-level signals, and 35.8% when a feature selection step kept only stable, predictive features. They explain that some system-level features drift over time or change with the cloud environment, so a model can look good in a random split and fall apart on later data.

A method warning. With a random split into training and testing sets, the same approach on Dataset 1 reached a catch rate of 91.64% at a 0.1% false-positive rate, against 36.50% when the model was trained on the past and tested on the following month. The authors say this gap is larger when features depend on time or environment, for example disks in the same rack sharing voltage changes, and conclude that cross-validation is not suitable for judging an online predictor. Lu et al. (FAST 2020), working on another operator's data, also warn that long test periods are needed before concluding anything about model quality and that porting a model between sites can lose a lot of accuracy.

What other fleets say. Backblaze's 2016 post found that 76.7% of failed drives, against 4.2% of operational drives, had one or more of five SMART counts above zero, so 23.3% of failed drives showed no warning in those counts. Lu et al. found models using SMART alone had a very high false-negative rate and that adding performance and location data cut it. Both are consistent with Xu et al.'s starting point that SMART data are an incomplete view of a disk that is starting to misbehave. In production the authors use CDEF to prefer healthier disks when allocating VMs and to live-migrate VMs off predicted disks, and they report an average of about 63,000 minutes of VM downtime saved per month.

Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.

Sources

  1. 01Improving Service Availability of Cloud Systems by Predicting Disk Error, Xu et al., USENIX ATC 2018 (USENIX PDF, pages 481 to 494) · accessed 2026-10-10
  2. 02What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 · accessed 2026-10-10
  3. 03Making Disk Failure Predictions SMARTer!, Lu, Luo, Patel, Yao, Tiwari, and Shi, FAST 2020 (USENIX PDF, pages 151 to 167) · accessed 2026-10-10