Azure disk-error prediction caught 36.5% of faulty disks at a 0.1% false-positive rate, and random cross-validation said 91.6%
Published 2026-10-10
In Xu et al. (USENIX ATC 2018), Microsoft's disk-error predictor, CDEF, found 36.50% of faulty disks on one test set when it was allowed to wrongly flag only 0.1% of healthy disks (29.67% to 41.09% across three test sets), with about 10,000 healthy disks per 3 faulty ones. Scored with a random split, the same task gave 91.64%, which is why the authors insist on testing forward in time. Adding system-level signals such as OS events to SMART raised the average catch rate at that false-positive rate from 27.6% to 30.3%, and to 35.8% with feature selection. In production at Azure the authors report about 63,000 fewer minutes of VM downtime a month. The data are one cloud system in one month, and labels come from engineers' root-cause analysis of service problems.
- System
- One Microsoft cloud system; disks in clusters that host virtual machines, not RAID storage clusters
- Window
- Training on October 2017; testing on November 2017 in three test sets
- Target
- Disk errors found through root-cause analysis of service issues, not only whole-disk failure
- Class balance
- About 10,000 healthy disks to 3 faulty disks per test set; about 300 of 1,000,000 disks become faulty per day
- Cross-checks
- Lu et al., FAST 2020 (380,000 disks, 64 sites); Backblaze, October 2016 (67,814 drives)
Numbers
| Measure | Value | Unit | Population or scope | Source |
|---|---|---|---|---|
| Healthy to faulty disks in each test set | about 10,000 to 3 | ratio | Three test sets from November 2017 | ATC 2018 |
| Disks that become faulty per day | about 300 of 1,000,000 | disks | The studied cloud system | ATC 2018 |
| CDEF true positive rate at 0.1% false positives, Dataset 1 | 36.50 | percent of faulty disks | Cost-sensitive ranking model | ATC 2018 |
| CDEF true positive rate at 0.1% false positives, Datasets 2 and 3 | 41.09 and 29.67 | percent of faulty disks | Same model | ATC 2018 |
| Random forest true positive rate at 0.1% false positives, Datasets 1 to 3 | 30.51, 34.11, and 18.81 | percent of faulty disks | SMART features only | ATC 2018 |
| SVM true positive rate at 0.1% false positives, Datasets 1 to 3 | 15.51, 21.71, and 7.20 | percent of faulty disks | SMART features only | ATC 2018 |
| Area under the ROC curve on Dataset 1, CDEF, random forest, SVM | 0.93, 0.85, and 0.53 | AUC | A random guess is 0.5 | ATC 2018 |
| Average true positive rate with SMART features alone | 27.6 | percent of faulty disks | At 0.1% false positives, three datasets | ATC 2018 |
| Average true positive rate with SMART plus system-level signals | 30.3 | percent of faulty disks | Same setting | ATC 2018 |
| Average true positive rate with feature selection added | 35.8 | percent of faulty disks | Same setting; this is CDEF | ATC 2018 |
| True positive rate on Dataset 1 with cross-validation instead of a time split | 91.64 | percent of faulty disks | At 0.1% false positives; the time split gave 36.50 | ATC 2018 |
| Average VM downtime saved after deployment | around 63,000 | minutes per month | Azure; 99.999% availability allows 26 seconds per month | ATC 2018 |
| Failed drives with one or more of five SMART counts above zero | 76.7 | percent of failed drives | Backblaze, 67,814 drives in 2016; 23.3% showed none | Backblaze 2016 |
| Operational drives with one or more of five SMART counts above zero | 4.2 | percent of operational drives | Same fleet | Backblaze 2016 |
| Best model on an unseen site, trained on 62 sites | above 0.90 | MCC score | Lu et al., 380,000 disks; CNN-LSTM, 10-day horizon, with SMART, performance, and location data | FAST 2020 |
This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 12 are from Xu et al. (ATC 2018). Rows 13 to 15 are cross-checks from other fleets and are different measures.
Method
The figures are copied from the USENIX PDF of Xu, Sui, Yao, Zhang, Lin, Dang, Li, Jiang, Zhang, Lou, Chintalapati, and Zhang, USENIX ATC 2018 (pages 481 to 494), not refit. The authors predict disk errors (sector errors, latency errors, timeouts, and partition errors that show up as VM downtime), not only complete disk failure. Faulty disks are those found through root-cause analysis of service issues by field engineers. They train on October 2017 data from a Microsoft cloud system and test on November 2017, split into three test sets, and they compare their cost-sensitive ranking model (CDEF) with random forest and SVM baselines on SMART data. They report the true positive rate (TPR) at a false positive rate (FPR) of 0.1% and the area under the ROC curve (AUC). Two other sources check the surrounding claims: Lu et al., FAST 2020 (another operator) on SMART-only prediction, and a Backblaze post of 6 October 2016 on how often failed drives show SMART warnings. Neither reproduces the Azure numbers.
Limits
The authors write that the data come from one cloud service system of one company, so the results may not generalize. Labels come from engineers' root-cause analyses, and drives sent for repair may return, so some labels are noisy, which the authors acknowledge. They report only FPR, TPR, and AUC, and say precision and recall are future work, so the number of false alarms in absolute terms is not stated. By this note's arithmetic, with about 10,000 healthy disks per 3 faulty ones, a 0.1% false-positive rate is about 10 healthy disks wrongly flagged while a 36.5% catch rate finds about 1 of the 3 faulty ones, so most flagged disks would be healthy; the paper does not report precision. The 91.64% cross-validation figure is a warning about method, not a claim about performance. The 63,000 minutes of VM downtime saved per month is the authors' own average after deployment, and the text I read does not describe a control group. The model flags a disk by ranking and a cutoff, and predicted disks are live-migrated and stress-tested before replacement, so the actions are cheap compared with replacing a disk. The cross-check sources are from other fleets with other failure definitions and do not test this model.
What the paper found
The answer. At a false-positive rate of 0.1%, the Azure predictor CDEF caught 36.50%, 41.09%, and 29.67% of faulty disks on the three test sets, against 30.51%, 34.11%, and 18.81% for a random forest and 15.51%, 21.71%, and 7.20% for an SVM, both trained on SMART data. On Dataset 1 the area under the ROC curve was 0.93 for CDEF, 0.85 for the random forest, and 0.53 for the SVM. The authors note that the SVM has low cost because it rarely raises false alarms, and that it does poorly at finding faulty disks.
What helped. The authors' main addition was system-level signals, such as events from the operating system, on top of SMART. On average the catch rate at a 0.1% false-positive rate was 27.6% with SMART features alone, 30.3% with SMART plus system-level signals, and 35.8% when a feature selection step kept only stable, predictive features. They explain that some system-level features drift over time or change with the cloud environment, so a model can look good in a random split and fall apart on later data.
A method warning. With a random split into training and testing sets, the same approach on Dataset 1 reached a catch rate of 91.64% at a 0.1% false-positive rate, against 36.50% when the model was trained on the past and tested on the following month. The authors say this gap is larger when features depend on time or environment, for example disks in the same rack sharing voltage changes, and conclude that cross-validation is not suitable for judging an online predictor. Lu et al. (FAST 2020), working on another operator's data, also warn that long test periods are needed before concluding anything about model quality and that porting a model between sites can lose a lot of accuracy.
What other fleets say. Backblaze's 2016 post found that 76.7% of failed drives, against 4.2% of operational drives, had one or more of five SMART counts above zero, so 23.3% of failed drives showed no warning in those counts. Lu et al. found models using SMART alone had a very high false-negative rate and that adding performance and location data cut it. Both are consistent with Xu et al.'s starting point that SMART data are an incomplete view of a disk that is starting to misbehave. In production the authors use CDEF to prefer healthier disks when allocating VMs and to live-migrate VMs off predicted disks, and they report an average of about 63,000 minutes of VM downtime saved per month.
Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.
Sources
- 01Improving Service Availability of Cloud Systems by Predicting Disk Error, Xu et al., USENIX ATC 2018 (USENIX PDF, pages 481 to 494) · accessed 2026-10-10
- 02What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 · accessed 2026-10-10
- 03Making Disk Failure Predictions SMARTer!, Lu, Luo, Patel, Yao, Tiwari, and Shi, FAST 2020 (USENIX PDF, pages 151 to 167) · accessed 2026-10-10