38% of failed datacenter SSDs showed none of the four SMART symptoms; the ones that did failed 3 to 20 times as often
Published 2026-10-10
In Narayanan et al. (SYSTOR 2016), a study of over half a million SSDs in Microsoft datacenters over nearly 3 years, 62% of devices that fail-stopped had shown at least one of four SMART symptoms (uncorrectable or CRC data errors, sector reallocations, program or erase failures, and SATA downshifts) and 38% had shown none. Devices with a symptom failed 3 to 20 times as often as devices without one, with data errors the strongest at up to 20 times. Of the models whose failure rate could be compared with vendor figures, two consumer models exceeded the published 0.61% to 0.73% annual failure rate. A multi-factor classifier reached 87% precision and 71% recall, but it was scored with 5-fold cross-validation on one operator's data. Maneas et al. (FAST 2020) and Backblaze (2016) show the same pattern in other fleets: symptoms raise risk and a large share of failures give no warning.
- Population
- Over half a million SSDs, five model groups, five large datacenters and several edge sites of one operator
- Window
- Nearly 3 years; 34-month window with 30 months analysed
- Failure definition
- Fail-stop: an SSD problem shuts a server down for repair or replacement; about 80% were replaced
- Symptoms
- Data errors (uncorrectable and CRC), reallocated sectors, program and erase failures, SATA downshifts
- Cross-checks
- NetApp SSDs, FAST 2020 (1.4 million drives); Backblaze hard disks, October 2016 (67,814 drives)
Numbers
| Measure | Value | Unit | Population or scope | Source |
|---|---|---|---|---|
| SSDs studied | over 500,000 | devices | 5 model groups; consumer MLC from vendor 1 and enterprise model 2-A | SYSTOR 2016 |
| Vendor-quoted annual failure rate for the consumer models | 0.61 to 0.73 | percent per year | Specification range shown in the paper's AFR chart | SYSTOR 2016 |
| Highest observed AFR above specification | as much as 70 | percent above the quoted rate | Consumer models 1-B and 1-C exceeded the range | SYSTOR 2016 |
| Uncorrectable bit error rate observed | 10^-11 to 10^-14 | errors per bit read | Models that report error counts and data volume; target is 10^-15 | SYSTOR 2016 |
| Replacement after a failure ticket, SSD versus hard disk | 79 versus 11 | percent of tickets | Datacenter tickets in the study | SYSTOR 2016 |
| AFR increase when a symptom is present | 3 to 20 | times | Four SMART symptom categories; data errors up to 20 times | SYSTOR 2016 |
| AFR increase with reallocations or SATA downshift | 4 | times | Compared with devices without the symptom | SYSTOR 2016 |
| Failed devices with at least one of the four symptoms | around 62 | percent of failed devices | 38% showed none | SYSTOR 2016 |
| Healthy devices with data errors | 1 | percent of healthy devices | About 12% of failed devices had them | SYSTOR 2016 |
| AFR increase with average writes per day | 2 to 4 | times | Older models 1-A, 1-B, and 1-C; not seen for 1-D | SYSTOR 2016 |
| Classifier precision | 87 | percent | Failed versus healthy devices; 5-fold cross-validation | SYSTOR 2016 |
| Classifier recall | 71 | percent | Same classifier | SYSTOR 2016 |
| Average annual replacement rate of enterprise SSDs | 0.22 | percent per year | NetApp, about 1.4 million SSDs; models from 0.07 to nearly 1.2 | FAST 2020 |
| Failed drives with one or more of five SMART counts above zero | 76.7 | percent of failed drives | Backblaze hard disks, 67,814 drives in 2016; 23.3% showed none | Backblaze 2016 |
| Operational drives with one or more of five SMART counts above zero | 4.2 | percent of operational drives | Same fleet | Backblaze 2016 |
This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 12 are from Narayanan et al. (SYSTOR 2016). Rows 13 to 15 are cross-checks from other fleets and are different measures.
Method
The figures are copied from the PDF of Narayanan, Wang, Jeon, Sharma, Caulfield, Sivasubramaniam, Cutler, Liu, Khessib, and Vaid, SYSTOR 2016 (Article 7, pages 1 to 11), not refit. The data cover over half a million SSDs in five large datacenters and several edge sites of one operator, in five model groups from two vendors (older 160 GB and newer 480 GB MLC drives), over nearly 3 years; failures are fail-stop events, where an SSD problem takes the server down and the device is replaced or repaired. The authors use a 34-month window, analyse the first 30 months, and keep the last 4 months only to classify devices. The annual failure rate (AFR) is the percentage of devices with failures divided by total device years. They compare AFR with and without four SMART-based symptoms, then build a multi-factor classifier that separates failed from healthy devices and report precision and recall with 5-fold cross-validation and SMOTE over-sampling. Two other sources check the surrounding claims: Maneas et al., FAST 2020 (1.4 million NetApp SSDs) and a Backblaze post of 6 October 2016 (hard disks). Neither tests the 87% and 71% figures.
Limits
The numbers are one operator's own fleet and workloads, and the model names are anonymized. Failure here means a fail-stop that takes the server down, about 80% of which ended in replacement, so it is not the same as the replacement rates in other studies. Several of the paper's results are in charts (the AFR by model, the AFR with and without each symptom), and this note uses only the figures printed in the text: the spec range of 0.61% to 0.73%, the 3 to 20 times range, and the 'as much as 70%' wording. The classifier was scored with random 5-fold cross-validation on devices, not with a split by time; Xu et al. (ATC 2018, covered in a separate note) found that random splits gave 91.64% against 36.50% when tested forward in time, so 87% and 71% may be optimistic for a deployed predictor. The text I read does not say whether the precision is measured before or after over-sampling. The correlations with write rate and space use differ by model, and the paper's own causal analysis is a model, not an experiment. The cross-check sources are other fleets with different definitions and do not test this classifier.
What the paper found
The answer. Symptoms raise risk but do not pick out most failures. In Narayanan et al., 62% of failed devices had shown at least one of the four symptoms, while 38% had shown none, so the authors conclude a pure symptom-based diagnosis is not very accurate. Devices with a symptom failed 3 to 20 times as often as those without, with data errors (uncorrectable and CRC) the strongest at up to 20 times, reallocations and SATA downshifts about 4 times, and program or erase failures 2.75 times. Data errors were rare on healthy devices (about 1%) and present on about 12% of failed ones, while reallocations were common among healthy devices too, so they are not a sufficient indicator by themselves.
Field rates against the datasheet. Vendor specifications for the consumer models quote an annual failure rate of 0.61% to 0.73%. Model 1-A was within that range, models 1-B and 1-C exceeded it, and model 1-D was well below it; the contributions list says some models were as much as 70% above specification. The authors read this as showing that real workloads change failure behaviour compared with a vendor's test conditions. The measured uncorrectable bit error rate was 10^-11 to 10^-14, at least an order of magnitude above the 10^-15 target. An SSD-related failure ticket led to a replacement 79% of the time, against 11% for hard disk tickets.
Workload and the classifier. For the older models 1-A, 1-B, and 1-C, AFR rose by 2 to 4 times as the average writes per day increased, while the newer, larger 1-D showed no clear relation. Combining tens of factors, including symptoms and workload measures such as total NAND writes, a classifier identified failed devices with 87% precision and 71% recall. Devices that matched a failure signature tended to fail within a month if the signature involved symptoms, and survived longer if it was based on workload factors alone.
What other fleets say. Maneas et al. found an average annual replacement rate of 0.22% across 1.4 million NetApp SSDs, ranging from 0.07% to nearly 1.2% by model, and that drives with a non-empty defect list were replaced significantly more often, the same direction as the symptoms here. Backblaze's 2016 post on hard disks found 76.7% of failed drives, against 4.2% of operational ones, with one or more of five SMART counts above zero, so 23.3% showed no warning. Those fleets, drive types and failure definitions differ, so the shares are not comparable, but all three show a warning signal that is real and incomplete.
Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.
Sources
- 01SSD Failures in Datacenters: What? When? and Why?, Narayanan et al., SYSTOR 2016 (author PDF, ACM article 7, pages 1 to 11) · accessed 2026-10-10
- 02A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, pages 137 to 149) · accessed 2026-10-10
- 03What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 · accessed 2026-10-10