Skip to content

Paper note

38% of failed datacenter SSDs showed none of the four SMART symptoms; the ones that did failed 3 to 20 times as often

Published 2026-10-10

In Narayanan et al. (SYSTOR 2016), a study of over half a million SSDs in Microsoft datacenters over nearly 3 years, 62% of devices that fail-stopped had shown at least one of four SMART symptoms (uncorrectable or CRC data errors, sector reallocations, program or erase failures, and SATA downshifts) and 38% had shown none. Devices with a symptom failed 3 to 20 times as often as devices without one, with data errors the strongest at up to 20 times. Of the models whose failure rate could be compared with vendor figures, two consumer models exceeded the published 0.61% to 0.73% annual failure rate. A multi-factor classifier reached 87% precision and 71% recall, but it was scored with 5-fold cross-validation on one operator's data. Maneas et al. (FAST 2020) and Backblaze (2016) show the same pattern in other fleets: symptoms raise risk and a large share of failures give no warning.

Population
Over half a million SSDs, five model groups, five large datacenters and several edge sites of one operator
Window
Nearly 3 years; 34-month window with 30 months analysed
Failure definition
Fail-stop: an SSD problem shuts a server down for repair or replacement; about 80% were replaced
Symptoms
Data errors (uncorrectable and CRC), reallocated sectors, program and erase failures, SATA downshifts
Cross-checks
NetApp SSDs, FAST 2020 (1.4 million drives); Backblaze hard disks, October 2016 (67,814 drives)

Numbers

SSD failure statistics as printed in Narayanan et al., SYSTOR 2016 (over half a million SSDs, nearly 3 years, one operator), rows 1 to 12, and in two cross-check sources, rows 13 to 15. AFR is the percentage of devices with failures divided by device years. Qualifiers such as 'as much as' and 'around' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
SSDs studiedover 500,000devices5 model groups; consumer MLC from vendor 1 and enterprise model 2-ASYSTOR 2016
Vendor-quoted annual failure rate for the consumer models0.61 to 0.73percent per yearSpecification range shown in the paper's AFR chartSYSTOR 2016
Highest observed AFR above specificationas much as 70percent above the quoted rateConsumer models 1-B and 1-C exceeded the rangeSYSTOR 2016
Uncorrectable bit error rate observed10^-11 to 10^-14errors per bit readModels that report error counts and data volume; target is 10^-15SYSTOR 2016
Replacement after a failure ticket, SSD versus hard disk79 versus 11percent of ticketsDatacenter tickets in the studySYSTOR 2016
AFR increase when a symptom is present3 to 20timesFour SMART symptom categories; data errors up to 20 timesSYSTOR 2016
AFR increase with reallocations or SATA downshift4timesCompared with devices without the symptomSYSTOR 2016
Failed devices with at least one of the four symptomsaround 62percent of failed devices38% showed noneSYSTOR 2016
Healthy devices with data errors1percent of healthy devicesAbout 12% of failed devices had themSYSTOR 2016
AFR increase with average writes per day2 to 4timesOlder models 1-A, 1-B, and 1-C; not seen for 1-DSYSTOR 2016
Classifier precision87percentFailed versus healthy devices; 5-fold cross-validationSYSTOR 2016
Classifier recall71percentSame classifierSYSTOR 2016
Average annual replacement rate of enterprise SSDs0.22percent per yearNetApp, about 1.4 million SSDs; models from 0.07 to nearly 1.2FAST 2020
Failed drives with one or more of five SMART counts above zero76.7percent of failed drivesBackblaze hard disks, 67,814 drives in 2016; 23.3% showed noneBackblaze 2016
Operational drives with one or more of five SMART counts above zero4.2percent of operational drivesSame fleetBackblaze 2016

This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 12 are from Narayanan et al. (SYSTOR 2016). Rows 13 to 15 are cross-checks from other fleets and are different measures.

Method

The figures are copied from the PDF of Narayanan, Wang, Jeon, Sharma, Caulfield, Sivasubramaniam, Cutler, Liu, Khessib, and Vaid, SYSTOR 2016 (Article 7, pages 1 to 11), not refit. The data cover over half a million SSDs in five large datacenters and several edge sites of one operator, in five model groups from two vendors (older 160 GB and newer 480 GB MLC drives), over nearly 3 years; failures are fail-stop events, where an SSD problem takes the server down and the device is replaced or repaired. The authors use a 34-month window, analyse the first 30 months, and keep the last 4 months only to classify devices. The annual failure rate (AFR) is the percentage of devices with failures divided by total device years. They compare AFR with and without four SMART-based symptoms, then build a multi-factor classifier that separates failed from healthy devices and report precision and recall with 5-fold cross-validation and SMOTE over-sampling. Two other sources check the surrounding claims: Maneas et al., FAST 2020 (1.4 million NetApp SSDs) and a Backblaze post of 6 October 2016 (hard disks). Neither tests the 87% and 71% figures.

Limits

The numbers are one operator's own fleet and workloads, and the model names are anonymized. Failure here means a fail-stop that takes the server down, about 80% of which ended in replacement, so it is not the same as the replacement rates in other studies. Several of the paper's results are in charts (the AFR by model, the AFR with and without each symptom), and this note uses only the figures printed in the text: the spec range of 0.61% to 0.73%, the 3 to 20 times range, and the 'as much as 70%' wording. The classifier was scored with random 5-fold cross-validation on devices, not with a split by time; Xu et al. (ATC 2018, covered in a separate note) found that random splits gave 91.64% against 36.50% when tested forward in time, so 87% and 71% may be optimistic for a deployed predictor. The text I read does not say whether the precision is measured before or after over-sampling. The correlations with write rate and space use differ by model, and the paper's own causal analysis is a model, not an experiment. The cross-check sources are other fleets with different definitions and do not test this classifier.

What the paper found

The answer. Symptoms raise risk but do not pick out most failures. In Narayanan et al., 62% of failed devices had shown at least one of the four symptoms, while 38% had shown none, so the authors conclude a pure symptom-based diagnosis is not very accurate. Devices with a symptom failed 3 to 20 times as often as those without, with data errors (uncorrectable and CRC) the strongest at up to 20 times, reallocations and SATA downshifts about 4 times, and program or erase failures 2.75 times. Data errors were rare on healthy devices (about 1%) and present on about 12% of failed ones, while reallocations were common among healthy devices too, so they are not a sufficient indicator by themselves.

Field rates against the datasheet. Vendor specifications for the consumer models quote an annual failure rate of 0.61% to 0.73%. Model 1-A was within that range, models 1-B and 1-C exceeded it, and model 1-D was well below it; the contributions list says some models were as much as 70% above specification. The authors read this as showing that real workloads change failure behaviour compared with a vendor's test conditions. The measured uncorrectable bit error rate was 10^-11 to 10^-14, at least an order of magnitude above the 10^-15 target. An SSD-related failure ticket led to a replacement 79% of the time, against 11% for hard disk tickets.

Workload and the classifier. For the older models 1-A, 1-B, and 1-C, AFR rose by 2 to 4 times as the average writes per day increased, while the newer, larger 1-D showed no clear relation. Combining tens of factors, including symptoms and workload measures such as total NAND writes, a classifier identified failed devices with 87% precision and 71% recall. Devices that matched a failure signature tended to fail within a month if the signature involved symptoms, and survived longer if it was based on workload factors alone.

What other fleets say. Maneas et al. found an average annual replacement rate of 0.22% across 1.4 million NetApp SSDs, ranging from 0.07% to nearly 1.2% by model, and that drives with a non-empty defect list were replaced significantly more often, the same direction as the symptoms here. Backblaze's 2016 post on hard disks found 76.7% of failed drives, against 4.2% of operational ones, with one or more of five SMART counts above zero, so 23.3% showed no warning. Those fleets, drive types and failure definitions differ, so the shares are not comparable, but all three show a warning signal that is real and incomplete.

Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.

Sources

  1. 01SSD Failures in Datacenters: What? When? and Why?, Narayanan et al., SYSTOR 2016 (author PDF, ACM article 7, pages 1 to 11) · accessed 2026-10-10
  2. 02A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, pages 137 to 149) · accessed 2026-10-10
  3. 03What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 · accessed 2026-10-10