Skip to content

Paper note

Enterprise SSD replacements ran at 0.22% a year across 1.4 million NetApp drives, with firmware the biggest swing

Published 2026-10-10

In Maneas et al. (FAST 2020), about 1.4 million SSDs in NetApp enterprise storage systems had an average annual replacement rate (ARR) of 0.22%, with models ranging from 0.07% to nearly 1.2%. That is well below the 4% to 10% of Google's flash drives removed within four years and the 2% to 9% a year the same paper quotes for hard disks, though the definitions of replacement differ. Within the NetApp data, the earliest firmware versions of some models were replaced up to 8 to more than 10 times as often as later ones, and the chance of a second replacement in the same RAID group within a week was 9.39% against 0.0504% in a random week. One third of replacements were predictive and another third were severe SCSI errors. Method limits are on the note.

Population
About 1.4 million SSDs in NetApp filers; 3 manufacturers, 18 models, 12 capacities
Window
10 data snapshots from January 2017 to May 2019; 30 months of telemetry
Failure definition
Replacement: a drive marked as failed and swapped out, with a reason type recorded for 60% of events
Metric
Annual replacement rate (ARR): replacements divided by device years
Cross-checks
Google flash drives, FAST 2016 (millions of drive days, 10 models); hard disks in backup systems, FAST 2015

Numbers

Enterprise SSD replacement statistics as printed in Maneas et al., FAST 2020 (NetApp, about 1.4 million SSDs, 18 models, 30 months, January 2017 to May 2019), rows 1 to 13, and two cross-check sources, rows 14 to 16. ARR is replacements divided by device years. A replacement is a drive marked as failed and swapped out. Qualifiers such as 'about' and 'more than' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
Average annual replacement rate0.22percent per yearAll studied NetApp SSDsFAST 2020
Annual replacement rate range across models0.07 to nearly 1.2percent per year18 modelsFAST 2020
Annual replacement rate, two similar 15 TB-class models0.53 versus 1.13percent per yearII-G 15TB versus II-C 15.3TBFAST 2020
Replacements caused by SCSI errors32.78percent of replacementsReason type normalized for the 40% of events with no reasonFAST 2020
Replacements where the drive became unresponsive0.60percent of replacementsSame normalizationFAST 2020
Predictive replacements (category D)about one thirdshare of replacementsDrive still operating when replacedFAST 2020
Length of the infant-mortality period20 to 40percent of a 5-year lifeeMLC and 3D-TLC drivesFAST 2020
Higher replacement rate with less than 1% rated life used1.25timeseMLC drivesFAST 2020
Drop in ARR from the earliest firmware to a later one8, more than 10, more than 2timesFamilies II-A (FV2 to FV3), II-F (FV2 to FV3), and I-B (FV1 to FV2)FAST 2020
Drives with an empty defect list99.04percent of drivesDefect list (g-list) of unrecoverable-error blocksFAST 2020
Chance of a RAID group replacement in a random week0.0504percent per weekAll RAID groupsFAST 2020
Chance of another replacement within a week of one9.39percent per weekSame RAID group; more than 180 times the random-week figureFAST 2020
Consecutive replacements within one week of each other52percent of consecutive replacementsRAID groups with more than one replacement; 46% within one dayFAST 2020
Google flash drives permanently removed within 4 yearsabout 5, about 10 for the worst modelspercent of drivesGoogle data centers, 10 models, over 6 years of production useGoogle, FAST 2016
Google: fraction of flash drives with uncorrectable errors in four yearsmore than 20percent of drivesSame populationGoogle, FAST 2016
Hard disks: failed disks found in their fourth year, three models63, 66, and 64percent of failed disksAbout 1 million SATA disks in one vendor's backup systemsFAST 2015

This table is a short extract of printed figures, not a copy of the papers and not the telemetry. Rows 1 to 13 are from Maneas et al. (FAST 2020). Rows 14 to 16 are cross-checks from other fleets and are different measures.

Method

The figures are copied from the USENIX PDF of Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (pages 137 to 149), not refit. The authors mined NetApp Active IQ support messages in 10 snapshots from January 2017 to May 2019, covering about 1.4 million SSDs from 3 manufacturers, 18 models, and 12 capacities, over a period of 30 months. A replacement is a drive marked as failed and swapped out, usually for a hot spare, and the paper uses replacement and failure interchangeably. The annual replacement rate is device failures divided by device years. Percentages by reason are normalized because the reason is missing for 40% of replacement events. Confidence intervals are 95%, and the authors run two-sample z-tests. Two sources are used as cross-checks of the order of magnitude and direction: Schroeder, Lagisetty, and Merchant, FAST 2016 (Google flash drives, a different fleet) and Ma et al., FAST 2015 (hard disks in another vendor's backup systems). Neither reproduces the NetApp numbers.

Limits

The data are NetApp's own telemetry for a sample of its SSD population, and drives come from three manufacturers whose names are anonymized. Replacement is an operator action, and the paper notes that about a third of replacements are predictive and that the drive was still working when replaced, so ARR is not the same as a data-loss rate. The reason type is missing for 40% of replacement events because of a data-collection issue, and the authors assume the rest are spread in proportion. The firmware effect holds within a drive family and for drives with less than 1% of rated life used, but the paper says the likely cause (bug fixes in later versions) is an explanation and not a measurement. The rated-life field is a truncated integer, so many drives show 0. Some results, such as 3D-TLC beyond 1% of rated life, rest on limited data and have wide confidence intervals, and two outlier models were excluded from some figures. The RAID follow-up figures count replacements within a group, not data loss. The Google flash figure (4% to 10% in four years) is a share of drives permanently removed, not an annual rate, so it is only roughly comparable to 0.22% a year. The hard-disk range of 2% to 9% a year is a figure the Maneas paper quotes from earlier work; this note did not re-open those two papers.

What the paper found

The answer. Across the whole NetApp population the annual replacement rate was 0.22%, and it varied by model from 0.07% to nearly 1.2%. The authors say these numbers are significantly lower than those reported for data center SSDs and common numbers for hard disks. Even for models with similar specifications the rate can differ a lot: 0.53% for II-G 15TB drives against 1.13% for II-C 15.3TB drives. Google's own flash study found that most models had about 5% of drives permanently removed within 4 years and the worst about 10%, which is a different measure with a similar direction: flash replacements are rare compared with hard disks, but not zero.

Why drives were replaced. The most common single reason was a SCSI-layer error, about 33% (32.78%) of replacements and one of the severe types. A drive becoming completely unresponsive accounted for only 0.60%. About one third of replacements were predictive (category D) and so the drive was still operating when it was swapped out. Lost writes and aborted or timed-out commands were about equally common. The wear limit was not the problem: the typical drive used less than 2% of its rated program-erase cycles, and the 99th and 99.9th percentile drives had used 15% and 33%.

Early life and firmware. Instead of the textbook bathtub curve, failure rates rose for 12 to 15 months, then fell slowly for another 6 to 12 months before levelling out, so drives spend 20% to 40% of a 5-year life in infant mortality, with rates 2 to 3 times larger than later in life. For eMLC drives, those with less than 1% of rated life used were 1.25 times more likely to be replaced than drives with more. Firmware mattered a lot: 70% of drives stayed on one firmware version across snapshots, and the earliest versions could have an order of magnitude higher ARR than later versions, such as a factor of 8 for family II-A (FV2 to FV3) and more than 10 for II-F, and more than 2 for I-B. The effect persisted when only drives that never changed firmware were compared. Families such as II-J also showed a later version with a higher rate, so a newer version is not always better.

Warning signs and RAID groups. The defect list (g-list) was empty for 99.04% of drives, and drives with at least one entry were replaced significantly more often, including for reasons other than prediction. This is the same direction as the reallocated-sector result for hard disks in Ma et al., and Google likewise reports that previous errors predict later uncorrectable errors. Inside a RAID group, the chance of another replacement within a week was 9.39% after a replacement, against 0.0504% in a random week, a factor of more than 180. Of the consecutive replacements in one group, 46% were at most one day apart and 52% within a week. Larger groups had more replacements, but no clear rise in the share of groups with a follow-up failure, so the authors find no evidence that group size drives multiple failures.

Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.

Sources

  1. 01A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, pages 137 to 149) · accessed 2026-10-10
  2. 02A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, USENIX FAST 2020 presentation page (abstract) · accessed 2026-10-10
  3. 03Flash Reliability in Production: The Expected and the Unexpected, Schroeder, Lagisetty, and Merchant, FAST 2016 (USENIX PDF) · accessed 2026-10-10
  4. 04RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk Failures, Ma et al., FAST 2015 (USENIX PDF) · accessed 2026-10-10