Enterprise SSD replacements ran at 0.22% a year across 1.4 million NetApp drives, with firmware the biggest swing
Published 2026-10-10
In Maneas et al. (FAST 2020), about 1.4 million SSDs in NetApp enterprise storage systems had an average annual replacement rate (ARR) of 0.22%, with models ranging from 0.07% to nearly 1.2%. That is well below the 4% to 10% of Google's flash drives removed within four years and the 2% to 9% a year the same paper quotes for hard disks, though the definitions of replacement differ. Within the NetApp data, the earliest firmware versions of some models were replaced up to 8 to more than 10 times as often as later ones, and the chance of a second replacement in the same RAID group within a week was 9.39% against 0.0504% in a random week. One third of replacements were predictive and another third were severe SCSI errors. Method limits are on the note.
- Population
- About 1.4 million SSDs in NetApp filers; 3 manufacturers, 18 models, 12 capacities
- Window
- 10 data snapshots from January 2017 to May 2019; 30 months of telemetry
- Failure definition
- Replacement: a drive marked as failed and swapped out, with a reason type recorded for 60% of events
- Metric
- Annual replacement rate (ARR): replacements divided by device years
- Cross-checks
- Google flash drives, FAST 2016 (millions of drive days, 10 models); hard disks in backup systems, FAST 2015
Numbers
| Measure | Value | Unit | Population or scope | Source |
|---|---|---|---|---|
| Average annual replacement rate | 0.22 | percent per year | All studied NetApp SSDs | FAST 2020 |
| Annual replacement rate range across models | 0.07 to nearly 1.2 | percent per year | 18 models | FAST 2020 |
| Annual replacement rate, two similar 15 TB-class models | 0.53 versus 1.13 | percent per year | II-G 15TB versus II-C 15.3TB | FAST 2020 |
| Replacements caused by SCSI errors | 32.78 | percent of replacements | Reason type normalized for the 40% of events with no reason | FAST 2020 |
| Replacements where the drive became unresponsive | 0.60 | percent of replacements | Same normalization | FAST 2020 |
| Predictive replacements (category D) | about one third | share of replacements | Drive still operating when replaced | FAST 2020 |
| Length of the infant-mortality period | 20 to 40 | percent of a 5-year life | eMLC and 3D-TLC drives | FAST 2020 |
| Higher replacement rate with less than 1% rated life used | 1.25 | times | eMLC drives | FAST 2020 |
| Drop in ARR from the earliest firmware to a later one | 8, more than 10, more than 2 | times | Families II-A (FV2 to FV3), II-F (FV2 to FV3), and I-B (FV1 to FV2) | FAST 2020 |
| Drives with an empty defect list | 99.04 | percent of drives | Defect list (g-list) of unrecoverable-error blocks | FAST 2020 |
| Chance of a RAID group replacement in a random week | 0.0504 | percent per week | All RAID groups | FAST 2020 |
| Chance of another replacement within a week of one | 9.39 | percent per week | Same RAID group; more than 180 times the random-week figure | FAST 2020 |
| Consecutive replacements within one week of each other | 52 | percent of consecutive replacements | RAID groups with more than one replacement; 46% within one day | FAST 2020 |
| Google flash drives permanently removed within 4 years | about 5, about 10 for the worst models | percent of drives | Google data centers, 10 models, over 6 years of production use | Google, FAST 2016 |
| Google: fraction of flash drives with uncorrectable errors in four years | more than 20 | percent of drives | Same population | Google, FAST 2016 |
| Hard disks: failed disks found in their fourth year, three models | 63, 66, and 64 | percent of failed disks | About 1 million SATA disks in one vendor's backup systems | FAST 2015 |
This table is a short extract of printed figures, not a copy of the papers and not the telemetry. Rows 1 to 13 are from Maneas et al. (FAST 2020). Rows 14 to 16 are cross-checks from other fleets and are different measures.
Method
The figures are copied from the USENIX PDF of Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (pages 137 to 149), not refit. The authors mined NetApp Active IQ support messages in 10 snapshots from January 2017 to May 2019, covering about 1.4 million SSDs from 3 manufacturers, 18 models, and 12 capacities, over a period of 30 months. A replacement is a drive marked as failed and swapped out, usually for a hot spare, and the paper uses replacement and failure interchangeably. The annual replacement rate is device failures divided by device years. Percentages by reason are normalized because the reason is missing for 40% of replacement events. Confidence intervals are 95%, and the authors run two-sample z-tests. Two sources are used as cross-checks of the order of magnitude and direction: Schroeder, Lagisetty, and Merchant, FAST 2016 (Google flash drives, a different fleet) and Ma et al., FAST 2015 (hard disks in another vendor's backup systems). Neither reproduces the NetApp numbers.
Limits
The data are NetApp's own telemetry for a sample of its SSD population, and drives come from three manufacturers whose names are anonymized. Replacement is an operator action, and the paper notes that about a third of replacements are predictive and that the drive was still working when replaced, so ARR is not the same as a data-loss rate. The reason type is missing for 40% of replacement events because of a data-collection issue, and the authors assume the rest are spread in proportion. The firmware effect holds within a drive family and for drives with less than 1% of rated life used, but the paper says the likely cause (bug fixes in later versions) is an explanation and not a measurement. The rated-life field is a truncated integer, so many drives show 0. Some results, such as 3D-TLC beyond 1% of rated life, rest on limited data and have wide confidence intervals, and two outlier models were excluded from some figures. The RAID follow-up figures count replacements within a group, not data loss. The Google flash figure (4% to 10% in four years) is a share of drives permanently removed, not an annual rate, so it is only roughly comparable to 0.22% a year. The hard-disk range of 2% to 9% a year is a figure the Maneas paper quotes from earlier work; this note did not re-open those two papers.
What the paper found
The answer. Across the whole NetApp population the annual replacement rate was 0.22%, and it varied by model from 0.07% to nearly 1.2%. The authors say these numbers are significantly lower than those reported for data center SSDs and common numbers for hard disks. Even for models with similar specifications the rate can differ a lot: 0.53% for II-G 15TB drives against 1.13% for II-C 15.3TB drives. Google's own flash study found that most models had about 5% of drives permanently removed within 4 years and the worst about 10%, which is a different measure with a similar direction: flash replacements are rare compared with hard disks, but not zero.
Why drives were replaced. The most common single reason was a SCSI-layer error, about 33% (32.78%) of replacements and one of the severe types. A drive becoming completely unresponsive accounted for only 0.60%. About one third of replacements were predictive (category D) and so the drive was still operating when it was swapped out. Lost writes and aborted or timed-out commands were about equally common. The wear limit was not the problem: the typical drive used less than 2% of its rated program-erase cycles, and the 99th and 99.9th percentile drives had used 15% and 33%.
Early life and firmware. Instead of the textbook bathtub curve, failure rates rose for 12 to 15 months, then fell slowly for another 6 to 12 months before levelling out, so drives spend 20% to 40% of a 5-year life in infant mortality, with rates 2 to 3 times larger than later in life. For eMLC drives, those with less than 1% of rated life used were 1.25 times more likely to be replaced than drives with more. Firmware mattered a lot: 70% of drives stayed on one firmware version across snapshots, and the earliest versions could have an order of magnitude higher ARR than later versions, such as a factor of 8 for family II-A (FV2 to FV3) and more than 10 for II-F, and more than 2 for I-B. The effect persisted when only drives that never changed firmware were compared. Families such as II-J also showed a later version with a higher rate, so a newer version is not always better.
Warning signs and RAID groups. The defect list (g-list) was empty for 99.04% of drives, and drives with at least one entry were replaced significantly more often, including for reasons other than prediction. This is the same direction as the reallocated-sector result for hard disks in Ma et al., and Google likewise reports that previous errors predict later uncorrectable errors. Inside a RAID group, the chance of another replacement within a week was 9.39% after a replacement, against 0.0504% in a random week, a factor of more than 180. Of the consecutive replacements in one group, 46% were at most one day apart and 52% within a week. Larger groups had more replacements, but no clear rise in the share of groups with a follow-up failure, so the authors find no evidence that group size drives multiple failures.
Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.
Sources
- 01A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, pages 137 to 149) · accessed 2026-10-10
- 02A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, USENIX FAST 2020 presentation page (abstract) · accessed 2026-10-10
- 03Flash Reliability in Production: The Expected and the Unexpected, Schroeder, Lagisetty, and Merchant, FAST 2016 (USENIX PDF) · accessed 2026-10-10
- 04RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk Failures, Ma et al., FAST 2015 (USENIX PDF) · accessed 2026-10-10