Skip to content

Paper note

In Alibaba's SSD fleet, 12.9% of failures were in the same node and 18.3% in the same rack within 30 minutes of another failure

Published 2026-10-11

In Han et al. (FAST 2021), a study of nearly one million SSDs of 11 models in Alibaba data centers over 2018 and 2019, 12.9% of the roughly 19,000 failures were intra-node failures and 18.3% were intra-rack failures, meaning another SSD in the same node or rack failed within 30 minutes. Once two SSDs in a node had failed, the chance of a further failure in that group was 26.3% to 64.3%, against a fleet annual failure rate of 1.16%. The strongest SMART attribute had a rank correlation of only 0.23 with these failures. Maneas et al. (FAST 2020, NetApp) found a drive replacement was 180 times more likely in the week after another in the same RAID group, and Ford et al. (OSDI 2010, Google) found 37% of node failures in bursts, so failures that cluster in space and time appear in all three fleets.

Population
Nearly one million SSDs, 11 models, 3 vendors, 200,000 nodes in 30,000 racks, Alibaba data centers
Window
Two years, January 2018 to December 2019
Failure definition
Administrator-checked trouble tickets, whole-drive or partial-drive failures; about 19,000 failed SSDs
Correlated failure definition
Another failure in the same node or rack within a threshold, 30 minutes by default
Cross-checks
NetApp SSD RAID groups, FAST 2020; Google storage cells, OSDI 2010

Numbers

SSD correlated-failure figures as printed in Han et al., FAST 2021 (Alibaba, nearly one million SSDs, two years), rows 1 to 13, and in two cross-check sources, rows 14 to 16. Qualifiers such as 'nearly' and 'about' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
SSDs studiednearly 1,000,000devices11 drive models from 3 vendors, Alibaba data centers
Nodes holding the SSDs200,000nodesSame fleet
Racks holding the SSDs30,000racksSame fleet
Length of the record2yearsJanuary 2018 to December 2019
Failed SSDs (trouble tickets)about 19,000drivesSame fleet
Annual failure rate of all SSDs1.16percent per yearSame fleet
Failures that are intra-node failures12.9percent of failuresFailures in one node within 30 minutes of each other
Failures that are intra-rack failures18.3percent of failuresFailures in one rack within 30 minutes of each other
Chance of one more failure in an intra-node failure group26.3 to 64.3percentGroup size 2 to 11 failures
Intra-node failures with an interval of one minute or less10.0percent of intra-node failuresSame fleet
Intra-rack failures with an interval of one month or less63.0percent of intra-rack failuresSame fleet
Highest rank correlation of a SMART attribute with correlated failures0.23Spearman correlation coefficientSMART attribute S187 (reported uncorrectable errors)
Intra-node failure share by application2.1 to 33.6percent of an application's SSD failuresEight largest applications
Chance of a drive replacement in a random week0.0504percentNetApp RAID groups, about 1.4 million SSDs
Chance of a replacement within a week of a previous one in the same RAID group9.39percentSame fleet
Node failures that are part of a burst of at least 2 nodes37percent of failuresTens of Google storage cells, 120 second window

This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 13 are from Han et al. (FAST 2021). Rows 14 and 15 are from Maneas et al. (FAST 2020) and row 16 from Ford et al. (OSDI 2010); they are cross-checks with different measures.

Method

The figures are copied from the USENIX PDF of Han, Lee, Xu, Liu, He, and Liu, FAST 2021, not refit. The data are SMART logs, trouble tickets, physical locations, and application labels for nearly one million SSDs of 11 models from three vendors in Alibaba data centers, from January 2018 to December 2019. The tickets are the ground truth for failures and each was checked by an administrator. An intra-node (intra-rack) failure is a failure in a node (rack) that happens within a threshold, 30 minutes by default, of another failure in the same node (rack). The paper then simulates redundancy schemes on the failure traces. Two other sources check the surrounding claim that failures cluster: Maneas et al., FAST 2020 (about 1.4 million NetApp SSDs in RAID groups) and Ford et al., OSDI 2010 (Google storage cells, node failures). Neither tests the SMART or simulation results.

Limits

The numbers are one operator's own fleet, workloads, and ticketing rules. The 12.9% and 18.3% shares depend on the 30 minute threshold, and the paper itself shows the shares change when the threshold is one minute or one month. Failures are trouble tickets, including partial drive failures. The authors note that their SMART logs have missing days, that the data has no repair details or operating system logs, and that 88.6% of nodes with two or more SSDs hold a single drive model, so a share of the spatial clustering may reflect a bad drive model or batch rather than a node or rack cause. The comparison with the annual failure rate is the paper's approximation for independent failures. The redundancy results come from a trace-driven simulation with the authors' repair assumptions, not from a live system. Maneas et al. count drive replacements in RAID groups and Ford et al. count node failures in a distributed file system, so those numbers use other definitions and are not comparable with the Alibaba shares.

What the paper found

The answer. In Han et al., SSD failures cluster in the same node and the same rack. 12.9% of failures were intra-node failures and 18.3% intra-rack failures, defined as another failure in the same scope within 30 minutes. Failure groups were sometimes large, and a group size can exceed what typical redundancy schemes tolerate, for example four failures. The conditional probability of one more failure in an intra-node failure group was 26.3% to 64.3% as the group grew from 2 to 11 failures, against an annual failure rate of 1.16% for the whole fleet, so a node that has already lost drives is much more likely to lose another.

Timing and drive factors. Failures in the same scope are often close in time: 10.0% of intra-node failures had an interval of one minute or less and 29.2% one month or less, while 14.4% of intra-rack failures were within one minute and 63.0% within one month. Intra-node and intra-rack failures were more likely in nodes and racks that held many SSDs of the same model, and the shares rose with drive age. Write-dominant applications had more SSD failures overall, and the intra-node share ranged from 2.1% to 33.6% across the eight largest applications.

SMART does not flag them. The highest rank correlation between any SMART attribute and intra-node or intra-rack failures, from S187 reported uncorrectable errors, was only 0.23, so the authors conclude SMART attributes are not good indicators for detecting that such correlated failures exist. In a simulation, schemes that tolerate independent failures, such as 3-way replication and several erasure codes, lost no data under independently generated failures at the fleet rate, but their reliability degraded under the real failure patterns because follow-on failures compete for repair bandwidth. Erasure coding showed higher reliability than replication on those patterns.

Other fleets show the same clustering. Maneas et al. found a 0.0504% chance of a drive replacement in a random week but 9.39% in the week after a previous replacement in the same RAID group, more than 180 times higher, and 52% of consecutive replacements within a week. Ford et al. found 37% of node failures in Google cells were part of a burst of at least 2 nodes. The three fleets, devices, and measures differ, so the sizes are not comparable, but all three reject the assumption of independent failures.

Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.

Sources

  1. 01An In-Depth Study of Correlated Failures in Production SSD-Based Data Centers, Han, Lee, Xu, Liu, He, and Liu, USENIX FAST 2021, pages 417 to 429 (USENIX PDF) · accessed 2026-10-11
  2. 02A Study of SSD Reliability in Large Scale Enterprise Storage Deployments, Maneas, Mahdaviani, Emami, and Schroeder, FAST 2020 (USENIX PDF, section 6 on correlated replacements) · accessed 2026-10-11
  3. 03Availability in Globally Distributed Storage Systems, Ford et al., USENIX OSDI 2010 (USENIX PDF, section 4 on correlated failures) · accessed 2026-10-11