80% of 315 verified fail-slow drives at Alibaba were caused by software scheduling, and 216 of them were in two clusters
Published 2026-10-11
In Perseus (Lu et al., FAST 2023), a fail-slow detector run on 248K Alibaba drives for 10 months found 304 fail-slow drives, and of 315 verified fail-slow drives in the authors' benchmark, 252 (80.0%) were caused by ill-implemented software scheduling and 63 (20.0%) by hardware, with 216 of the 252 in just two clusters of open-channel SSDs that shared one flaw. Isolating the drives cut node-level 99.99th percentile write latency by 48%. Gunawi et al. (FAST 2018), a separate set of 101 fail-slow reports from 12 institutions, found 39% of root causes were external factors such as configuration, environment, power, and temperature, and that 17% of incidents took months to detect. In both, a drive that runs slowly is often slow for a reason outside the drive.
- Population
- 248K Alibaba drives under monitoring; benchmark of 315 verified fail-slow drives and about 41K peers in 25 clusters
- Window
- 10 months of monitoring; 15 consecutive days of traces in the benchmark
- Fail-slow definition
- A component that keeps working but delivers lower-than-expected performance
- Root cause groups
- Software scheduling, hardware defects, environmental factors; verified by on-site engineers or manufacturers
- Cross-check
- 101 fail-slow hardware reports from 12 institutions, FAST 2018
Numbers
This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 12 are from Lu et al. (FAST 2023). Rows 13 to 15 are from Gunawi et al. (FAST 2018); they are cross-checks with different measures.
Method
The figures are copied from the USENIX PDF of Lu, Xu, Zhang, Zhu, Zhu, Wang, Zhu, Xue, Shu, Li, and Wu, FAST 2023, not refit. Perseus is a regression-based detector that compares each drive's latency with its peers in the same cluster and flags sustained slowdowns. The authors ran it for 10 months on 248K drives, then built a benchmark from 15 consecutive days of traces for 315 verified fail-slow drives (237 SSDs and 78 HDDs, each verified by on-site engineers or manufacturers) and about 41K normal peer drives in 25 clusters, and sorted the verified drives by root cause. Gunawi et al., FAST 2018, is a separate source: 101 reports of fail-slow hardware from 12 institutions, collected from operators and not from logs. Neither source tests the other's counts.
Limits
The numbers are one operator's cloud storage, drives, and scheduling software. The 80% software share is dominated by one cause: 216 of the 252 software-caused drives were SSDs in two open-channel SSD clusters, where the same logical drives on every node were slow, so the share describes a few clusters and not 252 independent drive faults. Only 15 of the 63 hardware-caused drives (9 HDDs and 6 SSDs) have a root cause returned by the vendor. The benchmark contains drives the detector flagged (304 of the 315) plus a few found otherwise, so the precision of 0.99 and recall of 1.00 are not tested on an unbiased sample. The latency cut is a node-level mean with a wide error range. Gunawi et al. is a convenience sample of reports that operators chose to write up, so its 39% is not a rate, and a report can hold several causes. The two sources count different things, drives and incident reports, so the numbers are not comparable.
What the paper found
The answer. In Lu et al., most verified fail-slow drives were not broken hardware. Of 315 verified fail-slow drives, 252 (80.0%) were slow because of ill-implemented software scheduling, 216 SSDs and 36 HDDs, and 63 (20.0%) were hardware cases, 42 HDDs and 21 SSDs. The caveat is how concentrated the 252 are: 216 were in two clusters of open-channel SSDs, where the same logical drives on every node were affected, so one scheduling flaw produced most of the drives. Excluding clusters with software-caused cases, the authors test the detector on a subset to show it is not specific to their stack.
Detection pays off. Run for 10 months on 248K drives, the detector found 304 fail-slow drives. Isolating them cut node-level write latency by 30.67% at the 95th percentile, 46.39% at the 99th, and 48.05% at the 99.99th, with wide error ranges (plus or minus 10.96 to 15.53 points). On the authors' benchmark it reached a precision of 0.99 and a recall of 1.00, well above threshold and peer-comparison baselines, but the benchmark is built mostly from drives the detector itself flagged.
Hardware cases are hard to explain. Only 15 of the 63 hardware cases (9 HDDs and 6 SSDs) came back from the vendor with a root cause. The authors list bad sectors that cause repeated remapping and a rotor with eccentricity as HDD causes, and say hardware defects are usually neither necessary nor sufficient for fail-slow, so a count of bad sectors cannot by itself be read as the cause.
Another fleet agrees on direction. In Gunawi et al., 101 reports from 12 institutions list causes inside the device, such as firmware bugs and device errors or wear-out, and outside it. 39% of root causes were external factors (configuration, environment, power, temperature), so the authors advise operators to troubleshoot online with full-stack monitoring. 17% of incidents took months to find, 13% days, 13% hours, 11% weeks, 1% minutes, and 45% had no recorded time. The sources count different things, but both show that a slow drive is often not a failing drive.
Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.
Sources
- 01Perseus: A Fail-Slow Detection Framework for Cloud Storage Systems, Lu, Xu, Zhang, Zhu, Zhu, Wang, Zhu, Xue, Shu, Li, and Wu, USENIX FAST 2023 (USENIX PDF, sections 5 and 6) · accessed 2026-10-11
- 02Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems, Gunawi et al., USENIX FAST 2018 (USENIX PDF, sections 2 to 3) · accessed 2026-10-11