Skip to content

Paper note

No clear SMART correlation with fail-slow NVMe SSDs: 4,584 of 779,978 drives, 10 later in the failure tickets

Published 2026-10-10

SMART attributes showed no clear correlation with fail-slow labels on the Alibaba NVMe SSDs in Lu et al. (USENIX ATC 2022). Under the 5-minute bar, Table 7 counts 4,584 slow drives out of 779,978 (0.59%). Table 5's 1.41% is the unweighted average of eight models, not that fleet share. The printed 6.05 times versus HDDs matches (1.41 - 0.20) / 0.20, not the quotient 1.41 / 0.20 (7.05). Only 10 of the 4,584 slow drives (around 0.22%) later appear in the failure tickets.

Fleet
Alibaba enterprise NVMe SSDs in all-flash nodes, 12 drives per node, not RAIDed, not hosting the OS. Fail-slow iostat covers more than half a million NVMe SSDs and more than 4 million HDDs.
Perf-log window
2020-11-16 to 2021-03-05, about 84 million entries. Daemon from 9 PM to 12 AM, 15-second averages, 720 records per drive per day.
5-minute census
Table 7: 779,978 NVMe SSDs scored, 4,584 slow, 4,439 in the failure tickets. Heavy-traffic drives are excluded from the Table 5 rates.
Whole study
Logs from over one million NVMe SSDs. SMART logs 2019-11-04 to 2020-11-14. Failure tickets 2019-11-04 to 2020-11-02. Those spans are not the perf-log window.
Fail-slow
A peer write-latency rule. Not a SMART attribute, and not a failure ticket.

Comparison

Fail-slow NVMe SSDs and peer HDDs in Lu et al., USENIX ATC 2022. Share columns are the unweighted model averages in Table 5 (eight enterprise NVMe models, three HDD models). Heavy-traffic drives are excluded from those rates. Perf logs run from 2020-11-16 to 2021-03-05, and the daemon records 15-second averages from 9 PM to 12 AM. Count columns are Table 7, the 5-minute contingency table only: slow drives against drives that later appear in the failure tickets. The 15-minute and 30-minute excess cells are not stated because the paper prints multiples only for the 5-minute bar (6.05, written as 1.41% to 0.20%) and the 60-minute bar (51, written as 0.52% to 0.01%). The H2 60-minute model cell is printed as <0.01%; the HDD average row used here is the printed 0.01. Units are in the column headers.
Minimum event length (minutes)NVMe model-average slow-drive share (percent of drives)HDD model-average slow-drive share (percent of drives)Printed excess versus HDD (times)Slow drives (count)Drives scored (count)Fleet slow-drive share (percent of drives)Later in the failure tickets (count)Drives in the failure tickets (count)Later ticketed (percent of slow drives)Later ticketed (percent of ticketed drives)Source
51.410.206.0545847799780.59104439around 0.22around 0.23ATC 2022
150.960.05not statednot statednot statednot statednot statednot statednot statednot statedATC 2022
300.750.03not statednot statednot statednot statednot statednot statednot statednot statedATC 2022
600.520.0151not statednot statednot statednot statednot statednot statednot statedATC 2022

This table copies printed average rows from Table 5 and the 5-minute census in Table 7. It is not a copy of the model rows and not the drive logs. The canonical file is the USENIX PDF linked in the source column. The checks below are arithmetic on those printed figures. Dong et al. is not a row: that paper cites the 1.41% sentence and does not remeasure it.

Checks on the printed figures

5-minute NVMe model shares (percent of drives)
I-A-2000 (15 nm) 4.44; II-A-1920 (20 nm) 1.25; II-C-1920 (32-layer) 0.52; II-C-4000 (32-layer) 0.06; II-D-1920 (64-layer) 0.17; II-D-3840 (64-layer) 0.48; III-B-1900 (48-layer) 3.04; III-B-3800 (48-layer) 1.31. Unweighted mean 1.40875, which Table 5 prints as 1.41. The paper does not print the drive count behind each model.
5-minute HDD model shares (percent of drives)
H1 0.32; H2 0.24; H3 0.04, all CMR. Unweighted mean 0.20, which Table 5 prints as 0.20.
Printed excess versus the quotient (times)
(1.41 - 0.20) / 0.20 = 6.05, matching the printed 6.05. The paper writes that multiple as 1.41% to 0.20% and does not print the subtraction. The quotient 1.41 / 0.20 = 7.05 is not printed. (0.52 - 0.01) / 0.01 = 51, matching the printed 51, written as 0.52% to 0.01%. The quotient 0.52 / 0.01 = 52 is not printed. H2's own 60-minute cell is printed as <0.01%; the HDD average row is 0.01%. The 15-minute and 30-minute multiples are not printed, so those cells say not stated. (0.96 - 0.05) / 0.05 and (0.75 - 0.03) / 0.03 are not filled in.
Fleet share versus the model average (percent of drives)
4,584 / 779,978 = 0.5877%, which Table 7 prints as 0.59%. That is not 1.41%. The text says around 5,000 fail-slow NVMe SSDs; Table 7 prints 4,584.
Failure-ticket transition
10 / 4,584 = 0.218%, which section 5.4 calls around 0.22% of slow drives. The abstract also prints 0.22% of the fail-slow drives. 10 / 4,439 = 0.225%, called around 0.23% of replaced drives. 10 / 779,978 = 0.00128%, and the slow-and-replaced cell is printed as <0.01% of drives scored. Mean gap 73 days. Median gap 67 days. Latest tickets are about five months after the last fail-slow event.
Event frequency, same excess pattern (events per 1,000 drives per hour)
Table 5 average row at 5 minutes: NVMe 48.00, HDD 2.45. (48.00 - 2.45) / 2.45 = 18.59, matching the printed 18.59 times. At 60 minutes: NVMe 5.56, HDD 0.08. (5.56 - 0.08) / 0.08 = 68.5, matching the printed 68.50 times. The quotients 48.00 / 2.45 = 19.59 and 5.56 / 0.08 = 69.5 are not printed.
Average event latency (microseconds)
Printed NVMe average row: 174.40 at 5 minutes, 155.17 at 15, 159.02 at 30, and 149.58 at 60. The paper summarizes those rows as around 160 microseconds and calls that SATA-SSD-level latency. It also says the top 1% of slowest events in several models reached around 22 milliseconds. These are iostat write latencies, not a link rate.
Vendor repair results, as the paper reports them
Not a fleet rate. The authors sent the 100 slowest SSDs, around the top 2% of identified slow drives, average event latency 4.4 milliseconds, back to vendors. The paper says 33 had bad capacitors, 46 had bad chips, and the root causes of the rest were unclear.

Method

The figures are copied from Lu et al., USENIX ATC 2022, not refit on another fleet. Section 5.1 defines fail-slow on write latency. A drive is suspicious when its 3-hour median write latency is above the cluster bar, the third quartile plus two interquartile ranges, which the paper describes as the top 0.04% of latency variances in the cluster. Drives whose IOPS or throughput also clears that bar are dropped as heavy traffic. The drive then has to show one event whose mean relative latency is at least 2. Relative latency is that drive's 15-second write latency divided by the median of the drive and its 11 intra-node peers (12 drives per node). The event must last longer than 20, 60, 120, or 240 records, and the paper labels those spans 5, 15, 30, and 60 minutes. Table 5's average rows are unweighted means of eight NVMe models or three HDD models. Table 7 is the contingency table for the 5-minute bar only, slow drives against later failure tickets. The checks under the table recompute the printed 6.05 and 51 from the two percentages in each i.e. clause, and they recompute the unweighted mean of the eight 5-minute model shares. They are checks on this page. The paper does not print the subtraction. Dong et al., USENIX ATC 2025, is used only as a later citation of the 1.41% sentence.

Limits

The fail-slow population is Alibaba's iostat study: more than half a million enterprise NVMe SSDs and more than 4 million HDDs, four months of monitoring. Table 1's perf logs run from 2020-11-16 to 2021-03-05, about 84 million entries. The daemon runs from 9 PM to 12 AM and stores a 15-second average, 720 records per drive per day. The NVMe drives sit in all-flash nodes, 12 per node, not RAIDed, and not hosting the OS, in multiple IDCs. The earliest model was deployed around May 2015 and the latest in July 2019. Table 7 scores 779,978 NVMe SSDs under the 5-minute bar. The whole study, not this fail-slow slice, uses logs from over one million NVMe SSDs, and the perf daemon covers only a subset of clusters. SMART logs run from 2019-11-04 to 2020-11-14 (about 1.8 million entries). Failure tickets run from 2019-11-04 to 2020-11-02 (about 20 thousand entries). In the transition check, the latest tickets are about five months after the last recorded fail-slow event. Not measured: read-latency fail-slow (the paper studies write latency, because a write waits for every replica); the other 21 hours of the day; drives dropped for heavy traffic; per-model drive counts, so the 1.41% cell is not a drive-weighted share of the 779,978; the Spearman coefficients, which are not printed; an annual fail-slow rate (the paper expects it to be higher and does not give one); any fleet other than Alibaba; and whether a fail-slow drive becomes fail-stop after that roughly five-month follow-up. The artifact appendix says the released performance logs cover around 97,000 NVMe SSDs and 141,000 SATA HDDs, and that Findings 1-8 and 10 can be checked from that release. Finding 9 is not in that list. The PDF points at https://tianchi.aliyun.com/dataset/dataDetail?dataId=128972 and calls the dataset open-source. It does not name a license. This note does not reproduce the logs.

What the signal catches, and what it misses

The 1.41% figure is a model average, not the share of drives scored in Table 7. The abstract says that on average 1.41% of NVMe SSDs are fail-slow within four-month monitoring, which is 6.05 times the HDD figure. Section 5.2 then ties that 1.41% to the 5-minute requirement and writes the multiple as 1.41% to 0.20%. At 60 minutes it writes 0.52% to 0.01%, 51 times. The reading that reproduces those two multiples is the excess over the HDD average, not the quotient: 7.05 and 52 are not in the paper. Table 7, under the same 5-minute bar, counts 4,584 slow drives out of 779,978 and prints that margin as 0.59%. The identification prose says around 5,000 fail-slow NVMe SSDs. Dong et al. restate the 1.41% sentence, place over one million NVMe SSDs in parentheses in that same sentence, and cite Lu et al. That paper does not restate the SMART result or the 0.22% transition. It is a citation, not a second measurement. Event frequency follows the same excess pattern: 48.00 versus 2.45 events per 1,000 drives per hour at 5 minutes, printed as 18.59 times, and 5.56 versus 0.08 at 60 minutes, printed as 68.50 times.

Fail-slow here is a write-latency fault, and the signal that catches it is a peer comparison, not SMART. Section 5.1 marks a drive suspicious when the 3-hour median write latency is above the third quartile plus two interquartile ranges. The paper calls a drive over that bar higher than the top 0.04% of latency variances in the cluster. A drive whose IOPS or throughput is also over the bar is dropped, so heavy traffic is not counted. The suspicious drive still needs one event at least twice as slow as the median of itself and its 11 intra-node peers, held for more than 20, 60, 120, or 240 records of 15 seconds (the 5, 15, 30, and 60 minute bars). The daemon records only those averages, and only from 9 PM to 12 AM. The log can prove that, inside that window, the median sat over the bar and a span cleared the 2 times peer rule. It cannot prove the drive was slow in the other 21 hours, and it never sees the drives the heavy-traffic rule dropped or the slowdowns that miss the 2 times bar or the minimum length. Section 6 says fail-slow has no clear oracle, the thresholds are empirical, and the strict bar may underestimate the impact. The paper expects the annual rate to be higher than this four-month, three-hours-a-day figure, and it does not print that annual rate. The same table's NVMe average event latency is 174.40, 155.17, 159.02, and 149.58 microseconds at the four bars. The paper calls that around 160 microseconds, SATA-SSD-level latency, and says the top 1% of events in several models reached around 22 milliseconds. Those are measured write latencies, not a bus rate.

The signal that does not catch this fault is SMART. Finding 9 says SMART attributes only exhibit negligible correlation with fail-slow metrics. On the last day of the four-month window the authors grouped drives by model, age, program/erase cycle, and workload, labeled each drive slow or not-slow, and ran a Spearman rank correlation. Under the 5-minute bar that is 40 groups. None showed a clear correlation with any SMART attribute. The paper names Critical Warning, P/E error, and CRC error as examples, and it does not print a coefficient for any attribute. The same null result held at 15, 30, and 60 minutes, and when slow meant more than one event. The artifact appendix says Findings 1-8 and 10 can be checked from the released dataset. Finding 9 is not on that list. The released performance logs cover around 97,000 NVMe SSDs and 141,000 SATA HDDs, not the half-million drives in the iostat study.

A later failure ticket also misses almost all of these drives inside the window the paper has. Table 7's replaced column is drives that appear in the failure tickets. Ten drives are in both the slow row and that column. Section 5.4 calls that around 0.22% of the 4,584 slow drives and around 0.23% of the 4,439 replaced drives, and the cell is printed as <0.01% of the 779,978 drives scored. The abstract's 0.22% is that slow-drive share, not a fleet rate. Mean and median gaps are 73 and 67 days. The latest tickets are about five months older than the last recorded fail-slow event. Finding 10 says the transition was rarely observed, at least not within about five months. That is not a claim that fail-slow never becomes fail-stop. The ticket shows a later failure record for those 10 drives. It does not show that the other 4,574 stayed healthy, and it does not show that the slowdown caused the ticket. Separately, the paper reports vendor repair results for the 100 slowest SSDs, around the top 2% of identified slow drives, average event latency 4.4 milliseconds: 33 bad capacitors, 46 bad chips, and the rest unclear. Those are the vendors' findings as the paper reports them, not a rate for the 4,584.

These counts are forNVMeSSDs. TheSMART failure-signal noteis a different population: Google ATA disks from 2005 to 2006, not this Alibaba fleet.

Sources

  1. 01NVMe SSD Failures in the Field: the Fail-Stop and the Fail-Slow, Lu, Xu, Zhang, Zhu, Wang, Zhu, Xue, Li, and Wu, USENIX ATC 2022 (open-access PDF, pages 1005-1020) · accessed 2026-10-10
  2. 02Understanding and Detecting Fail-Slow Hardware Failure Bugs in Cloud Systems, Dong, Hua, Zhang, Chen, and Chen, USENIX ATC 2025 (open-access PDF). Cites the 1.41% sentence from Lu et al. and is not a second measurement. · accessed 2026-10-10