Skip to content

Paper note

34.4% of RASR failures are not the SSD: 20.5% human mistakes, 13.9% cables, 31.2% labeled Failed Device

Published 2026-10-10

In Xu et al., USENIX ATC 2019, 34.4% of Reported As “SSD-Related” (RASR) failures are not caused by the SSD device. That 34.4% is human mistakes (20.5%) plus faulty cables (13.9%), inferred from which repair worked. It is not the complement of a measured device failure: Table 11 labels 31.2% Failed Device because replacing the SSD is the last resort, 11.9% Transient, and 22.5% Undetermined (FSCK 16.5% plus Data Check 6.0%). The paper prints 20.5 and 34.4. It does not print 22.5 or the addition 20.5 + 13.9.

Fleet
Around 450,000 SSDs in 7 Alibaba Cloud datacenters, three years of deployment, no calendar start or end printed. All five models use SATA and MLC NAND. The abstract says near half a million SSDs over three years.
Hardware tickets
Over 150K failure tickets. Table 4: SSD 5.6% (RASR, about 10K), HDD 22.1%, network 34.0%, unknown 23.7%, memory 8.5%, motherboard 5.4%, CPU 0.7%. No ticket has more than one type tag.
What a ticket records
Timestamp, related hardware component, and a log snippet. Tagged SSD-related when an SSD appears to be missing. Tagged Unknown when the logs name no clear component.
What a repair log records
Symptom, diagnosis procedure, and the successful fix. The root cause is inferred. Replacing the SSD is the last resort, after cheaper attempts.
Denominator of the 34.4%
Percent of RASR failures, not percent of drives and not percent of all hardware tickets. The 6-month OIOP run and the 89 later Drive Unfound cases are not this set.

Comparison

Working fixes of RASR failures in Xu et al., USENIX ATC 2019, Table 11. RASR means Reported As “SSD-Related.” The population is the RASR subset of hardware failure tickets from around 450,000 SATA MLC SSDs in 7 Alibaba Cloud datacenters, spanning three years of deployment. The paper prints no calendar start or end. RASR is 5.6% of over 150K hardware tickets, about 10K instances. Each share is the percent of those RASR failures repaired by that fix, not a percent of drives and not a percent of all hardware tickets. The root cause is inferred from the successful fix. The repair log does not state it. Human mistakes are Slot Check 20.1 plus Mount Options Check 0.4, which section 6 prints as 20.5. Faulty cable is Replacing Cable, 13.9. Those two causes are the introduction’s 34.4% that are not caused by the SSD device. The paper does not print 20.5 + 13.9. Undetermined is FSCK 16.5 and Data Check 6.0; their sum 22.5 is not printed. Replacing SSD, 31.2, is labeled Failed Device and is the last resort. Units are in the column headers.
Working fixShare of RASR failures (percent of RASR failures)Root cause inferred from the successful fixSource
Rebooting11.9TransientATC 2019
Mount Options Check0.4Human MistakeATC 2019
FSCK16.5UndeterminedATC 2019
Data Check6.0UndeterminedATC 2019
Slot Check20.1Human MistakeATC 2019
Replacing Cable13.9Faulty CableATC 2019
Replacing SSD31.2Failed DeviceATC 2019

This table is a derived extract of Table 11. It is not a copy of the PDF and not the failure tickets or repair logs. The canonical file is the USENIX PDF in the source column. The checks recompute sums from the printed cells. The paper prints the 20.5% human-mistake total and the 34.4% introduction figure. It does not print 20.5 + 13.9, the 22.5% undetermined sum, or a Table 11 total of 100.

Checks on the printed figures

Table 11 sum (percent of RASR failures)
11.9 + 0.4 + 16.5 + 6.0 + 20.1 + 13.9 + 31.2 = 100.0. The paper does not print that total.
Human mistakes, printed 20.5 (percent of RASR failures)
Slot Check 20.1 + Mount Options Check 0.4 = 20.5, matching section 6. The paper describes those two fixes as plugging the device into the wrong slot and incorrect configuration. It does not print the addition.
The introduction’s 34.4 (percent of RASR failures)
20.5 + 13.9 = 34.4. Section 6 calls human mistakes and faulty cables the two main non-SSD causes. The introduction calls 34.4% not caused by the SSD device, with wrong drive slots (20.1%) as the example. The paper does not print 20.5 + 13.9.
Undetermined, not printed as 22.5 (percent of RASR failures)
FSCK 16.5 + Data Check 6.0 = 22.5. Both rows are labeled Undetermined. The paper does not print 22.5. Transient is the separate Rebooting row, 11.9, not part of that sum.
Not a device-failure complement (percent of RASR failures)
100.0 − 34.4 = 65.6, which is 11.9 + 22.5 + 31.2. The only Failed Device cell is Replacing SSD, 31.2. The paper does not print 65.6 and does not call the remainder device failures. Section 6 says 31.2% is the largest single fix and is still the last resort because of device and labor cost. Section 3.3 says every RASR failure is eventually fixed by replacing the SSD if earlier attempts do not work.
About 10K versus over 150K (count of tickets)
The contribution list says 5.6% are RASR failures (about 10K instances) among about over 150K failure tickets. Section 3.1 says over 150K tickets in total. 0.056 × 150,000 = 8,400. 10,000 / 0.056 is about 178,600, which is still over 150K. The paper does not print either product. Both the ticket total and the RASR count are approximate.
Table 4 sum (percent of hardware failure tickets)
CPU 0.7 + memory 8.5 + network 34.0 + motherboard 5.4 + HDD 22.1 + SSD 5.6 + unknown 23.7 = 100.0. The paper prints 22.1 + 5.6 = 27.7 for HDD and SSD together. Unknown is not a RASR outcome.
Table 3 symptom sum (percent of RASR failures)
Node Unbootable 2.6 + File System Unmountable 7.4 + Drive Unfound 53.7 + Buffer IO Error 17.3 + Media Error 19.0 = 100.0. Drive Unfound is the majority symptom. These are not Table 11 rows. Table 6’s within-symptom fix rates use a different denominator and are not copied here.
UCRC Heavy group, Drive Unfound only (percent of that group)
Table 12 Heavy: rebooting 4.4, slot check 7.6, replacing cable 70.6, replacing SSD 17.4. 4.4 + 7.6 = 12.0, matching the printed 12.0% fixed by the first two candidates. The four cells sum to 100.0. Light: 25.0 + 35.7 + 24.2 + 15.1 = 100.0, with cable replacement at 24.2. Section 6.1.2 sets the Heavy bar at 17 or more UCRC errors, the top 20%. These percents are not shares of all RASR failures.

Method

The cells are copied from Table 11 of Xu et al., USENIX ATC 2019, not refit on another fleet. Section 2.2: on an abnormal event the monitoring daemon reports a failure ticket with the timestamp, the related hardware component, and a log snippet. The ticket is tagged SSD-related when an SSD appears to be missing, and Unknown when the logs name no clear hardware component. Section 3.1 collects over 150K hardware failure tickets. RASR is 5.6% of them, about 10K instances, and no failure event is tagged with more than one type. Section 3.3: for each RASR failure, administrators try a pre-defined sequence of fix candidates until the failure disappears. The order is cost, then past effectiveness. Software fixes are preferred over slot check, cable replacement, and SSD replacement. All RASR failures are eventually fixed by replacing the SSD if the earlier attempts do not work. Section 6: the repair log does not explicitly state the root cause, so the paper infers a potential root cause from the successful fix. If a failure can only be fixed by replacing the SSD, the root cause is likely a failed SSD. Table 11 lists seven working fixes, the percent of RASR failures each one repaired, and that inferred cause. The introduction’s 34.4% is the two non-SSD causes in section 6: human mistakes, printed as 20.5% (Slot Check 20.1% and Mount Options Check 0.4%), and faulty cables, printed as 13.9% (Replacing Cable). The checks under the table recompute 20.1 + 0.4, 20.5 + 13.9, 16.5 + 6.0, and the seven-row sum from the printed cells. The paper prints 20.5 and 34.4. It does not print 22.5, the addition 20.5 + 13.9, or a row total of 100.

Limits

The population is the storage systems in 7 Alibaba Cloud datacenters: around 450,000 SSDs spanning three years of deployment (section 2.1). The introduction says around 450,000 SSDs over 3 years’ deployment, and the abstract says near half a million SSDs over three years of usage. The paper prints no calendar start or end. All five models in the dataset use SATA interfaces and MLC NAND cells. Table 11 percentages are shares of RASR failures, about 10K instances, not shares of the 450,000 drives and not shares of the over 150K hardware tickets. Unknown, 23.7% in Table 4, is a tag on that whole ticket set, not a RASR outcome. Table 6 rates are shares inside one symptom. Table 12 rates are shares of Drive Unfound inside the Heavy or Light UCRC group. Not measured: an explicit root-cause field; whether a failure an early fix cleared would also have been cleared by replacing the SSD; a device-failure rate taken independently of that cost-ordered procedure; the Spearman coefficient for UCRC, which section 6.1 says is a strong correlation and does not print; the sensitivity of the 17-error threshold, which section 6.1.2 leaves as future work; any fleet other than these Alibaba systems. The paper does not release the tickets or name a data license for them. Section 2.4 says the DFS and service software are not open-source. The 6-month One Interface One Purpose deployment (about 100K SSDs, 3 wrong-slot cases versus an average of 47 on a comparable system without it) and the 89 later Drive Unfound cases (21.1% less repair time, on-site engineer feedback, not included in the dataset) are later actions, not rows in Table 11. The proceedings header says open access to the ATC 2019 proceedings is sponsored by USENIX. No Creative Commons license is printed on the opened PDF. This note does not reproduce the tickets or the repair logs.

What the signal catches, and what it misses

The 34.4% figure is two inferred causes added together, not a count of drives that tested healthy. The introduction says that correlating the RASR failures with the repair logs shows a significant number (34.4%) are not caused by the SSD device, and that plugging SSDs into the wrong drive slots, a human mistake, accounts for 20.1%. Section 6 then names the two main non-SSD causes: human mistakes contribute 20.5%, Slot Check plus Mount Options Check, and faulty interconnections fixed by replacing the cable account for 13.9%. Slot Check 20.1 plus Mount Options Check 0.4 equals the printed 20.5, and 20.5 plus 13.9 equals the printed 34.4. The paper does not print those additions. The other Table 11 rows sit outside that 34.4: Rebooting 11.9 (Transient), FSCK 16.5 and Data Check 6.0 (Undetermined; 16.5 + 6.0 = 22.5, not printed), and Replacing SSD 31.2 (Failed Device). Together those are 65.6%. The paper does not print 65.6 and does not call it a device-failure rate. Replacing SSD is the largest single fix, and section 6 still calls it the last resort.

The faults inside the 34.4% are a wrong slot and a bad cable, and the daemon tags both as the SSD. Section 2.2 tags the ticket SSD-related when an SSD appears to be missing. The signal that catches the wrong-slot fault is Slot Check in the repair log. Section 6.2 says plugging the device into the wrong slot, fixed by Slot Check, accounts for 20.1% of RASR failures. That signal misses Mount Options Check, the 0.4% configuration mistakes, and every failure the slot check does not clear. The signal that catches the faulty cable, before the procedure spends its first two attempts, is the Ultra-DMA CRC count. Section 6.1 compares five device errors — UCRC, raw bit error rate, uncorrectable errors, program errors, and end-to-end errors — and says only UCRC has a strong correlation with faulty interconnection. The coefficient is not printed. Seventeen or more UCRC errors, the top 20% of drives, is the Heavy group. On Drive Unfound only, Table 12 shows 70.6% of that Heavy group fixed by replacing the cable, and 12.0% fixed by rebooting or a slot check. The same signal misses a UCRC count from a transient such as a voltage spike, the Light group (cable replacement succeeds for 24.2%), RASR symptoms other than those Drive Unfound rows, and the sensitivity of the threshold, which the paper leaves as future work.

A failure ticket can prove a timestamp, the hardware component the daemon recorded, and a log snippet. It cannot prove the root cause. If the logs name no clear component, the tag is Unknown: 23.7% of all hardware tickets in Table 4, beside CPU 0.7%, memory 8.5%, network 34.0%, motherboard 5.4%, HDD 22.1%, and SSD 5.6%. Section 3.1 says no failure event is tagged with more than one type. Section 2.4 says the daemons may fail to record an event, for example a network failure during log collection, or mistag one, for example a node crash tagged Unknown because the logs were insufficient. A repair log can prove the symptom, the diagnosis procedure, and which candidate in the sequence worked. Section 6 says it does not explicitly state the root cause. The inference the paper states is narrow: if a failure can only be fixed by replacing the SSD, the root cause is likely a failed SSD. What stayed unknown is the Undetermined pair, 22.5%, and the Transient row, 11.9%, where a reboot cleared the symptom and no component was isolated. Section 3.3 orders the candidates by cost and then by past effectiveness, prefers software fixes over slot check, cable replacement, and SSD replacement, and says every RASR failure is eventually fixed by replacing the SSD if the earlier attempts do not work. An early fix that worked was not followed by the later fixes, so the log does not show that those later fixes would have failed. Replacing the SSD shows that the earlier attempts did not work and that the machine was put back. It does not show a separate test that the drive was the only cause.

These shares are RASR tickets from around 450,000 SATA MLC SSDs in 7 Alibaba Cloud datacenters over three years of deployment. The paper states no calendar start or end. They are not a drive-failure rate for that fleet. They are not the later Alibaba NVMe fail-slow census, 4,584 of 779,978 drives with perf logs from 2020-11-16 to 2021-03-05. They are not the Google SMART counts on ATA disks from December 2005 through August 2006, and not the 101 fail-slow reports behind the 112 root-cause occurrences in Gunawi et al. Section 6.2’s One Interface One Purpose deployment, about 100K SSDs over 6 months, reports 3 wrong-slot RASR failures against an average of 47 on a comparable system that still used one interface for every purpose. That is a later population, not a remeasurement of the 20.1%. The 89 Drive Unfound cases and the 21.1% shorter repair time are engineer feedback on cases the paper says are not in the dataset. The DFS and the service software are not open source.

These RASR shares are SATA MLC tickets, not theNVMe fail-slow note(4,584 of 779,978 drives). TheSMART failure-signal noteis Google ATA disks from December 2005 through August 2006. Thefail-slow hardware noteis 101 reports and 112 root-cause occurrences, not this ticket set.

Sources

  1. 01Lessons and Actions: What We Learned from 10K SSD-Related Storage System Failures, Xu, Zheng, Qin, Xu, and Wu, Proceedings of the 2019 USENIX Annual Technical Conference, July 10–12, 2019, Renton, WA, USA, ISBN 978-1-939133-03-8, pages 961–975 (open-access PDF). Open access to the proceedings is sponsored by USENIX. No Creative Commons license is printed on the opened PDF. · accessed 2026-10-10