Skip to content

Paper note

Network is the largest fail-slow hardware type: 29 of 112 occurrences in 101 reports

Published 2026-10-10

Network is the largest hardware type in Table 3 of Gunawi et al., FAST 2018: 29 of 112 root-cause occurrences, ahead of CPU 25, disk 23, SSD 22, and memory 13. Those 112 occurrences come from 101 fail-slow reports at 12 institutions, incidents reported from 2000 through 2017, and a report can name more than one cause. Device errors account for 40 occurrences and firmware issues for 20. Temperature, power, environment, and configuration sum to 44 of 112 (39.3%), which rounds to the proceedings’ statement that 39% of root causes are external factors, not 39% of the 101 reports.

Reports
101 operator narratives of unexpected fail-slow hardware from 12 institutions. Incidents reported between 2000 and 2017. Only 30 reports predate 2010. One institution-level root cause is one report even if it happened many times.
Occurrences
112 root-cause cells in proceedings Table 3. Larger than 101 because a report can name more than one root cause. Not a second study and not a count of drives.
Deployments
Proceedings Table 2, the same node counts as ;login: Table 1. Companies: more than 10,000, 150, 100, more than 1,000, more than 10,000 nodes. Universities: 300, more than 100, more than 1,000, 500. National labs: more than 1,000, more than 10,000, more than 10,000.
Excluded slowdowns
Ordinary random I/O that slows disks, and routine SSD garbage collection. Only unexpected degradation is in the 101.
Edition used for the table
FAST 2018 proceedings and the ;login: Summer 2018 reprint. Not the journal edition’s 114 reports from 14 institutions.

Comparison

Root-cause occurrences by hardware type in Gunawi et al., FAST 2018 proceedings Table 3, also printed in ;login: Summer 2018 Table 3. The population is 101 operator reports of unexpected fail-slow hardware from 12 institutions. Incidents were reported between 2000 and 2017, and only 30 reports predate 2010. Deployments range from 100 nodes to more than 10,000. Each cell is a count of root-cause occurrences, not a percent of reports, not a percent of drives, and not a failure rate. A report can name more than one cause, so the total row is 112 rather than 101. Unknown means the operators could not pinpoint the cause and replaced the hardware. The paper’s headers are SSD, Disk, Mem, Net, and CPU. Units are in the column headers.
Hardware typeDevice errors (count of root-cause occurrences)Firmware issues (count of root-cause occurrences)Temperature (count of root-cause occurrences)Power (count of root-cause occurrences)Environment (count of root-cause occurrences)Configuration (count of root-cause occurrences)Unknown, hardware replaced (count of root-cause occurrences)All causes (count of root-cause occurrences)Source
SSD1061131022FAST 2018
Disk833051323FAST 2018
Memory900120113FAST 2018
Network1092042229FAST 2018
CPU325643225FAST 2018
All five types40201181878112FAST 2018

This table is a derived extract of open-access USENIX proceedings Table 3. ;login: Summer 2018 Table 3 prints the same cells. It is not a copy of either PDF, not the raw dataset, and not the later journal edition. The canonical proceedings file is the USENIX PDF in the source column. The checks under the table recompute row and column sums and 44/112 from those cells. The paper does not print the division.

Checks on the printed figures

Column sums (count of root-cause occurrences)
SSD 10+6+1+1+3+1+0 = 22. Disk 8+3+3+0+5+1+3 = 23. Memory 9+0+0+1+2+0+1 = 13. Network 10+9+2+0+4+2+2 = 29. CPU 3+2+5+6+4+3+2 = 25. Those are the printed column totals.
Row sums (count of root-cause occurrences)
Device errors 10+8+9+10+3 = 40. Firmware 6+3+0+9+2 = 20. Temperature 1+3+0+2+5 = 11. Power 1+0+1+0+6 = 8. Environment 3+5+2+4+4 = 18. Configuration 1+1+0+2+3 = 7. Unknown 0+3+1+2+2 = 8. Grand total 40+20+11+8+18+7+8 = 112.
External share versus the printed 39 percent
Temperature 11, power 8, environment 18, and configuration 7 sum to 44. 44/112 = 0.39286, which is 39.3% to one decimal place and rounds to 39%. Proceedings Table 1 prints “39% root causes are external factors.” The paper does not print 44 or the division. Unknown is not in that sum: 8 unknown occurrences would make 52/112, which is 46.4% and does not match 39%. 39% of 101 reports would be about 39 reports. That count is not printed. ;login: also says “39% of the cases were caused by external root causes,” which is a different wording and is not treated here as a count of the 101 reports.
Two different “not pinpointed” figures
Section 2 says 12% of reports did not pinpoint a root cause. Twelve percent of 101 is about 12 reports. Table 3’s unknown row is 8 occurrences, where operators replaced the hardware without a cause. This page does not treat 8 and 12% as the same quantity. The firmware row of 20 is also not the “firmware bugs from 5 institutions” example in that section.
Introduction examples, not table cells
Disk throughput to 100 KB/s from vibration, described as a drop of three orders of magnitude. SSD operations stalling for seconds from firmware bugs. Memory cards at 25% of normal speed from a loose NVDIMM connection. CPUs at 50% speed from lack of power. Network-card performance at Kbps from buffer corruption and retransmission. Section 7 says the anecdotes could not answer how much performance was degraded across the reports. These are not bus rates and not measured throughputs for the 101.

Method

The cells are copied from Table 3 of Gunawi et al. in the FAST 2018 proceedings, not refit on logs. ;login: Summer 2018 Table 3 prints the same numbers, with the rows spelled out as device errors, firmware bugs, temperature, power, environment, configuration, and unknown. Section 2 of the proceedings collects 101 operator reports of unexpected degradation from 12 institutions. Incidents were reported between 2000 and 2017, and only 30 reports predate 2010. A repeated root cause at one institution is one report, so one report can stand for many instances. The same root cause at another institution is counted again. The proceedings say 66% of root causes are unique, 22% are duplicates, and 12% of reports did not pinpoint a root cause. A duplicated incident is reported by 2.4 institutions on average. The example in that sentence is firmware bugs from 5 institutions, driver bugs from 3, and the remaining duplicated issues from 2. Those are institution counts, not the firmware row of 20. Known slowdowns are excluded, including random I/O that slows disks and ordinary SSD garbage collection. There were no analyzable hardware-level performance logs. Internal rows are device errors and firmware issues. External rows are temperature, power, environment, and configuration. Unknown means the operators could not pinpoint the cause and replaced the hardware. A report can carry more than one root cause, which is why the table totals 112 occurrences rather than 101 reports. The paper’s column order is SSD, disk, memory, network, and CPU. The 44 and 44/112 on this page are the sum of the four external rows. The paper does not print that division. Proceedings Table 1 prints “39% root causes are external factors.” ;login: prints both “Thirty-nine percent of root causes are external” and, in the online-diagnosis paragraph, “39% of the cases were caused by external root causes.”

Limits

The population is those 101 reports, not a drive census. The 12 deployments are proceedings Table 2, repeated as ;login: Table 1: five companies (more than 10,000, 150, 100, more than 1,000, and more than 10,000 nodes), four universities (300, more than 100, more than 1,000, and 500), and three national labs (more than 1,000, more than 10,000, and more than 10,000). The date range of the reports is 2000–2017, with 30 reports before 2010. Not measured: how often fail-slow hardware occurs, how much performance was degraded across the reports, or a correlation with device age, model, size, or vendor. Proceedings Section 7, Limitations (and Failed Attempts), calls the lack of quantitative analysis the major limitation and says the reports are anecdotes. The introduction’s examples — disk throughput to 100 KB/s from vibration, SSD stalls of seconds from firmware bugs, a network card collapsing to Kbps from buffer corruption and retransmission, memory at 25% of normal speed from a loose NVDIMM connection, and CPUs at 50% speed from lack of power — are single stories, not Table 3 cells and not a measured distribution. The author PDF of the journal version is a different edition: 114 reports from 14 institutions, incidents reported between 2000 and 2018, and “32% root causes are external factors,” while one sentence still says the Table 3 total (112) is larger than the 114 reports. Its header is an unfilled ACM template (Vol. 9, No. 4, Article 39, March 2010, © 2009). A text extract of DOI 10.1145/3242086 shows ACM Transactions on Storage 14(3), article 23, published 3 October 2018, free access, and an abstract of 114 reports from 14 institutions. The HTML landing page returned an interstitial and was not read as a full page. This note does not retabulate that edition. Republishing the journal text, including its table, requires prior permission from permissions@acm.org; abstracting with credit is permitted. The FAST ’18 errata slip adds a funding acknowledgement for this paper and does not change Table 3. No Creative Commons license is printed on the opened proceedings pages, the ;login: article, or the conference page. USENIX says the proceedings are freely available once the event begins. Both USENIX texts point at a raw partial dataset at http://ucare.cs.uchicago.edu/projects/failslow/. This note does not republish that file and does not verify its license or contents.

What the signal catches, and what it misses

The named comparison is root-cause occurrences, not a failure rate. Network is the largest hardware type at 29 of 112, then CPU 25, disk 23, SSD 22, and memory 13. Device errors (40) and firmware issues (20) are the two largest cause rows. On SSD, disk, and network the device-error cells are 10, 8, and 10, and the firmware cells are 6, 3, and 9. Memory is not a firmware story in this table: 9 device errors and 0 firmware. CPU is not mainly an internal-device story either: the CPU column is device errors 3, firmware 2, temperature 5, power 6, environment 4, configuration 3, and unknown 2. The 112 is the cell total. The caption of Table 3 says a report can have multiple root causes, environment and power or temperature among them, so 112 is larger than the 101 reports.

The external fault is a device that is functionally up but slow because of temperature, power, environment, or configuration. Proceedings Table 1 says 39% of root causes are external factors, so troubleshooting fail-slow hardware must be done online. The signal that catches that class is online diagnosis with full-stack monitoring, not a bench retest. ;login: says some reports suggest operators took days or even months because the problem could not be reproduced in offline testing, and that not all hardware is monitored. Its example is network cables: several organizations watched traffic flow as a proxy for cable health instead of monitoring the cables, and blame went to the main components first. The proceedings’ own network examples of external causes include loose network cables and pinched fiber. Offline reproduction misses a cause that appears only in the live rack. Full-stack monitoring still misses whatever component nobody instrumented. Section 3.5 says pinpointing took hours or even months, for reasons that include no full-stack visibility, environment conditions, and cascading causes. The minute-to-month percentage split in that section is not a row in the table on this page.

The internal fault is a device error or wear-out, or a firmware issue. Proceedings Table 1 tells vendors that when error masking becomes more frequent, the device should throw a more explicit signal instead of running with a high overhead, and that device-level performance statistics should be collected and reported, for example via S.M.A.R.T. That suggestion is not a measurement of SMART in these 101 reports. Section 2 says there were no analyzable hardware-level performance logs, which is why this is an interview-style study rather than a log study. The signal that table cannot show is a counter that was never collected. What remains unknown is the unknown row: 8 occurrences in which operators could not pinpoint the cause and replaced the hardware. That is separate from the 12% of reports that did not pinpoint a root cause. Replacing the part shows that the machine was put back, not which fault it was, and not whether the same fault will recur. The SMART failure-signal note and the Alibaba NVMe fail-slow note are other populations. Their percentages are not rates for these 12 institutions.

One report can prove that an institution recorded an unexpected degradation and, when the operators named it, a root cause tied to a hardware type. It cannot prove how often that fault occurs, how large the slowdown was across the 101, or a tie to age, model, size, or vendor. Section 7 states those as the questions the anecdotes could not answer. The introduction’s 100 KB/s disk, seconds-long SSD stall, Kbps network card, 25% memory, and 50% CPU examples are illustrations of that kind of story, not columns in Table 3. The journal author PDF’s 114 reports, 14 institutions, and 32% external figure are a later edition with an unfilled ACM template header. They are not substituted into this table. The FAST ’18 errata slip does not revise the cells.

These 101 reports are not the Alibaba iostat study. TheNVMe fail-slow notecounts 4,584 slow drives out of 779,978 under a 5-minute bar. TheSMART failure-signal noteis Google ATA disks from December 2005 through August 2006. Neither population is these 12 institutions.

Sources

  1. 01Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems, Gunawi et al., Proceedings of the 16th USENIX Conference on File and Storage Technologies (FAST ’18), February 12–15, 2018, Oakland, CA, ISBN 978-1-931971-42-3, pages 1–14 (open-access PDF) · accessed 2026-10-10
  2. 02;login: Summer 2018, Vol. 43, No. 2, Fail-Slow at Scale, Gunawi et al. (USENIX PDF). Reprints the operational scale and Table 3. · accessed 2026-10-10
  3. 03Fail-Slow at Scale, USENIX FAST ’18 presentation page. Papers and proceedings are freely available once the event begins. No Creative Commons license is printed on the page. · accessed 2026-10-10
  4. 04FAST ’18 errata slip. Adds a Department of Energy acknowledgement for this paper (contract DE-AC02-06CH11357) and does not change Table 3. The figure correction on the slip belongs to a different paper. · accessed 2026-10-10
  5. 05Author PDF of Fail-Slow at Scale (cdn.jsdelivr.net, ucare-uchicago/ucare-html). Unfilled ACM template header: Vol. 9, No. 4, Article 39, March 2010, © 2009. Body text says 114 reports from 14 institutions and 32% external root causes. Not the source of the table on this page. Posting the journal text requires permission from permissions@acm.org. · accessed 2026-10-10
  6. 06Fail-Slow at Scale, ACM Transactions on Storage 14(3), article 23, DOI 10.1145/3242086. A text extract of this record shows free access, published 3 October 2018, and an abstract of 114 reports from 14 institutions. The HTML landing page returned an interstitial in this run. Not the source of the table on this page. · accessed 2026-10-10