Skip to content

Stat analysis

When the cable, backplane or adapter fails instead of the disk: what link faults look like in logs and field data

Published 2026-10-10

In NetApp's 2008 field study of about 39,000 storage systems (about 1,800,000 disks, 44 months, January 2004 to August 2007), physical interconnect failures (host adapters, cables, shelf power and backplanes) made up 27 to 68% of storage subsystem failures, against 20 to 55% for the disks themselves, and in one class the whole subsystem failed at 4.6% a year while its disks failed at 0.9%. A link fault reaches the operator as a disk that vanishes or times out, so it is often replaced as a bad disk. The public signals that separate the two are specific: on Linux, an ICRC flag in the ATA error register and SError bits such as BadCRC point to the cable, connector or power path, a CRC count in SMART (attribute 199) counts interface errors, and a drive that is missing with no media errors points the same way. What these signals cannot do is name the failing part, because a bad cable, a bad port and a marginal power supply write the same bits; the kernel documentation says only 'often'.

Subsystem field evidence
About 39,000 NetApp systems, about 1,800,000 disks in about 155,000 shelf enclosures, 44 months from January 2004 to August 2007 (Jiang et al., FAST 2008).
Replacement-log evidence
Hardware replacement logs from HPC1, COM1 and COM2 systems, 2001 to 2006 (Schroeder and Gibson, FAST 2007).
Host-side signals
Linux libata error line (Emask, ICRC, SError bits) and SMART attribute 199 as documented in the opened pages.
Not covered
SAS expander and PHY counters, NVMe PCIe link errors, connector wear rates, and any link failure rate for hardware made after 2007.

Where storage subsystem failures came from in the NetApp logs

Failure events by type from Table 1 of Jiang et al. (FAST 2008), January 2004 to August 2007. The 'computed' columns are this page's division of the interconnect and disk event counts by the class total; they include the one problematic disk family the paper later excludes in its Figure 4(b), so they differ from the paper's 27 to 68% range, which is stated for the data without that family.
System classDisk failures (events)Physical interconnect failures (events)Interconnect share of 4 types, computed (%)Disk share of 4 types, computed (%)Source
Near-line (SATA)10,1054,88827.356.5FAST '08 paper
Low-end3,2304,33844.232.9FAST '08 paper
Mid-range8,9897,94937.342.2FAST '08 paper
High-end8,2407,39542.647.5FAST '08 paper

Disk failure rate is not the subsystem failure rate

Annualized failure rates (AFR) stated in the text of Jiang et al. Disk H, a problematic family, is excluded from the first two rows as in the paper's Figure 4(b). The multipathing rows compare one path with two independent paths in mid-range and high-end systems; the paper gives the physical interconnect AFR with its error term, and the mid-range and high-end pairs are assigned here in the order the text lists them.
ComparisonValue A (% per year)Value B (% per year)What the paper concludesSource
Low-end: disk AFR (A) versus whole subsystem AFR (B)0.94.6Disks are about 20% of the subsystem AFR; disk failure rate is not indicative of subsystem failure rateFAST '08 paper
Near-line: disk AFR (A) versus whole subsystem AFR (B)1.93.4Near-line has the higher disk AFR but the lower subsystem AFR than low-endFAST '08 paper
Mid-range interconnect AFR: single path (A) versus dual path (B)1.82 ± 0.040.91 ± 0.09Second path cuts interconnect failures by 50 to 60%FAST '08 paper
High-end interconnect AFR: single path (A) versus dual path (B)2.13 ± 0.070.90 ± 0.06Whole-subsystem AFR falls 30 to 40%; backplane faults are not covered by a second pathFAST '08 paper
Low-end interconnect AFR: shelf model A versus shelf model B, same disk model2.66 ± 0.232.18 ± 0.13Shelf model changes interconnect failures but not other types; difference significant at 99.5%FAST '08 paper

Cable and backplane lines in hardware replacement logs

Share of hardware replacements by component from Table 3 of Schroeder and Gibson (FAST 2007). Abbreviations come from service data and may not mean the same thing across sets. Only the rows relevant to the link are shown; the disk rows are included for scale.
Data setComponentShare of replacements (%)NoteSource
COM2Hard drive49.1Most replaced component in this setFAST '07 paper
COM2RAID card4.1Adapter in the I/O pathFAST '07 paper
COM2SCSI cable2.2Same share as fans and CPUs in this listFAST '07 paper
COM1SCSI board0.6Disk drives are 18.1% in the same listFAST '07 paper
HPC1SCSI backplane0.3Hard drives are 30.6% in the same listFAST '07 paper

One interconnect failure as the storage layers logged it

The example log in Figure 3 of Jiang et al., re-tabulated; the seconds column is this page's subtraction from the first timestamp. It shows the order of events in one system, not a measured delay for any other.
Clock time (PDT)Elapsed since first event (s)Event as loggedLayerSource
05:43:360fci.device.timeout: adapter 8 encountered a device timeout on device 8.24Fibre ChannelFAST '08 paper
05:43:5014fci.adapter.reset: resetting Fibre Channel adapter 8; scsi.cmd.abortedByHost on device 8.24Fibre Channel, SCSIFAST '08 paper
05:44:1236scsi.cmd.selectionTimeout: targeted device did not respond; I/O will be retriedSCSIFAST '08 paper
05:44:2246scsi.cmd.noMorePaths: no more paths to device, all retries have failedSCSIFAST '08 paper
05:46:22166raid.config.filesystem.disk.missing: disk 8.24 is missingRAIDFAST '08 paper

Host-side signals that point at the link, and what each leaves open

Signals as described in the opened Linux ATA wiki page (archived, marked obsolete) and the Wikipedia SMART attribute table. The wording in the 'what the source says' column is paraphrased closely; 'often' is the source's word.
SignalWhat the source saysWhat it cannot tellUnitSource
Emask 0x10 (ATA bus error)Chip-to-device bus error in the libata exception lineWhich part of the path failedbitmasklibata error wiki
ICRC in the error registerInterface CRC error during an Ultra DMA transfer, often a bad cable or power problem, possibly a wrong Ultra DMA mode set by the driverCable versus power versus driver settingflag per commandlibata error wiki
SError bits (BadCRC, 10B8B, Dispar, PHYRdyChg, CommWake)Set by the SATA host interface on link errors; not normal unless a drive was hot-plugged, and usually points strongly to hardware, often a bad cable or a bad or inadequate power supplyCable versus adapter versus power supplybits in a registerlibata error wiki
SMART 199, UltraDMA CRC Error CountCount of errors in data transfer via the interface cable as determined by ICRCWhether the count is still rising; the raw format is vendor-definedcountSMART attributes
SMART 183, SATA Downshift Error CountOn some Western Digital, Samsung and Seagate drives, number of link-speed downshifts such as 6 Gbit/s to 3 Gbit/s; on others the same ID counts runtime bad blocksWhich meaning applies without the vendor's definitioncountSMART attributes
Missing drive, no media errorsNetApp's physical interconnect failure type: affected disks appear to be missing from the systemAdapter, cable, enclosure power or backplaneeventFAST '08 paper

Reading the numbers

In the largest opened field data, a disk-only view misses more than a quarter of subsystem failures. NetApp's physical interconnect failures were 27 to 68% of subsystem failures and disk failures 20 to 55%. In the low-end class the subsystem AFR was 4.6% against 0.9% for disks. The same disk models showed disk AFRs that varied little between systems (disk D-2: 0.6 to 0.77%) while their subsystem AFRs varied from 2.2 to 4.9%, so the difference sat in the shelf, adapter and cabling around the disk. This is a field result for 2004 to 2007 Fibre Channel and SATA shelves; it shows that link faults were a major source of failures then, not what share they are in a given system now.

Link faults are often counted as disk faults. The NetApp authors note that system administrators often replace a disk when it becomes unavailable, and that replacement-rate studies therefore end up close to the subsystem failure rate, not the disk failure rate. Schroeder and Gibson add that after a drive is blamed, tests decide whether it is replaced, that a vendor reported finding no problem in 43% of returned disks (a personal communication quoted in the paper), and that the hardware lines for cables and backplanes in the replacement logs are small: 2.2% for SCSI cables in COM2, 0.3% for SCSI backplanes in HPC1. Both can be true: a replacement log lists the part that was swapped, so a swapped disk that was never faulty is counted as a disk.

The log trail of a link failure is quick at the top and slow at the bottom. In the NetApp example the adapter timed out, reset after 14 seconds, found the device not responding after 36 seconds, ran out of paths after 46 seconds, and the RAID layer logged the disk as missing 166 seconds after the first event. The authors say a physical interconnect failure can be recovered by retries or tolerated by multipathing, so many never reach the RAID layer; what is seen there is only the part that was exposed. They also report strong clustering: after one failure in a shelf, a second of the same type was 10 to 25 times more likely than independence predicts for the non-disk types, because the disks share adapters, cables and terminators. For an operator this means that several disks vanishing from one shelf or one controller at nearly the same time is a link or shelf signature, not a coincidence of disks.

On a Linux host the useful discriminators are an ICRC flag, SError link bits, the Emask value and SMART attribute 199 reported together with no media errors, which the kernel documentation says often means a cable, connector or power problem. They cannot go further. The same bits can come from an inadequate power supply, a wrong transfer mode or an adapter, and no opened source says how often each cause occurs. Two caveats apply to the SMART side: attribute meanings are manufacturer-specific, and ID 183 means link downshifts on some drives and runtime bad blocks on others. A redundant path reduced interconnect failures by 50 to 60% in the NetApp data but did not remove them, because the authors name shelf backplane errors, which a second path does not cover, and a shared physical host adapter behind two logical ones, as likely reasons.

Related: SCSI sense and kernel-log signatures, what fails most often on HDDs, SSDs and links and the SMART failure-signal note.

Method

Everything here was read on 2026-10-10 from documents opened that day: the FAST 2008 paper by Jiang, Hu, Zhou and Kanevsky (full text, sections 2 to 5 and Table 1); the FAST 2007 paper by Schroeder and Gibson (Table 3 and sections 2.1 and 6); the archived Linux ATA wiki page on libata error messages; and the English Wikipedia article on S.M.A.R.T. (attribute table). Numbers are copied from the text and tables of those documents. The two shares marked as computed in the first table are this page's own arithmetic on the event counts in Table 1 of the NetApp paper, not figures the paper prints. Values that exist only inside figures are not used.

Limits

Both field studies are old: the NetApp logs end in August 2007 and cover Fibre Channel and SATA shelves behind NetApp controllers, and the Schroeder data end in 2006 and are hardware replacement logs from HPC and internet-service clusters. Neither covers SAS expanders, NVMe over PCIe, or modern SATA backplanes, and no opened source gives a link failure rate for any of those. Replacement logs record what was swapped, not the root cause (Schroeder and Gibson say so), so a cable swap may be missing from the disk share. The kernel text is from an archived wiki marked obsolete and describes causes with the word 'often'; no opened source quantifies how often an ICRC flag is a cable and how often it is power or a driver setting. The Wikipedia table is a secondary summary: the vendor definitions it cites were not opened, and it says the meaning of attributes varies by manufacturer. The log timeline is a figure's example from one system, not a measured distribution. SAS-specific link counters were not read and are not covered. The page covers public failure analysis only; nothing here concerns encryption, drive security features, or recovering anyone else's media.

Sources

  1. 01Jiang, Hu, Zhou, Kanevsky: Are Disks the Dominant Contributor for Storage Failures? FAST 2008 (paper PDF) · accessed 2026-10-10
  2. 02Schroeder, Gibson: Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you? FAST 2007 (HTML paper) · accessed 2026-10-10
  3. 03Libata error messages, Linux ATA wiki (archived, marked obsolete content) · accessed 2026-10-10
  4. 04Self-Monitoring, Analysis and Reporting Technology, Wikipedia (ATA attribute table) · accessed 2026-10-10