When the cable, backplane or adapter fails instead of the disk: what link faults look like in logs and field data
Published 2026-10-10
In NetApp's 2008 field study of about 39,000 storage systems (about 1,800,000 disks, 44 months, January 2004 to August 2007), physical interconnect failures (host adapters, cables, shelf power and backplanes) made up 27 to 68% of storage subsystem failures, against 20 to 55% for the disks themselves, and in one class the whole subsystem failed at 4.6% a year while its disks failed at 0.9%. A link fault reaches the operator as a disk that vanishes or times out, so it is often replaced as a bad disk. The public signals that separate the two are specific: on Linux, an ICRC flag in the ATA error register and SError bits such as BadCRC point to the cable, connector or power path, a CRC count in SMART (attribute 199) counts interface errors, and a drive that is missing with no media errors points the same way. What these signals cannot do is name the failing part, because a bad cable, a bad port and a marginal power supply write the same bits; the kernel documentation says only 'often'.
- Subsystem field evidence
- About 39,000 NetApp systems, about 1,800,000 disks in about 155,000 shelf enclosures, 44 months from January 2004 to August 2007 (Jiang et al., FAST 2008).
- Replacement-log evidence
- Hardware replacement logs from HPC1, COM1 and COM2 systems, 2001 to 2006 (Schroeder and Gibson, FAST 2007).
- Host-side signals
- Linux libata error line (Emask, ICRC, SError bits) and SMART attribute 199 as documented in the opened pages.
- Not covered
- SAS expander and PHY counters, NVMe PCIe link errors, connector wear rates, and any link failure rate for hardware made after 2007.
Where storage subsystem failures came from in the NetApp logs
| System class | Disk failures (events) | Physical interconnect failures (events) | Interconnect share of 4 types, computed (%) | Disk share of 4 types, computed (%) | Source |
|---|---|---|---|---|---|
| Near-line (SATA) | 10,105 | 4,888 | 27.3 | 56.5 | FAST '08 paper |
| Low-end | 3,230 | 4,338 | 44.2 | 32.9 | FAST '08 paper |
| Mid-range | 8,989 | 7,949 | 37.3 | 42.2 | FAST '08 paper |
| High-end | 8,240 | 7,395 | 42.6 | 47.5 | FAST '08 paper |
Disk failure rate is not the subsystem failure rate
| Comparison | Value A (% per year) | Value B (% per year) | What the paper concludes | Source |
|---|---|---|---|---|
| Low-end: disk AFR (A) versus whole subsystem AFR (B) | 0.9 | 4.6 | Disks are about 20% of the subsystem AFR; disk failure rate is not indicative of subsystem failure rate | FAST '08 paper |
| Near-line: disk AFR (A) versus whole subsystem AFR (B) | 1.9 | 3.4 | Near-line has the higher disk AFR but the lower subsystem AFR than low-end | FAST '08 paper |
| Mid-range interconnect AFR: single path (A) versus dual path (B) | 1.82 ± 0.04 | 0.91 ± 0.09 | Second path cuts interconnect failures by 50 to 60% | FAST '08 paper |
| High-end interconnect AFR: single path (A) versus dual path (B) | 2.13 ± 0.07 | 0.90 ± 0.06 | Whole-subsystem AFR falls 30 to 40%; backplane faults are not covered by a second path | FAST '08 paper |
| Low-end interconnect AFR: shelf model A versus shelf model B, same disk model | 2.66 ± 0.23 | 2.18 ± 0.13 | Shelf model changes interconnect failures but not other types; difference significant at 99.5% | FAST '08 paper |
Cable and backplane lines in hardware replacement logs
| Data set | Component | Share of replacements (%) | Note | Source |
|---|---|---|---|---|
| COM2 | Hard drive | 49.1 | Most replaced component in this set | FAST '07 paper |
| COM2 | RAID card | 4.1 | Adapter in the I/O path | FAST '07 paper |
| COM2 | SCSI cable | 2.2 | Same share as fans and CPUs in this list | FAST '07 paper |
| COM1 | SCSI board | 0.6 | Disk drives are 18.1% in the same list | FAST '07 paper |
| HPC1 | SCSI backplane | 0.3 | Hard drives are 30.6% in the same list | FAST '07 paper |
One interconnect failure as the storage layers logged it
| Clock time (PDT) | Elapsed since first event (s) | Event as logged | Layer | Source |
|---|---|---|---|---|
| 05:43:36 | 0 | fci.device.timeout: adapter 8 encountered a device timeout on device 8.24 | Fibre Channel | FAST '08 paper |
| 05:43:50 | 14 | fci.adapter.reset: resetting Fibre Channel adapter 8; scsi.cmd.abortedByHost on device 8.24 | Fibre Channel, SCSI | FAST '08 paper |
| 05:44:12 | 36 | scsi.cmd.selectionTimeout: targeted device did not respond; I/O will be retried | SCSI | FAST '08 paper |
| 05:44:22 | 46 | scsi.cmd.noMorePaths: no more paths to device, all retries have failed | SCSI | FAST '08 paper |
| 05:46:22 | 166 | raid.config.filesystem.disk.missing: disk 8.24 is missing | RAID | FAST '08 paper |
Host-side signals that point at the link, and what each leaves open
| Signal | What the source says | What it cannot tell | Unit | Source |
|---|---|---|---|---|
| Emask 0x10 (ATA bus error) | Chip-to-device bus error in the libata exception line | Which part of the path failed | bitmask | libata error wiki |
| ICRC in the error register | Interface CRC error during an Ultra DMA transfer, often a bad cable or power problem, possibly a wrong Ultra DMA mode set by the driver | Cable versus power versus driver setting | flag per command | libata error wiki |
| SError bits (BadCRC, 10B8B, Dispar, PHYRdyChg, CommWake) | Set by the SATA host interface on link errors; not normal unless a drive was hot-plugged, and usually points strongly to hardware, often a bad cable or a bad or inadequate power supply | Cable versus adapter versus power supply | bits in a register | libata error wiki |
| SMART 199, UltraDMA CRC Error Count | Count of errors in data transfer via the interface cable as determined by ICRC | Whether the count is still rising; the raw format is vendor-defined | count | SMART attributes |
| SMART 183, SATA Downshift Error Count | On some Western Digital, Samsung and Seagate drives, number of link-speed downshifts such as 6 Gbit/s to 3 Gbit/s; on others the same ID counts runtime bad blocks | Which meaning applies without the vendor's definition | count | SMART attributes |
| Missing drive, no media errors | NetApp's physical interconnect failure type: affected disks appear to be missing from the system | Adapter, cable, enclosure power or backplane | event | FAST '08 paper |
Reading the numbers
In the largest opened field data, a disk-only view misses more than a quarter of subsystem failures. NetApp's physical interconnect failures were 27 to 68% of subsystem failures and disk failures 20 to 55%. In the low-end class the subsystem AFR was 4.6% against 0.9% for disks. The same disk models showed disk AFRs that varied little between systems (disk D-2: 0.6 to 0.77%) while their subsystem AFRs varied from 2.2 to 4.9%, so the difference sat in the shelf, adapter and cabling around the disk. This is a field result for 2004 to 2007 Fibre Channel and SATA shelves; it shows that link faults were a major source of failures then, not what share they are in a given system now.
Link faults are often counted as disk faults. The NetApp authors note that system administrators often replace a disk when it becomes unavailable, and that replacement-rate studies therefore end up close to the subsystem failure rate, not the disk failure rate. Schroeder and Gibson add that after a drive is blamed, tests decide whether it is replaced, that a vendor reported finding no problem in 43% of returned disks (a personal communication quoted in the paper), and that the hardware lines for cables and backplanes in the replacement logs are small: 2.2% for SCSI cables in COM2, 0.3% for SCSI backplanes in HPC1. Both can be true: a replacement log lists the part that was swapped, so a swapped disk that was never faulty is counted as a disk.
The log trail of a link failure is quick at the top and slow at the bottom. In the NetApp example the adapter timed out, reset after 14 seconds, found the device not responding after 36 seconds, ran out of paths after 46 seconds, and the RAID layer logged the disk as missing 166 seconds after the first event. The authors say a physical interconnect failure can be recovered by retries or tolerated by multipathing, so many never reach the RAID layer; what is seen there is only the part that was exposed. They also report strong clustering: after one failure in a shelf, a second of the same type was 10 to 25 times more likely than independence predicts for the non-disk types, because the disks share adapters, cables and terminators. For an operator this means that several disks vanishing from one shelf or one controller at nearly the same time is a link or shelf signature, not a coincidence of disks.
On a Linux host the useful discriminators are an ICRC flag, SError link bits, the Emask value and SMART attribute 199 reported together with no media errors, which the kernel documentation says often means a cable, connector or power problem. They cannot go further. The same bits can come from an inadequate power supply, a wrong transfer mode or an adapter, and no opened source says how often each cause occurs. Two caveats apply to the SMART side: attribute meanings are manufacturer-specific, and ID 183 means link downshifts on some drives and runtime bad blocks on others. A redundant path reduced interconnect failures by 50 to 60% in the NetApp data but did not remove them, because the authors name shelf backplane errors, which a second path does not cover, and a shared physical host adapter behind two logical ones, as likely reasons.
Related: SCSI sense and kernel-log signatures, what fails most often on HDDs, SSDs and links and the SMART failure-signal note.
Method
Everything here was read on 2026-10-10 from documents opened that day: the FAST 2008 paper by Jiang, Hu, Zhou and Kanevsky (full text, sections 2 to 5 and Table 1); the FAST 2007 paper by Schroeder and Gibson (Table 3 and sections 2.1 and 6); the archived Linux ATA wiki page on libata error messages; and the English Wikipedia article on S.M.A.R.T. (attribute table). Numbers are copied from the text and tables of those documents. The two shares marked as computed in the first table are this page's own arithmetic on the event counts in Table 1 of the NetApp paper, not figures the paper prints. Values that exist only inside figures are not used.
Limits
Both field studies are old: the NetApp logs end in August 2007 and cover Fibre Channel and SATA shelves behind NetApp controllers, and the Schroeder data end in 2006 and are hardware replacement logs from HPC and internet-service clusters. Neither covers SAS expanders, NVMe over PCIe, or modern SATA backplanes, and no opened source gives a link failure rate for any of those. Replacement logs record what was swapped, not the root cause (Schroeder and Gibson say so), so a cable swap may be missing from the disk share. The kernel text is from an archived wiki marked obsolete and describes causes with the word 'often'; no opened source quantifies how often an ICRC flag is a cable and how often it is power or a driver setting. The Wikipedia table is a secondary summary: the vendor definitions it cites were not opened, and it says the meaning of attributes varies by manufacturer. The log timeline is a figure's example from one system, not a measured distribution. SAS-specific link counters were not read and are not covered. The page covers public failure analysis only; nothing here concerns encryption, drive security features, or recovering anyone else's media.
Sources
- 01Jiang, Hu, Zhou, Kanevsky: Are Disks the Dominant Contributor for Storage Failures? FAST 2008 (paper PDF) · accessed 2026-10-10
- 02Schroeder, Gibson: Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you? FAST 2007 (HTML paper) · accessed 2026-10-10
- 03Libata error messages, Linux ATA wiki (archived, marked obsolete content) · accessed 2026-10-10
- 04Self-Monitoring, Analysis and Reporting Technology, Wikipedia (ATA attribute table) · accessed 2026-10-10