Media error or link error? What SCSI sense codes and Linux ATA log lines can and cannot tell apart
Published 2026-10-10
A SATA drive that fails to read a sector and a SATA cable that corrupts a transfer are separated by two bits in the ATA error register: UNC (0x40, a media error, reported only after the drive's own retries) maps to SCSI sense key 03h MEDIUM ERROR with ASC/ASCQ 11h/04h, while ICRC with ABRT (0x84) maps to sense key 0Bh ABORTED COMMAND with 47h/00h and, in Linux, to the log text 'ATA bus error' (Emask 0x10). The separation is lossy: the Linux translation table collapses every link CRC into one SCSI parity code although the T10 list has separate codes for data-phase CRC (47h/01h) and UDMA CRC (08h/03h), a bus error clears the device and media bits of the same command, and a timeout (Emask 0x4) says only that the controller got no answer. Linux changes link speed after more than 3 bus or timeout errors in 10 minutes, but media errors are not counted by any of those rules.
- Path analysed
- SATA drive behind a Linux libata host: ATA status and error registers, then libata's own error mask, then the SCSI sense data the SCSI layer sees
- Standard reference
- T10 numeric ASC/ASCQ assignment list, as of 2024-11-28
- Kernel reference
- torvalds/linux master, HEAD 3857c2fe5449 at access on 2026-10-10
- Not covered
- SAS-native sense reporting from real drives, hardware RAID controllers, tape drive logs, any measured rates
From ATA error bits to SCSI sense data
| ATA register value (hex) | ATA meaning (include/linux/ata.h) | Sense key (hex) | ASC/ASCQ (hex) | T10 name of that code | Class | Source |
|---|---|---|---|---|---|---|
| error 0x40 | UNC: uncorrectable media error | 03h MEDIUM ERROR | 11h/04h | UNRECOVERED READ ERROR - AUTO REALLOCATE FAILED | media | Linux kernel source |
| error 0x84 | ICRC with ABRT: interface CRC error | 0Bh ABORTED COMMAND | 47h/00h | SCSI PARITY ERROR (T10 lists data-phase CRC separately as 47h/01h) | link | Linux kernel source |
| error 0x01 | AMNF: address mark not found | 03h MEDIUM ERROR | 13h/00h | ADDRESS MARK NOT FOUND FOR DATA FIELD | media | Linux kernel source |
| error 0x10 | IDNF: ID not found | 05h ILLEGAL REQUEST | 21h/00h | LOGICAL BLOCK ADDRESS OUT OF RANGE | addressing (a media-format cause is not distinguished) | Linux kernel source |
| status DF bit | Device fault | 04h HARDWARE ERROR | 44h/00h | INTERNAL TARGET FAILURE | device | Linux kernel source |
| status BSY bit | Device busy, error bits invalid | 0Bh ABORTED COMMAND | 00h/00h | NO ADDITIONAL SENSE INFORMATION | unknown | Linux kernel source |
Codes T10 assigns that separate link from media
| ASC/ASCQ (hex) | T10 name | Applies to disks (D) | Class | Source |
|---|---|---|---|---|
| 11h/00h | UNRECOVERED READ ERROR | yes | media | T10 ASC/ASCQ list |
| 11h/01h | READ RETRIES EXHAUSTED | yes | media | T10 ASC/ASCQ list |
| 0Ch/02h | WRITE ERROR - AUTO REALLOCATION FAILED | yes | media | T10 ASC/ASCQ list |
| 47h/01h | DATA PHASE CRC ERROR DETECTED | yes | link | T10 ASC/ASCQ list |
| 47h/03h | INFORMATION UNIT iuCRC ERROR DETECTED | yes | link | T10 ASC/ASCQ list |
| 08h/03h | LOGICAL UNIT COMMUNICATION CRC ERROR (ULTRA-DMA/32) | yes | link | T10 ASC/ASCQ list |
| 44h/00h | INTERNAL TARGET FAILURE | yes | device internal | T10 ASC/ASCQ list |
| 5Dh/00h | FAILURE PREDICTION THRESHOLD EXCEEDED | yes | prediction flag, not a fault | T10 ASC/ASCQ list |
What the kernel log line means
| Emask (hex) | String printed | Meaning per kernel wiki | Kernel action per libata-eh.c | Source |
|---|---|---|---|---|
| 0x10 | ATA bus error | Chip-to-device bus error | Reset scheduled; DEV, MEDIA and INVALID bits of the same command are cleared as probably spurious | Kernel ATA wiki |
| 0x8 | media error | Software detected a media error | Set from the UNC or AMNF error bit; no reset by this rule | Linux kernel source |
| 0x4 | timeout | Controller failed to respond to an active command; wiki says any number of causes | Reset scheduled; counted toward speed-down | Kernel ATA wiki |
| 0x2 | HSM violation | Hardware did not respond as the host state machine expects | Reset scheduled; counted toward speed-down | Kernel ATA wiki |
| 0x1 | device error | Error delivered directly by the drive; many of them often means a hardware problem | Counted as unknown-device error only if not MEDIA or INVALID | Kernel ATA wiki |
When Linux reacts to link-type errors
| Counted errors | Window (min) | Threshold (errors) | Action | Source |
|---|---|---|---|---|
| ATA bus or timeout/HSM | 10 | more than 3 | Lower the link speed | Linux kernel source |
| Timeout/HSM or unknown device | 10 | more than 3 | Turn NCQ off | Linux kernel source |
| Unknown device | 10 | more than 6 | Lower the link speed | Linux kernel source |
| ATA bus, timeout/HSM and unknown device, combined | 5 | more than 6 | Fall back to PIO | Linux kernel source |
| Media errors (UNC, AMNF) | none | none | Category 0 (ATA_ECAT_NONE); none of the rules reads it | Linux kernel source |
Reading the numbers
The first thing to read in a failing SATA read is not the sense key but the ATA error register. UNC (0x40) is reported by the drive only after it has used its own retries, and libata turns it into 03h MEDIUM ERROR with 11h/04h; ICRC with ABRT (0x84) is a transfer that arrived damaged and becomes 0Bh ABORTED COMMAND with 47h/00h. The kernel wiki reads these the same way: UNC as 'often due to bad sectors on the disk', ICRC as 'often either a bad cable or power problem'. The two sources agree on the split but both say 'often', so it is a pointer, not a diagnosis.
Three details limit the split. First, the SCSI side loses information: libata maps every interface CRC to 47h/00h SCSI PARITY ERROR, although T10 lists data-phase CRC (47h/01h), iuCRC (47h/03h) and UDMA CRC (08h/03h) as separate codes, so a tool that only reads sense data cannot tell which link layer was involved. Second, when a command ends with an ATA bus error, libata-eh.c clears the device, media and invalid-argument bits of that command; a media problem on the same command is therefore not logged beside the bus error. Third, IDNF is shown as 05h ILLEGAL REQUEST with 21h/00h (LBA out of range), a name that points to addressing, not to a damaged sector.
A timeout (Emask 0x4) and an HSM violation (0x2) are the least informative lines. The kernel wiki says a timeout may have 'any number of causes', and the same lines drive the speed-down counters as bus errors do. The T10 code 5Dh/00h, FAILURE PREDICTION THRESHOLD EXCEEDED, is a drive-side flag, not an observed fault, and it appears in neither the libata translation table nor the kernel's sense-key handling read for this note.
The kernel reacts differently to the two classes. After more than 3 bus or timeout errors in 10 minutes it lowers the link speed, and after more than 6 combined errors in 5 minutes it falls back to PIO. Media errors get no such treatment: the scsi_error.c handler returns a final disposition for a MEDIUM ERROR with ASC 11h, 13h or 14h and marks it as a medium error, and tries again only for other ASC values. The practical reading is that a log full of bus errors with a link that drops speed is a different situation from a log with scattered UNC lines, and the code treats them as such; whether a given drive's cause is a cable, a backplane, power or the drive's own interface is not something these fields settle.
Related: the SMART failure-signal note,the NVMe health-log noteand what fails most often on HDDs, SSDs and links.
Method
Every code and number was copied on 2026-10-10 from a source that was opened that day: the Linux kernel files drivers/ata/libata-scsi.c (ATA-to-SCSI sense translation table), drivers/ata/libata-eh.c (error strings, bus-error handling, speed-down rules), drivers/scsi/scsi_error.c (per-sense-key disposition), include/scsi/scsi_proto.h (sense key values) and include/linux/ata.h (error register bits), read from the master branch (HEAD 3857c2fe5449541e24afc5efdb0f81a8a8f9a3a0 at access time); the T10 numeric ASC/ASCQ list (dated by T10 as of 2024-11-28); and the archived ata.wiki.kernel.org page on libata error messages. The kernel tree and the T10 list are independent publications: one says what Linux does, the other what the standard assigns. Tables give each code in the form the source prints it (hex) and name the source in the last column. Where Linux and T10 disagree, the table shows both.
Limits
Linux behaviour is read from source on the master branch on one date; older kernels, other operating systems, hardware RAID firmware and drives that report native SCSI sense (SAS) are not covered, and no log from a real failing drive was analysed, so no failure rate or detection rate is stated. The T10 list gives code names and the device types they apply to; it does not say which physical fault produces which code, so the class column in Table 2 is a reading of the code name, not a T10 statement. The archived kernel wiki says its content is obsolete; it was used only for the Emask and error-bit meanings, which the current libata-eh.c strings match. The comment above the speed-down rules in libata-eh.c says more than 8 combined bus, timeout or unknown-device errors in 5 minutes trigger a PIO fallback, while the code tests more than 6; the table follows the code. Drive-internal causes behind a timeout or a device fault are not visible in any of these fields. Nothing here covers NVMe (see the NVMe health-log note), tape drive logs, or any prediction of failure.
Sources
- 01Linux kernel: drivers/ata/libata-scsi.c, drivers/ata/libata-eh.c, drivers/scsi/scsi_error.c, include/scsi/scsi_proto.h, include/linux/ata.h (master, HEAD 3857c2fe5449) · accessed 2026-10-10
- 02T10: SCSI ASC/ASCQ assignments, numeric sorted listing (as of 2024-11-28) · accessed 2026-10-10
- 03Libata error messages, ata Wiki (archived, marked obsolete) · accessed 2026-10-10