Research notes / Recovery semantics
Storage timeouts: an expired timer is not a cancelled write
11 October 2026 / Primary-source review and executable illustration
A timeout says that one layer stopped waiting for completion. It does not, by itself, prove that a write never happened, that outstanding I/O has stopped, or that retrying the operation is safe.
First name the timer's owner
An application deadline, a command timeout and a path-recovery policy answer different questions. During an incident, record which layer produced the error before interpreting it as a disk failure. A deadline in a caller is not automatically a cancellation contract with every layer below it.
| Boundary | Question | Still needs evidence |
|---|---|---|
| Application deadline | How long does this caller wait? | The operation's final outcome and cancellation acknowledgment. |
| Command timeout | When does the driver begin handling an uncompleted command? | Abort, retry, reset and final completion status. |
| Path availability | Is another route usable? | How pending work behaves when all routes disappear. |
| Logical operation | Is this submission the same request as before? | Duplicate suppression and ordering across recovery. |
SCSI error handling is a process, not an instant
Linux's SCSI EH documentation allows a timeout handler to restart the timer or initiate abort/recovery. Timed-out commands can remain active below the midlayer until recovery deals with them. Entering host recovery also blocks new commands to that host. [1, sections 1.2-1.4]
The documented escalation includes device, bus and host resets. A timeout message therefore does not identify a physical media defect, and recovery scope may be wider than one command. Check the deployed kernel, low-level driver and transport rather than treating this architectural documentation as a universal timing guarantee. [1, section 2.1]
For evidence at adjacent layers, see SCSI sense and kernel logs and link failure signatures.
All paths down can mean waiting, not a prompt error
Upstream no_path_retry supports a positive retry count, fail for no queueing, or queue for indefinite queueing. A numeric value is not a number of seconds. The effective configuration, installed version and path checker matter. Do not turn it into an end-to-end deadline by reading the name alone. [2]
Red Hat documents that queue_if_no_path can leave processes issuing I/O blocked until a path returns. It also warns that an eh_deadline-triggered HBA reset affects every target path on that adapter. A shorter setting can therefore have a wider impact, not merely reduce one request's latency. [3, sections 9.1 and 10.1]
For Fibre Channel, fast_io_fail_tmo concerns failing I/O after a remote-port problem; dev_loss_tmo concerns removal after connection loss. Their scope differs from a command timer. These SCSI/DM-Multipath options must not be copied wholesale into native NVMe multipath configurations. [2]
An apparently harmless retry can restore an old value
Hesela's deterministic example uses two logical writers and a lost response. Repeating an assignment is harmless in isolation; inserting another writer changes that argument. The table is generated from tested code during the build, with no client-side script.
| Event | Blind replay | Tracked retry | Records lost |
|---|---|---|---|
| A writes draft; its response is lost | draft | draft | draft |
| B writes approved and receives success | approved | approved | approved |
| A retries the original logical operation | draft | approved | draft |
The tracked retry returns A's earlier result without changing B's value. It does not claim that A's value is current. Assigning the retry a new request identity also defeats this model's duplicate check. The example has no disks, crashes or persistent log: losing its map is deliberately a failure case, not a recovery algorithm.
What RIFL adds beyond a retry loop
RIFL (SOSP 2015) combines stable RPC identity, completion records made durable atomically with mutations, record migration and lease-based reclamation. Its guarantee assumes reliable client state; an upstream client retrying as a new request remains a separate problem. [4, sections 2-4 and 9]
The paper evaluates RAMCloud over InfiniBand with three-way replication and its log cleaner disabled. This is evidence about a particular RPC design, not certification that Linux block retries implement it. We reviewed the design and evaluation; we did not reproduce RIFL or reuse its performance figures. [4, section 7]
A useful incident record preserves uncertainty
Hesela's synthesis: record unknown outcome when that is all the evidence supports. Neither a healthy replacement path nor a new application leader answers whether an earlier operation took effect.
- Identify the original logical request, resource and timing layer; retain its identity across a genuine retry.
- Separate observed failure from assumed non-execution, and record any later completion or recovery evidence.
- Inventory command, transport, multipath and application policies, including the all-paths-down case.
- Test delayed responses, path loss, recovery and duplicate-state loss only on disposable resources; verify final data, not just error latency.
This is a design-review checklist, not a tuning recipe. Changing queueing or resetting an adapter can disrupt live work. Use the supported procedure for the actual storage stack.
Related boundaries: fencing and stale writers and why retrying a failed fsync does not establish durability. Corpus: I/O timeout, no_path_retry and retry deduplication.
Primary sources and limits
- Linux kernel documentation: SCSI EH, sections 1.2-1.4 and 2.1.
- multipath-tools: multipath.conf(5), no_path_retry and FC timeout options; pinned commit d53932bdbee02f440518ea384e63d7c4d3f66685.
- Red Hat Enterprise Linux 9: Configuring device mapper multipath, sections 9.1 and 10.1.
- Collin Lee, Seo Jin Park, Ankita Kejriwal, Satoshi Matsushita and John Ousterhout. Implementing Linearizability at Large Scale and Low Latency. SOSP 2015; DOI 10.1145/2815400.2815416; sections 2-4, 7 and 9.
Reviewed 11 October 2026. No hardware fault injection, failure-rate estimate or universal timeout value is claimed. Protocol- and version-specific behavior still needs validation.