Skip to content

Research notes / Shared storage

Storage fencing: why a lost heartbeat does not stop writes

11 October 2026 / Primary-source review

A node that stops answering heartbeats may still reach shared disks. Even a node that has stopped can leave I/O in flight. Safe failover needs an enforced boundary on old access, not just a decision that another node should take over.

The request can outlive the lock holder

Minuet, a FAST 2009 paper, examines a mismatch: a distributed lock manager can transfer ownership correctly while the storage target still accepts a delayed request from the previous owner. Its proposed guard checks session information at the target. [1, sections 2.2 and 3.2]

Hesela illustration, not a measured trace. A and B are hosts; X is one shared record.
OrderEventWhat it does not establish
1A submits an update to X; delivery is delayed.The request has not necessarily disappeared if A fails.
2The coordinator grants B the lock formerly held by A.The storage target has not necessarily learned this decision.
3B writes a newer value to X.A's outstanding request has not necessarily been cancelled.
4A's old request arrives at the target.Without an applicable guard, it can overwrite B's value.

This is an ordering problem, not evidence of defective media. A checksum can validate the bytes of the wrong update. The missing question is whether that writer was still entitled to change X. See also integrity, identity and lost writes.

Path choice, cluster authority and enforcement differ

Routing: SCSI uses ALUA; NVMe uses ANA. Linux native NVMe multipath prefers an optimized path over a non-optimized one before applying its path-selection policy. Neither the term "optimized" nor redundant connectivity means exclusive application ownership. [6]

Coordination: cluster membership and quorum guide recovery decisions. Fencing: the selected mechanism prevents an excluded node from affecting the protected resource. Pacemaker treats quorum and fencing as separate concerns; its documented fencing rules also have explicit exceptions. "A majority exists" is not evidence that a particular storage path has been blocked. [5]

Registration is not exclusive write ownership

SCSI persistent reservations separate registering a key from establishing an access policy. A list of registered keys alone does not tell an operator which hosts may write. Inspect the reservation type as well. [7]

Selected Linux PR access rules for ordinary reads and writes, not every management command. [3]
Reservation typeReadersWriters
Write ExclusiveAll initiatorsHolder only
Exclusive AccessHolder onlyHolder only
Write Exclusive, Registrants OnlyAll initiatorsRegistered initiators
Exclusive Access, Registrants OnlyRegistered initiatorsRegistered initiators

ClusterLabs' fence_mpath uses the registrants-only write policy. Its pinned implementation removes the failed node's key using preempt-and-abort, then verifies that the key is absent and a reservation remains on every listed device. A remaining registrant is still allowed to write: cluster software must coordinate those writers. [4]

Linux distinguishes preempt from preempt-and-abort: the latter also aborts outstanding commands associated with the displaced connection. That distinction belongs in a recovery review, not just a checklist that says "reservations enabled." [3]

The reservation scope is the LUN, not an individual partition. Cover every shared device that the recovery procedure depends on; fencing one LUN does not establish exclusion from another. [8]

"Persistent" needs a power-loss qualification

The SCSI interface has an explicit APTPL flag for persistence through power loss. Do not infer the active configuration from the feature's name. Linux's simplified PR interface expects power-loss survival and multipath-wide coverage, although those behaviors are optional in SPC. Raw command tools and other stacks need their own capability and configuration checks. [7] [3]

A token is useful only when the recipient checks it

Burrows' OSDI 2006 Chubby paper describes a sequencer carrying the lock name, mode and generation. A client passes it with protected operations; the receiving service validates it and rejects an invalid request. The paper also describes lock-delay as an imperfect fallback for services without sequencer support, not an equivalent correctness guarantee. [2, section 2.4]

Hesela's design implication: put the stale-request check where the protected state changes, and include that check's own recovery state in the design. Merely attaching a number to a request does not enforce anything. Application fencing tokens and SCSI reservation keys are different mechanisms; do not treat an arbitrary registration key as a monotonic lock generation.

What the papers establish, and what they do not

Minuet evaluated a modified iSCSI prototype, not off-the-shelf arrays: 39 Emulab nodes, 100 Mbps links, five-minute runs averaged over three iterations, using chunkmap and B-tree workloads. The authors identify target changes and application rejection handling as deployment requirements. This supports a design argument, not a present-day throughput promise. [1, sections 4-6]

Chubby is an operational design report about coarse-grained coordination, not a storage-device fencing benchmark. We reviewed its rationale and sequencer semantics; we did not reproduce its service or Minuet. No new failure-rate estimate, timing bound or hardware certification follows from this review. [2]

Questions to answer before enabling recovery

This is a design-review checklist, not a command sequence for live disks:

  • Which exact resource and writer are excluded, and at which enforcement point?
  • What happens to already-issued I/O, every alternate path and every dependent LUN?
  • Which reservation type or request-generation check actually constrains access?
  • What verifies completion, and what does recovery do if exclusion cannot be confirmed?
  • How are power loss, controller failover and node rejoining handled without restoring stale access?

Use the supported cluster and storage procedure and validate disruptive cases only on disposable test resources. Clearing reservations during diagnosis can remove the protection being investigated.

Corpus: I/O fencing, SCSI persistent reservations, fencing token, ALUA, ANA and multipath I/O.

Primary sources

  1. Ermolinskiy, Moon, Chun and Shenker. Minuet: Rethinking Concurrency Control in Storage Area Networks. FAST 2009, sections 2.2, 3.2 and 4-6.
  2. Mike Burrows. The Chubby lock service for loosely-coupled distributed systems. OSDI 2006, sections 2.1 and 2.4.
  3. Linux kernel documentation: Block layer support for Persistent Reservations.
  4. ClusterLabs fence_mpath source and embedded documentation, commit 65a423c846a91deff49271ffaeb8f7d4ed90ce50.
  5. ClusterLabs, Pacemaker Explained 3.0: Fencing, including Fencing and Quorum.
  6. Linux kernel documentation: Linux NVMe multipath, introduction and policies.
  7. Douglas Gilbert, sg3_utils: sg_persist(8), registration, reservation and APTPL semantics.
  8. Red Hat Enterprise Linux 6, Fence Configuration Guide 4.28: SCSI Persistent Reservations (historical guide, LUN scope).

Reviewed 11 October 2026. Historical papers and documentation are used within the scopes stated above. No unverified DOI is asserted.