Skip to content

Research notes / Crash recovery

RAID write hole: why a bitmap is not a write journal

Published 11 October 2026 / Primary-source review and algebraic model

A RAID bitmap remembers where writes may be incomplete. It does not store the bytes needed to replay them. Partial Parity Log and a data journal address different recovery obligations.

The key distinction is between recovering untouched data, preserving an in-flight update, and committing an application transaction. Calling all three "power-loss protection" hides the question an engineer needs answered.

One missing drive is not the only condition

In a simple RAID5 stripe, parity obeys P = A XOR B XOR C. If B disappears, its value follows from B = P XOR A XOR C. That reconstruction assumes P matches the surviving data.

Changing A also changes P. An interruption between those member updates can leave a mismatched pair. If B is then unavailable before consistency is restored, the same XOR operation can silently produce the wrong B, even though B was never part of the interrupted write.

Linux MD documents this dirty-and-degraded hazard and normally refuses to assemble such an array. Here, "dirty" means redundancy may be inconsistent; "degraded" means a member is missing. An assembly override does not create missing evidence. [3]

A four-state counterexample

Start with one bit per chunk: A=0, B=1, C=1, P=0. Change only A to 1, so the intended new parity is 1. For this illustration, each member update is atomic, but data and parity may persist independently. B becomes unavailable at reconstruction.

BeforeA 0 / B 1 / C 1P 0: consistent
Only A persistsA 1 / B 1 / C 1P 0: stale
B missing0 XOR 1 XOR 1 = 0Wrong B: it should be 1
Hesela's bit-level example. These are logical states, not measured disks, timing samples or a physical drive layout.
Computed from the downloadable model for the initial state above. PPL columns assume a valid partial-parity record survived.
Persisted updatesA / P after crashNaive recovered BPPL recovered BNew A present?
Neither0 / 01 (correct)1 (correct)No
Only P0 / 10 (wrong)1 (correct)No
Only A1 / 00 (wrong)1 (correct)Yes
A and P1 / 11 (correct)1 (correct)Yes

Both mixed states reconstruct B incorrectly with stale parity. A normal resynchronization can recompute P if all data chunks remain readable. Merely knowing that the stripe is dirty cannot supply B once it is missing.

What PPL saves, and what it does not

Linux MD's PPL records the XOR of unmodified chunks before dispatching member updates. It lives on the stripe's parity member, without a separate journal drive. The documented implementation is for RAID5 and does not combine PPL with a write-intent bitmap. [1]

In our example, that record is Q = B XOR C = 0. Use the A that actually survived: P_repaired = A_after XOR Q. Recovering B then cancels A: B = P_repaired XOR A_after XOR C = Q XOR C = 1. This works whether A is old or new.

That last sentence is the limit as well as the benefit. The algebra protects B; it says nothing about whether the intended A arrived. The kernel documentation explicitly excludes a guarantee for in-flight data and describes the case of losing a modified member as unprotected by PPL recovery. [1]

Four mechanisms, four different promises

Linux MD documentation, not a ranking of every RAID product.
MechanismRecovery informationImportant boundary
Write-intent bitmapRegion markers for synchronizationNo replay payload [3]
PPLPartial parity of untouched chunksNot an in-flight data journal [1]
Write-through journalData and parity logged before member updatesCompletion waits for array members [2]
Write-back journalData logged before later destagingCompleted I/O can depend on the journal device [2]

In Linux MD write-through mode, losing the journal does not discard already completed writes, but removes write-hole protection for subsequent operation. In write-back mode, loss of the journal can lose completed data not yet destaged. The extra device is therefore part of the durability design, not just a performance accessory. [2]

The log must survive, and the application still has a protocol

A recovery proof that begins "the log is durable" has a real precondition. Linux distinguishes flushing previous cached writes from FUA, which ties completion of the flagged write to nonvolatile storage. Remapping layers must preserve these semantics. A normal device completion is not universally that durability boundary. [4]

Zheng and colleagues injected power faults into SSDs and checked recovered contents, including ordering and partial-record failures. Their historical device study tests a lower layer than our XOR model. It is evidence for validating storage assumptions, not a measured failure rate for PPL or modern journal devices. [6, sections 3-4]

At the other end of the stack, Pillai and colleagues used block reordering and application-specific recovery checks to study persistence properties. The OSDI14 work separates ordering, atomicity and durability; its application analysis also considers expectations beyond some documented guarantees. We use that distinction, not its vulnerability totals as a ranking of current software. [5, sections 2-4]

Hesela's engineering conclusion: a parity-consistent array can still hold an incomplete multi-block application update. A filesystem or database must enforce its own recovery protocol. Closing the RAID write hole alone does not turn every file operation into an atomic transaction.

Reproduce the reasoning

Primary-source review of version-pinned Linux v6.12 documentation and FAST13/OSDI14 research, plus an original exhaustive one-bit XOR illustration. No hardware fault injection or Linux MD implementation test was performed.

The model enumerates 32 cases: eight initial bit triples and four persistence masks, with A always changed. Naive B reconstruction is wrong in 16 cases; the assumed intact partial-parity record recovers B in all 32. The intended A remains absent in 16.

None of these counts is a probability. We assign no likelihood to a crash time or write order. There is no measured population and no statistical confidence interval to estimate.

All states, CSV, JSON provenance and runnable model. The page and data use the same model bytes. This is an algebraic illustration of a guarantee boundary, not a reproduction of either paper or a verification of Linux MD.

The model assumes atomic member writes, an intact durable partial-parity record and a missing untouched data member. State counts are not probabilities. Historical papers do not establish current product failure rates. Torn writes, corrupted logs, a missing modified member, concurrent updates and application acknowledgements are outside the model.

Questions for an actual array

  • Which consistency policy is active on this exact array, kernel and metadata format?
  • What does write completion promise, and which device still holds the only current copy?
  • Does recovery preserve untouched data, replay new data, or both under the specified failure?
  • Have the cache, flush, firmware and application recovery assumptions been validated together?

These are review questions, not instructions to force-assemble a damaged array or interrupt production power. Qualification belongs on isolated, disposable test media with expected contents and recovery rules defined first.

Related: SSD power-fault evidence, fsync failures and NAND paired pages.

Definitions: RAID write hole, Partial Parity Log, write-intent bitmap and RAID 5.

Primary sources

  1. Linux v6.12: Partial Parity Log
  2. Linux v6.12: RAID 4/5/6 cache
  3. Linux v6.12: RAID arrays, dirty/degraded assembly and bitmap metadata
  4. Linux v6.12: Explicit volatile write back cache control
  5. Pillai, Chidambaram, Alagappan, Al-Kiswany, Arpaci-Dusseau and Arpaci-Dusseau. All File Systems Are Not Created Equal: On the Complexity of Crafting Crash-Consistent Applications. OSDI 2014, pp. 433-448, sections 2-4.
  6. Zheng, Tucek, Qin and Lillibridge. Understanding the Robustness of SSDs under Power Fault. FAST 2013, pp. 271-284, sections 3-4.

Accessed 11 October 2026. Linux claims are pinned to v6.12 documentation, not a claim about every version or vendor. Paper figures are not reproduced.