Skip to content

Research notes / RAID performance

RAID write penalty is not a speed ratio

11 October 2026 / Source review and reproducible accounting

Four or six member operations describe a particular small-write path. They do not establish one-quarter or one-sixth of full-stripe write speed. Count reads and writes first; measure time separately.

The missing denominator

Chen and colleagues' survey compares throughput per dollar against RAID 0, under an explicit array-cost model. Its Table 3 is not a small-write versus full-stripe speed ratio. It also allows a different small-write path for narrow arrays. The surrounding discussion distinguishes throughput from response time. [1, sections 3.3 and 4.1]

We corrected the Hesela read-modify-write record: its previous wording lost that comparison basis. The arithmetic below makes the denominator inspectable without pretending to predict a controller's benchmark.

Two ways to update parity

Read-modify-write (RMW) reads the overwritten data and old parity, then writes the replacement data and updated parity. It does not normally rewrite every unchanged data block. For XOR parity, one update isP_new = P_old XOR D_old XOR D_new. [1, section 3.2.5]

Reconstruct-write combines the incoming data with the unchanged data from the same stripe, recomputes parity and writes only changed data plus parity. Williams describes this path in the 2006 Linux MD design. [2, section 4]

The name does not mean rebuilding a failed drive. Both paths here operate on a healthy array. Missing members require separate cases; this model does not choose a degraded-write strategy.

An explicit cold-cache model

Let k be the number of data blocks in one parity row,p the parity blocks (one for RAID5, two for P+Q), andu the complete data blocks replaced together, from one to k. All blocks have the same size. No old block is cached, and each member-block transfer counts as one operation, regardless of command merging.

  • RMW: u + p reads and u + p writes.
  • Reconstruct-write: k - u reads and u + p writes.
  • Equal operation count when 2u + p = k; RMW uses fewer below that boundary.

This is Hesela's accounting derivation, not a latency model or a RAID implementation. P+Q updates need their coding arithmetic; two parity blocks are not two identical XOR copies. The classic one-block RMW cases total four and six member operations. [1, sections 3.2.5-3.2.7]

Computed example: four data members plus one parity member; 4 KiB per block. Counts describe one row update, not seconds or measured IOPS.
Changed blocksRMW reads / writesReconstruct reads / writesFewer operations
1 of 42 / 23 / 2RMW
2 of 43 / 32 / 3Reconstruct
3 of 44 / 41 / 4Reconstruct
4 of 45 / 50 / 5Reconstruct

At two changed blocks, reconstruction already needs fewer transfers: five versus six. At all four, it needs zero old-data reads and five writes. The full-row RMW entry is deliberately forced for comparison; it is not a claim that a controller would choose it.

All 56 cases, CSV, JSON and model code cover two, four and eight data members, one or two parity members, and both strategies. The table above is generated from the same model at build time.

Hold the payload constant

Consider four different 4 KiB blocks in the same four-data-block row. Four isolated RMW updates, with no cache reuse, read 32 KiB and write 32 KiB to deliver 16 KiB of host data. Supplying all four together reads zero old bytes and writes 20 KiB, including parity.

That is sixteen versus five member-block transfers for the same payload. The write-byte ratios are 2 and 1.25. Counting reads too gives total-transfer ratios of 4 and1.25. Calling all of these numbers "write amplification" without a definition hides what was measured.

For this example the transfer-count reduction is 68.75%, but it isnot a measured speedup. Independent member requests can overlap. Queueing, seeks, transfer size, parity computation and the completion boundary influence elapsed time. These counts also stop at the member-device interface: they do not count an SSD's internal NAND programs.

Full coverage must reach the parity layer

A host write whose length equals a stripe's payload is not enough by itself: its address must cover the relevant stripe boundaries, and splitting or dispatch timing can change what the parity layer sees. Linux MD's journal can collect updates, but space pressure can also force partially populated stripes out. [3]

The model's 4 KiB blocks are equal slices at the same offset across members, not a claim that the configured RAID chunk size is 4 KiB. Covering every data slice removes parity prereads for that row; covering the entire configured stripe applies the same reasoning to all its rows.

Sector alignment is another layer. A partial physical-sector update on a 512e drive may need its own read/merge/rewrite, even when the RAID accounting is otherwise understood. Partition alignment does not turn every application update into a full-stripe write. [4]

Aggregation also moves the durability question. MD write-back completion can precede destaging to the array members; its journal then holds acknowledged data. Journal-device failure is therefore a separate risk. Read avoidance is not a substitute for write-hole protection. [3]

What this establishes, and what it does not

The 1994 survey is an analytical review, not a present-day SSD fleet trial. Williams' 2006 implementation paper discusses an early offload attempt that did not yield significant improvement and explains serialization constraints; it does not benchmark the current kernel. Linux's rolling documentation describes behavior, not a universal throughput guarantee. [2, sections 4-5]

Our deterministic examples have no sampled population, observation period or confidence interval. They exclude metadata, journals, flush commands, cache hits, merging, retries, failures and lower-layer write amplification. The smaller count is the model's choice, not necessarily a real controller's faster choice.

For an actual comparison, record stripe geometry and address alignment, read/write bytes at host and members, cache/journal mode, queue depth, achieved bandwidth and latency, and when a completion becomes durable. Keep warm-cache and cold-cache results separate.

Corpus: reconstruct-write, full-stripe write and write amplification.

Primary sources

  1. Peter M. Chen, Edward K. Lee, Garth A. Gibson, Randy H. Katz and David A. Patterson. RAID: High-Performance, Reliable Secondary Storage. ACM Computing Surveys 26(2), 145-185 (1994). Author manuscript: sections 3.2.5-3.3.2, Table 3 and 4.1.
  2. Dan J. Williams. MD RAID Acceleration: Support for Asynchronous DMA/XOR Engines. Linux Symposium 2006, Volume Two, printed pp. 409-414; sections 3-5. The PDF filename uses different pagination.
  3. Linux kernel documentation: RAID 4/5/6 cache, write-back mode and implementation. Rolling documentation; accessed 11 October 2026.
  4. Seagate: Transition to Advanced Format 4K Sector Hard Drives, 512-byte emulation and alignment.

Sources reviewed 11 October 2026. Original accounting code and generated data are CC0; cited publications retain their rights.