Research notes / Erasure coding
Local repair is not whole-node repair
11 October 2026 / Source review and reproducible accounting
A code can read six fragments to recover one missing data fragment and still have a whole-node repair cost above six. Parity must also be repaired, and the denominator matters. Locality is a repair property, not a complete durability guarantee.
Local parity changes who must be read
Huang and colleagues' 2012 construction divides data into local groups and adds global parity. In its 12+2+2 configuration, six data fragments share each local parity. One missing data fragment can be rebuilt from five peers and that parity. The corresponding whole-fragment RS 12+4 repair reads twelve helpers. Both store sixteen fragments for twelve data fragments. [1, sections 2-3]
Here, 12+2+2 means data + local parity + global parity. It is not a Ceph profile string. Ceph uses l for its locality parameter and can group coding chunks as well as data chunks. Translate the construction before translating the parameter names. [3]
Three answers to three different questions
Hesela's deterministic model counts complete helper-fragment reads. Each repair is independent, with one missing fragment and all its helpers intact. A global-parity repair reads the twelve data fragments; a local-parity repair reads its six data members. No encoder or storage system is run.
| Construction | One lost data fragment: helpers | One lost data fragment: MiB read | Average per lost fragment | Per lost data-equivalent |
|---|---|---|---|---|
| RS 12+4 | 12 | 768 | 12 | 16 |
| LRC 12+2+2 | 6 | 384 | 6.75 | 9 |
The LRC accounting sum is 12 x 6 + 2 x 6 + 2 x 12 = 108. Dividing by sixteen stored fragments gives 6.75; dividing by twelve payload fragments gives 9. The RS sum is 16 x 12 = 192, giving 12 and 16, respectively.
These are independent repair costs summed across fragment roles, not sixteen simultaneous erasures of one stripe. For a node holding many stripes, the normalized comparison assumes uniform placement of data, local parity and global parity. A node with a different mix needs a weighted calculation.
For this model, data-only reads fall by 50%; the data-normalized total falls by 43.75%. Those are arithmetic consequences of the stated inputs, not measured speedups. The three columns follow the degraded-read, average-repair and normalized-repair distinctions in Kolosov and colleagues' methodology. [2, section 3]
All five configurations, CSV, JSON and reproduction code. The table above is generated from the same tested model during the site build.
Four parity fragments do not promise any four losses
The paper's LRC construction tolerates any three erased fragments, but not every set of four. Its maximally recoverable property is relative to its local/global topology; it is not the MDS guarantee of the RS baseline. [1, section 2.2]
A concrete counterexample: lose three data fragments from one group and that group's local parity. Only the two global equations constrain the three missing data values. The other group's local parity adds no information about those values. Counting stored parity without asking which unknowns it covers misses the problem.
This is an erasure argument: the unavailable locations are known. Silent corruption, faulty metadata, correlated placement and detection delays need separate treatment. Neither a parity count nor this calculator estimates annual data-loss risk.
What the second paper actually tested
Kolosov and colleagues compared LRC approaches in Ceph on EC2. Their experiment removed one OSD and measured recovery, with additional storage, network and foreground-load cases. Less repair data did not translate proportionally into less elapsed time. [2, sections 5-6]
An important qualification is in their implementation notes: Ceph's LRC plugin used Pyramid codes, not the exact Azure construction. The authors matched single-node read behavior, changed recovery reads and disabled rebalancing for the experiment. That does not establish identical multi-erasure tolerance or stock-Ceph performance. [2, section 5 and footnote 1]
A comparison worth keeping
Hesela's synthesis: a useful repair report should make these choices explicit.
- State the construction, fragment placement and erasure pattern, not just a parity count.
- Separate data-only recovery from all-fragment and full-node recovery.
- Report helper reads, bytes read, bytes transferred and reconstructed writes separately.
- State whether comparisons hold payload, storage capacity or fault tolerance constant.
- Measure completion time under the actual bottleneck and foreground workload.
Our equal-overhead example does not hold worst-case erasure tolerance equal. It excludes partial-fragment repair, read sharing, compression, retries, cache effects, CPU cost, placement and scheduling. The 64 MiB size is an illustrative input. No sampling or confidence interval applies to deterministic arithmetic.
Corpus: locally repairable code, repair locality, repair-read amplification and erasure coding.
Primary sources and scope
- Cheng Huang, Huseyin Simitci, Yikang Xu, Aaron Ogus, Brad Calder, Parikshit Gopalan, Jin Li and Sergey Yekhanin. Erasure Coding in Windows Azure Storage. USENIX ATC 2012, pp. 15-26; sections 2-3.
- Oleg Kolosov, Gala Yadgar, Matan Liram, Itzhak Tamo and Alexander Barg. On Fault Tolerance, Locality, and Optimality in Locally Repairable Codes. USENIX ATC 2018, pp. 865-877; sections 3, 5-6.
- Ceph: Locally Repairable Erasure Code Plugin (development documentation; parameter names are implementation-specific).
Sources reviewed 11 October 2026. This is a literature synthesis and an accounting model, not a reproduction of either cluster experiment or a claim about Azure's current internals.