Research notes / NAND error correction
LDPC in NAND: the read cost of confidence
Published 11 October 2026 · Literature analysis, not an SSD benchmark
An error-correction engine needs evidence, not just compute. In NAND, obtaining more informative observations can require extra sensing. The useful question is not simply whether an SSD has LDPC, but what its complete recovery path costs.
A bit decision is not a confidence measurement
Wang and colleagues' 2011 study uses repeated comparisons at different reference voltages to identify a finer interval for a cell's threshold voltage. This gives a decoder information beyond a single hard decision. Their model uses Gaussian noise; its results are simulations, not measurements of a production SSD. [1, sections II-IV]
LDPC describes a family of codes defined by sparse parity-check matrices, not a fixed number of retries or a universal correction guarantee. The sensing interface, selected code and decoder must be specified together. The separate read-retry analysis explains why resubmitting a host request is not the same operation.
Three costs hide behind one read
Zhao and colleagues separate sensing, flash-to-controller transfer and decoding work. Their progressive-precision strategy starts with hard decisions and escalates after failure. A design favoring hard decoding need not be best for soft decoding; their proposed code-structure adaptation addresses that tension. [2]
The evaluation uses a hypothetical two-bit-per-cell device, three rate-8/9 codes protecting 2 kB each, Monte Carlo simulation and 65 nm ASIC design. Those are valuable engineering models, but neither shipping-firmware evidence nor measured latency for today's 3D NAND.
Our practical inference: a report giving only decoder throughput leaves acquisition cost unknown. Conversely, a count of sensing operations omits iteration limits, transfers and possible overlap. Neither quantity alone establishes end-to-end read latency.
The obscure trap: a better channel metric can lose
Wang and colleagues' 2014 manuscript compares finite LDPC codes under quantized Gaussian and retention models. For one code, maximizing channel mutual information also minimizes frame errors. For a code containing small absorbing sets, another quantization performs better even though it carries less mutual information. [3, sections IV and V.D]
An absorbing set is a decoder-graph structure that can sustain troublesome error patterns. The result is a counterexample to treating an information-theoretic optimum as a guaranteed optimum for every implemented decoder. It is not evidence that less accurate sensing generally improves SSDs.
Confidence depends on the model
Cai and colleagues describe soft information as a log-likelihood ratio (LLR) conditioned on the observed voltage interval. Its sign favors a bit value under the chosen convention; its magnitude expresses relative confidence. Their example recovery flow gathers additional observations for soft decoding after hard decoding fails. [4, section 6.2]
Our inference for reproducibility: preserve the mapping from observations to confidence values. A paper's decoder result cannot be reproduced by naming the code family while omitting that mapping.
Keep the numerator attached to the result
NIST defines BER through erroneous bits relative to transmitted bits and explicitly includes storage. BER alone does not specify a before-correction measurement point. [5] Cai's comparison distinguishes raw errors from uncorrectable outcomes; its UBER convention divides codeword failure rate by codeword length. That is not a count of every wrong bit inside a failed word. [4, section 6.3]
| Reported quantity | Record with it | Do not substitute |
|---|---|---|
| Raw bit errors | Bits examined, sensing settings and pre-ECC boundary | Host data-loss probability |
| Failed codewords | Words attempted, word length, decoder and stopping rule | Number of erroneous bits |
| Additional sensing | Observations per recovery and when escalation occurs | Decoder iterations |
| Request latency | Workload, sample count, percentiles and completion outcome | A claim about internal recovery without telemetry |
What this review establishes
The sources support distinct mechanisms and design tradeoffs. They do not establish a universal retry budget, an SSD replacement threshold or a vendor ranking. No hardware experiment, LDPC decoder reproduction or new failure-rate estimate was performed here. The checklist and cross-paper interpretation are our synthesis.
For an independent implementation study, disclose code construction, rate, block length, channel assumptions, confidence mapping, quantization, decoder schedule, stopping criteria and seeds. Report unsuccessful cases as well as successful decodes. Compare sensing and end-to-end costs under the same workload.
Structured definitions: LDPC, soft-decision decoding, log-likelihood ratio and bit error rate. For host-visible evidence, see NVMe health-log limits.
Primary sources
- Wang, Courtade, Shankar and Wesel. Soft Information for LDPC Decoding in Flash: Mutual-Information Optimized Quantization. IEEE GLOBECOM 2011, sections II-IV.
- Zhao, Dong, Sun, Zheng and Zhang. Reducing latency overhead caused by using LDPC codes in NAND flash memory. EURASIP JASP 2012:203. DOI: 10.1186/1687-6180-2012-203.
- Wang et al. Enhanced Precision Through Multiple Reads for LDPC Decoding in Flash Memories. Author manuscript arXiv:1309.0566v4, 2014, sections IV and V.D.
- Cai, Ghose, Haratsch, Luo and Mutlu. Error Characterization, Mitigation, and Recovery in Flash Memory Based Solid-State Drives. Author manuscript arXiv:1706.08642, 2017, sections 6.2-6.3.
- NIST CSRC glossary: bit error rate, citing CNSSI 4009-2022.
Reviewed 11 October 2026, Europe/Warsaw. No third-party figures or datasets reproduced.