Research notes / NAND recovery
NAND read-retry: when a successful SSD read takes longer
Published 11 October 2026 · Literature review, not a device benchmark
A slow read can still return correct data. Before calling it media failure, separate recovery inside the flash device from an error reported to the host.
Read-retry is not simply resubmitting a host request
Linux's raw-NAND interface describes retry modes that change sensing thresholds when a page has too many bit errors for correction. The page is read again after the adjustment. Its ECC callbacks distinguish corrected bitflips from errors beyond correction capability. These are separate outcomes, not interchangeable failure counters. [1]
The documented read_retries member counts supported modes, not observed retries. Nor is nand_setup_read_retry a universal tuning interface for an NVMe SSD. Keep the interface layer attached to any counter or control you report.
What the read-retry study actually measured
Park and colleagues characterized 160 48-layer 3D TLC chips, sampling 120 blocks per chip and testing 11,059,200 pages. Their FPGA platform varied wear, temperature and retention age; long retention conditions used accelerated aging. Their subsequent SSD performance evaluation used MQSim simulation with twelve workloads, informed by those chip measurements. It was not a fleet trial of retail SSDs. [2, sections 4 and 7.1]
The paper explores overlapping retry stages and shortening sensing time within an ECC margin. Its useful lesson is architectural: retry cost depends on the controller's recovery path, not just the number of retries. Reported speedups need their baseline and workload; they are not a firmware upgrade promise.
The less obvious variable: time since programming
Luo and colleagues studied single-vendor 3D MLC chips at 20 degrees Celsius. Retention observations spanned seven minutes to 24 days. They found faster initial degradation, called early retention loss, and dependence on neighboring cell state, called retention interference. Read-reference settings therefore cannot be interpreted from wear alone. [3, sections 4.1, 4.3 and 4.4]
This is historical device evidence, not a universal shelf-life prediction. Time since a physical page was programmed is also a different quantity from drive power-on age. A benchmark description that gives only the drive's age leaves the data's retention history unspecified.
Read disturb is a different mechanism
In Cai and colleagues' planar MLC experiments, reading one wordline applied pass-through voltage to unread cells in the same block. Repeated exposure could shift their thresholds. This differs from charge loss while data waits. Their voltage-tuning characterization emulated pass-voltage changes through read-reference control because the commercial interface did not directly expose that setting. [4, sections 2 and 3.1]
Do not turn that experiment into an instruction to lower arbitrary SSD voltages. The study itself describes a tradeoff: lowering pass-through voltage can reduce disturbance but introduce other read errors. Neither a blanket ban on scrubbing nor unlimited rereading follows from the mechanism.
Match the claim to the observation
| Observation | What it establishes | What remains unproven |
|---|---|---|
| Correct data returned slowly | Completion time and readback for that request | Which internal recovery path ran |
| Raw-NAND ECC reports corrected bits | Correction at the documented interface | An application-visible data-loss event |
| A vendor retry counter increases | Only the event defined by that vendor | A portable failure probability |
| A chip experiment changes error rates | A result for its tested conditions | A current fleet's replacement rate |
A latency trace alone cannot identify NAND retry. Investigate competing explanations such as queueing, thermal throttling, background maintenance and the storage path. Start with the existing thermal analysis and NVMe health-log limits; do not convert a clean health summary into proof that every internal read was cheap.
A useful investigation record
- Record the model, firmware, interface and exact meaning of each available counter.
- Keep read outcomes separate from latency percentiles; state request size, queue depth, duration and sample count.
- Record write history and temperatures when known. Mark physical retention age unknown when it cannot be observed.
- Compare like-for-like workloads before and after a supported intervention. Preserve the original evidence.
This checklist is our synthesis, not validation of a particular product. Avoid forced aging, voltage changes or destructive refresh experiments on production data. A device with read errors calls for the existing backup and recovery procedure, not an improvised benchmark.
Method and limits
We reviewed the papers' mechanisms, sampling and evaluation methods, plus the Linux API contract. No hardware experiment or simulator reproduction was performed. These laboratory samples do not estimate current failure prevalence or establish a replacement threshold.
Machine-readable definitions: read-retry, early retention loss, read disturb and data retention.
Primary sources
- Linux MTD NAND driver API: nand_setup_read_retry, ECC callbacks and read_retries
- Park, Kim, Chun, Orosa, Kim and Mutlu. Reducing Solid-State Drive Read Latency by Optimizing Read-Retry. ASPLOS 2021, sections 4 and 7.1. DOI: 10.1145/3445814.3446719
- Luo, Ghose, Cai, Haratsch and Mutlu. Improving 3D NAND Flash Memory Lifetime by Tolerating Early Retention Loss and Process Variation. POMACS 2018, sections 4.1, 4.3 and 4.4
- Cai, Luo, Ghose, Haratsch, Mai and Mutlu. Read Disturb Errors in MLC NAND Flash Memory: Characterization, Mitigation, and Recovery. DSN 2015, sections 2 and 3
Accessed 11 October 2026, Europe/Warsaw. Original figures are not reproduced.