Skip to content

Paper note

An example BCH code corrects raw bit error rates up to 1e-3, yet Google's field SSDs had median rates of 6e-10 to 3e-8, and raw rate did not predict uncorrectable errors

Published 2026-10-11

An example BCH code in the survey of Cai et al. (Proceedings of the IEEE, 2017) stays within the 1e-16 uncorrectable error target up to a raw bit error rate of 1.0e-3, and an example soft LDPC code of the same rate reaches 5.0e-3, at a cost of 0.28% more reads in the survey's worst-case example; Zhao et al. (FAST 2013) find about 3 times for six extra sensing levels, and that soft decoding can add over 100% to worst-case read response time at 10,000 program/erase cycles, cut to below 20% by their techniques. In Google's field data (Schroeder et al., FAST 2016), median raw bit error rates of first-generation drive models were 5.8e-10 to over 3e-8 and the 99th percentile reached 2.7e-5, so typical drives sat many orders of magnitude below the example BCH limit, yet 2 to 6 per 1,000 drive days had an uncorrectable error and raw rate did not predict them.

Design sources
Zhao et al., 25nm MLC chips and a simulated SSD; Cai et al., a survey with example BCH and LDPC codes of similar rate
Field source
Google data centers, several first-generation and later drive models, MLC and SLC
Error measures
Raw bit error rate before ECC; uncorrectable bit error rate after ECC; uncorrectable errors per drive day
Targets
1e-16 uncorrectable bit error rate (Cai et al.); 1e-15 decoding failure probability (Zhao et al.)

Numbers

NAND error correction figures as printed in Cai et al. (rows 1 to 4), Zhao et al. (rows 5 to 9), and Schroeder et al. (rows 10 to 13). Values are as printed; qualifiers such as 'over' and 'below' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
Target uncorrectable bit error rate1e-16errors per bit readContemporary SSD makers, as cited by Cai et al.
Highest raw bit error rate an example BCH code corrects at that target1.0e-3raw bit errors per bitCode rate 0.935, one hard read
Highest raw bit error rate an example soft LDPC code corrects at that target5.0e-3raw bit errors per bitCode rate 0.936, soft decoding with extra reads
Extra reads of the example combined hard and soft LDPC scheme0.28percent more reads than BCHIf 0.01% of codewords fail and soft LDPC needs seven extra reads
Raw bit error rate tolerance of soft LDPC relative to BCHabout 3times25nm MLC chip model, six extra sensing levels, 1e-15 target
Redundancy of the LDPC code in the simulation512bytes per 4 KB of user dataRate 8/9 code
Worst-case read response time increase with soft decoding at 10,000 program/erase cyclesover 100percent (lower bound)Read-intensive traces, simulated SSD
Worst-case read response time increase with the authors' three techniquesbelow 20percent (upper bound)Same traces
Hard-decision LDPC decoding failure probability across the life of the flash1e-15 to 1e-2probability per readEarly life to late life
Median raw bit error rate of first-generation drive models5.8e-10 to over 3e-8raw bit errors per bit readGoogle data centers, several drive models
99th percentile raw bit error rate across models2.2e-8 to 2.7e-5raw bit errors per bit readSame drives; SLC-B lowest, MLC-D highest
Drive days with an uncorrectable error2 to 6per 1,000 drive daysSame drives
Drives with non-final read errors that retries can recoverunder 2percent of drives (upper bound)Same drives

This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 4 are from Cai et al. (Proceedings of the IEEE 2017), rows 5 to 9 from Zhao et al. (FAST 2013), and rows 10 to 13 from Schroeder et al. (FAST 2016). The sources measure different things.

Method

The figures are copied from the PDFs of Zhao, Zhao, Sun, Zhang, Zhang, and Zheng, FAST 2013, from the arXiv version of Cai, Ghose, Haratsch, Luo, and Mutlu, Proceedings of the IEEE 2017, and from Schroeder, Lagisetty, and Merchant, FAST 2016, not refit. Zhao et al. measure 25nm MLC NAND chips and drive a trace-based simulator of an SSD with a rate-8/9 LDPC code. Cai et al. is a survey whose BCH and LDPC strength curves are an example for similar codeword lengths and code rates. Schroeder et al. report raw bit error rates and uncorrectable errors from drives in Google data centers. The comparison of code limits with field error rates is this page's own and is made only to show orders of magnitude.

Limits

The code limits are examples: one BCH code at rate 0.935 and one LDPC code at rate 0.936 in a survey, and a simulated rate-8/9 code in Zhao et al., not the codes inside any specific drive, and Schroeder et al. do not say which ECC their drives use beyond that all models of a generation use the same one. Raw bit error rates in the field data are medians and percentiles of drive months, so the worst pages inside a drive, after wear and retention, can be far above the median and are what ECC must handle. Cai et al. cite Zhao et al. among their sources, so the two design sources are not fully independent; Schroeder et al. is. The 3 times and 5 times tolerance gains use different codes and sensing levels and are not comparable. Zhao et al.'s latency results come from simulation of 25nm MLC and workloads from 2013 and may not match current flash. Uncorrectable errors in the field are counted per drive day, not per bit, so they are not an uncorrectable bit error rate.

What the paper found

The answer. ECC strength is a trade between the raw bit errors a code can correct and the extra reads it needs. In the survey of Cai et al., an example BCH code of rate 0.935 stays inside the uncorrectable error target of 1e-16 up to a raw bit error rate of 1.0e-3, while hard-decision LDPC at a similar rate is about the same, and soft LDPC decoding reaches 5.0e-3, up to five times more. The cost is latency: soft decoding takes several extra reads, and in the survey's example, 0.01% failed codewords and seven extra reads mean up to 0.28% more reads for twice BCH's correction strength.

The latency bill. Zhao et al. (FAST 2013) model a rate-8/9 LDPC code with 512 bytes of redundancy per 4 KB, find soft decoding with six extra sensing levels tolerates about 3 times the raw bit error rate of BCH at a 1e-15 target, and show that once 25nm MLC chips reach 10,000 program/erase cycles, soft-decision decoding is invoked often enough to add over 100% to the worst-case read response time for read-intensive workloads. Their three techniques, including look-ahead sensing and data placement interleaving, cut that to below 20%, at the cost of about 102.9% more flash access power for look-ahead sensing.

The field gap. In Google's data centers, median raw bit error rates of first-generation models ranged from 5.8e-10 to over 3e-8, and 99th percentile rates from 2.2e-8 to 2.7e-5, with MLC drives orders of magnitude above SLC. Those values sit below the example BCH limit of 1.0e-3, which is why most reads never need strong decoding. Even so, uncorrectable errors hit 2 to 6 of every 1,000 drive days, and the authors found no correlation between a model's or a drive's raw bit error rate and its share of drive days with uncorrectable errors.

What to take from it. A code limit and a field median answer different questions. The limit says how much raw damage a page can absorb before data is lost, the median says how damaged a typical page is, and the uncorrectable errors that matter come from the tail and from other failure modes. The three sources agree that raw bit error rate alone does not describe risk, but they measure different things on different flash and codes, so the sizes are not comparable.

Background on drive types is on the drive technology card. Warning signs before a failure are also covered in the SMART failure signals analysis.

Sources

  1. 01LDPC-in-SSD: Making Advanced Error Correction Codes Work Effectively in Solid State Drives, Zhao, Zhao, Sun, Zhang, Zhang, and Zheng, USENIX FAST 2013, pages 243 to 256 (USENIX PDF) · accessed 2026-10-11
  2. 02Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives, Cai, Ghose, Haratsch, Luo, and Mutlu, Proceedings of the IEEE, September 2017 (arXiv 1706.08642, section 6) · accessed 2026-10-11
  3. 03Flash Reliability in Production: The Expected and the Unexpected, Schroeder, Lagisetty, and Merchant, USENIX FAST 2016 (USENIX PDF, sections 3 and 4) · accessed 2026-10-11