Skip to content

Paper note

NAND erase blocks grew from 64 pages to as many as 576, and in one simulation write amplification rose from 2.1 to 5.1 with block size

Published 2026-10-11

Flash is read and programmed in pages but erased only in whole blocks. In the 2008 example part of Agrawal et al. (ATC 2008) a 256 KB block held 64 pages of 4 KB, and erasing it took 1.5 ms against 200 microseconds to program a page, with 100,000 erase cycles per block. Liu et al. (FAST 2018) report 128 to 192 pages per block in 2D NAND and up to 576 in 3D NAND, and in their block-level FTL simulation write amplification rose from 2.1 with 72-page blocks to 5.1 with 576-page blocks, because every garbage collection has to copy more valid pages before an erase. Cai et al. (Proceedings of the IEEE, 2017) describe pages of 8 to 16 KB in blocks of hundreds of pages, and Grupp et al. (FAST 2012) found each extra bit per cell cuts program/erase lifetime by 10 to 20 times, so block size, density, and wear all push in the same direction.

Block anatomy
Page (read and program unit), block (erase unit), plane, die, superblock
2008 example part
4 KB page, 256 KB block, 1.5 ms erase, 100,000 cycles (Agrawal et al.)
2D and 3D NAND
128 to 192 pages per block in 2D; up to 576 in 3D (Liu et al.)
Simulation
Four block sizes at equal capacity, 12 write-dominant traces, block-level and hybrid FTL
Cell density
Each extra bit per cell: 4 times write latency, 10 to 20 times lower program/erase lifetime (Grupp et al.)

Numbers

NAND erase block figures as printed in Agrawal et al. (rows 1 to 8), Liu et al. (rows 9 to 13), Cai et al. (row 14), and Grupp et al. (rows 15 and 16). Qualifiers such as 'over' and 'up to' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
Page size in the 2008 example part4KB2 GB die in a 2 x 2 GB package
Erase block size in the 2008 example part256KBSame part
Time to read a page into the register25microsecondsSame part
Time to program a page from the register200microsecondsSame part
Time to erase a block1.5millisecondsSame part
Erase cycles a block endures100,000program/erase cyclesCurrent-generation flash in 2008
Average latency with 4 KB logical pages, TPC-C trace200microsecondsSimulated SSD, eight packages
Average latency with 256 KB logical pages, TPC-C traceover 20milliseconds (lower bound)Same simulation, logical page equal to the block
Pages per block in 2D NAND128 to 192pagesTypical planar flash
Pages per block in 3D NANDup to 576pages (upper end)Vertical-channel 3D NAND
Write amplification with 72-page blocks2.1ratio of flash writes to host writesBlock-level FTL, write-intensive traces, same capacity
Write amplification with 576-page blocks5.1ratio of flash writes to host writesSame simulation
Write latency cut from partial erase44.3 to 47.9percentBlock-level FTL (44.3%) and hybrid FTL (47.9%)
Page size in a modern SSD8 to 16KBNAND flash memory in SSDs
Write latency increase per extra bit stored per cell4timesFlash chips measured by the authors
Program/erase lifetime loss per extra bit stored per cell10 to 20times lowerSame chips

This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 8 are from Agrawal et al. (ATC 2008), rows 9 to 13 from Liu et al. (FAST 2018), row 14 from Cai et al. (Proceedings of the IEEE 2017), and rows 15 and 16 from Grupp et al. (FAST 2012). The sources use different parts and years.

Method

The figures are copied from the PDFs of Agrawal, Prabhakaran, Wobber, Davis, Manasse, and Panigrahy, ATC 2008; Liu, Kotra, et al., FAST 2018; Grupp, Davis, and Swanson, FAST 2012; and the arXiv version of Cai, Ghose, Haratsch, Luo, and Mutlu, Proceedings of the IEEE 2017, not refit. Agrawal et al. use a simulator built from the parameters of a specific 2008 flash part. Liu et al. simulate SSDs with four block sizes at the same capacity (blocks per plane and pages per block of 15104 by 72, 7552 by 144, 3776 by 288, and 1888 by 576) on 12 write-dominant traces with a block-level and a hybrid FTL. Grupp et al. measured flash chips. Cai et al. is a survey. The derived value of 64 pages per block is 256 KB divided by 4 KB.

Limits

The numbers come from different years and parts: a 2008 part, a 2018 3D NAND simulation, and a 2012 chip study, and current drives have different page sizes, erase times, and endurance, so the 2008 times and the 100,000 cycle limit are not current figures. The write amplification result is from a trace-driven simulation with a block-level FTL; the authors note that page-level FTLs do not show the same increase,, so the 2.1 to 5.1 rise does not describe every drive design. Partial erase is a proposal that needs chip support and is not a shipping feature. Pages-per-block ranges are as printed and vary by vendor and generation. Grupp et al. report lifetime effects per extra bit for the chips they tested. The sources do not give erase times for 3D NAND beyond saying they are longer.

What the paper found

The answer. A NAND erase block is the unit of erase, and its size sets the cost of garbage collection. Flash is read and programmed per page, from 2 KB to 32 KB as printed by Liu et al. and 8 to 16 KB in the survey of Cai et al., but a page can be rewritten only after its whole block is erased. In the 2008 part of Agrawal et al. a block was 256 KB, 64 pages of 4 KB, and it took 1.5 ms to erase against 25 microseconds to read and 200 microseconds to program a page, with 100,000 erase cycles per block. Pages must also be written in order inside a block.

Why block size matters. Liu et al. report 128 to 192 pages per block in typical 2D NAND and up to 576 in 3D NAND, at least 3 times as many as 2D, because stacking layers adds pages to a block and not blocks to a die. In their simulation with a block-level FTL at fixed capacity, average write amplification on write-intensive traces rose from 2.1 with 72-page blocks to 5.1 with 576-page blocks, since each erase forces a copy of every valid page in the victim block. Their proposed partial erase cut write latency by 44.3% for a block-level FTL and 47.9% for a hybrid FTL. They also note that page-level FTLs avoid most of this growth.

The logical page trap. Agrawal et al. show the same arithmetic from the host side: when the logical page was a full 256 KB block, the average latency of a TPC-C trace was over 20 ms, against 200 microseconds with 4 KB logical pages, because every smaller write became a read-modify-write of the whole unit. The lesson is that mapping unit and erase unit are different decisions, and a mismatch costs two orders of magnitude.

Density and wear. Grupp et al. found each additional bit per cell increases write latency by 4 times and cuts program/erase lifetime by 10 to 20 times, with shrinking returns in density of 2 times, 1.5 times, and 1.3 times. Together with bigger blocks, this makes each erase both more expensive in time and more consequential for wear. The four sources use different parts, years, and methods, so the sizes are not comparable, but they agree on the direction.

Background on drive types is on the drive technology card. Warning signs before a failure are also covered in the SMART failure signals analysis.

Sources

  1. 01Design Tradeoffs for SSD Performance, Agrawal, Prabhakaran, Wobber, Davis, Manasse, and Panigrahy, USENIX ATC 2008 (USENIX PDF, section 2 and table 1) · accessed 2026-10-11
  2. 02PEN: Design and Evaluation of Partial-Erase for 3D NAND-Based High Density SSDs, Liu, Kotra, et al., USENIX FAST 2018 (USENIX PDF, sections 2 and 6) · accessed 2026-10-11
  3. 03The Bleak Future of NAND Flash Memory, Grupp, Davis, and Swanson, USENIX FAST 2012 (USENIX PDF, introduction) · accessed 2026-10-11
  4. 04Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives, Cai, Ghose, Haratsch, Luo, and Mutlu, Proceedings of the IEEE, September 2017 (arXiv 1706.08642, section 2) · accessed 2026-10-11