Skip to content

Research notes / NAND programming

NAND paired pages: why a later write can damage earlier data

Published 11 October 2026 / Literature review

Finishing one NAND-page program does not necessarily isolate that data from later programming. Two pages can share the same cells, so a fault during the second operation can damage the first page.

This is a physical failure mechanism, not a claim that every SSD violates its durability contract. Controller design and the exact protection guarantee matter.

A pair is shared storage, not adjacent host addresses

Linux's mtd_pairing_info describes pages sharing cells, with a pair identifier and a group identifying the bit position. Its example pairs page 0 with page 4, not page 1. The comment explicitly allows three pages for TLC despite retaining the word "pair." [1]

That API concerns MTD flash write units within an erase block. It is not a map from an NVMe namespace's logical blocks to physical NAND. Do not infer page pairing from consecutive file offsets or LBAs.

First operationPage A storedOne bit per shared cell
Later operationPage B programmedSame cells, additional bit
Fault boundaryCheck A and BNot just the interrupted page
Conceptual two-bit example, not a voltage plot or a physical address map. An interrupted second program can affect earlier data. [2]

Power loss can corrupt a completed earlier program

Tseng, Grupp and Swanson used an FPGA test platform to interrupt operations on 11 raw flash chips from five vendors. Their MLC experiments include corruption of a first page after power was cut while programming its partner. They call this retroactive data corruption. [3, section 4.1.2]

The outcome was not simply "less time to write means more errors": timing effects were non-monotonic. This is useful experimental evidence against assuming that only the active page is at risk. It is not a probability estimate for today's managed SSDs.

A second vulnerability does not require a power cut

Cai and colleagues characterized multiple 15-19 nm MLC chips with direct FPGA access. Their two-step programming study shows an exposed interval: neighboring programming and read disturb can alter partially programmed cells before the second step. [4, sections 3-4]

In the studied path, the chip rereads lower-page data internally without passing it through the controller's ECC engine. A wrong value can then influence the final program. This does not mean ECC is absent from the SSD. It means where correction runs matters.

The proposed mitigations include retaining lower-page data in controller memory and correcting a reread on a buffer miss. Those are architectural proposals, not host commands or proof that a specific current drive implements them. [4, section 6.1]

What does "power-loss protection" protect?

Micron's 2014 white paper distinguishes protection of existing media contents from protection of writes still buffered or in progress. Its client and enterprise examples promise different coverage. Treat that as a dated vendor description, not a universal classification of current products. [2, pages 3-4]

Hesela's practical interpretation: ask for the exact model's guarantee and validation evidence. A capacitor photograph or a generic PLP label cannot substitute for an explicit contract.

Engineering review questions, not test results or a failure classifier.
BoundaryQuestion to resolveInsufficient evidence
Existing dataAre previously committed contents preserved during later interrupted operations?Checking only the newest write
Buffered dataWhich acknowledged writes must survive under the documented cache and flush policy?A generic successful write completion
Address mappingCan recovered logical addresses still locate the correct data?The device merely reappearing after reboot
Internal correctionWhich internal transfers receive error correction?The presence of an ECC engine

For qualification, define expected post-recovery contents before a test, preserve the operation history, and verify older data as well as the active write range. Power-interruption testing belongs on disposable, isolated test media. We performed no such test for this review.

Method and limits

Literature review of raw-flash experiments, a vendor white paper and version-pinned Linux source. The diagnostic matrix is Hesela's synthesis; no hardware experiment was performed.

Historical chip-level evidence is not a present-day SSD failure rate. Pairing, physical layout and controller protections are device-specific; host logical block addresses do not expose a NAND pairing map.

The two papers address different conditions and generations. Their results are not pooled. No current-product vulnerability, affected-drive count or firmware fix is inferred here.

Continue with SSD power-fault evidence, NAND read-retry and fsync error handling.

Machine-readable definitions: NAND page pairing, program interference and MLC.

Primary sources

  1. Linux v6.12, include/linux/mtd/mtd.h: mtd_pairing_info and mtd_pairing_scheme
  2. Micron. How Micron SSDs Handle Unexpected Power Loss. October 2014, pages 1-4; vendor white paper
  3. Tseng, Grupp and Swanson. Understanding the Impact of Power Loss on Flash Memory. DAC 2011, sections 3 and 4.1.2. DOI: 10.1145/2024724.2024733
  4. Cai, Ghose, Luo, Mai, Mutlu and Haratsch. Vulnerabilities in MLC NAND Flash Memory Programming: Experimental Analysis, Exploits, and Mitigation Techniques. HPCA 2017, sections 2-4 and 6. DOI: 10.1109/HPCA.2017.61

Accessed 11 October 2026, Europe/Warsaw. Original paper figures are not reproduced.