Skip to content

Stat analysis

How do worn-out SSDs actually die, and do Percentage Used and Available Spare warn you?

Published 2026-10-10

Worn-out SSDs did not die the same way in the one public endurance run that logged it: in The Tech Report's experiment (published 12 March 2015) six consumer SSDs wrote from about 700 TB to just over 2.4 PB before failing, four gave warning through SMART life indicators or software messages, two (Samsung 840 and 840 Pro) died without warning while their SMART reserves still looked healthy, and none of the six was usable afterwards: they bricked on a power cycle or were no longer detected. The NVMe health log offers two life gauges, Percentage Used and Available Spare, but its own definitions limit them: Percentage Used is a vendor estimate that may exceed 100 and at 100 'may not indicate an NVM subsystem failure', and at Available Spare below its threshold an alert 'may occur'. In Google's FAST 2016 field data the chance of an uncorrectable error grew roughly linearly with program-erase cycles, with no sharp rise at the rated limit, so a gauge reaching 100 is not a cliff edge.

Endurance run
Six SATA SSDs (five models, two HyperX 3K units), written continuously for about 18 months; final article dated 12 March 2015 (The Tech Report).
Field data
Ten drive models, SLC, MLC and eMLC, many millions of drive days, one operator's datacenters (Schroeder et al., FAST 2016).
Public signals
NVMe SMART log fields Available Spare, Available Spare Threshold and Percentage Used, plus critical-warning bits for spare, reliability and read-only media.
Not covered
Data retention after wear, QLC drives, failure rates by Percentage Used value, and vendor warranty terms.

Six consumer SSDs written to death: what each log and symptom showed

From the two pages of The Tech Report's final article. Host writes are in terabytes (TB) or petabytes (PB) as the article prints them.
Drive (flash type)Writes at death (TB or PB)Warning before deathState after the endSource
Intel 335 Series 240 GBMedia wear indicator ran out shortly after 700 TBWear indicator ran out and the drive shifted to read-only; only 1 reallocated sectorData accessible until a reboot, then the drive bricked itself on the power cycle (by design, per the article)Tech Report, page 1
Kingston HyperX 3K 240 GB (incompressible data)Refused to write after 728 TBDeclining SMART life indicator and warning messages; many failures and reallocated sectors between 600 and 728 TBData accessible at first; no response after a reboot (Kingston: it will not boot once its NAND reserve is exhausted)Tech Report, page 1
Samsung 840 Series 250 GB (TLC)Died before 1 PBNone before death; reallocated sectors from 200 TB and uncorrectable errors around 300 TB and again toward 900 TB, while SMART suggested plenty of reserveDied without warningTech Report, page 1
Corsair Neutron GTX 240 GBStill working after 1.2 PB; failed after a later rebootThousands of reallocated sectors and many warning messages after 1.1 PB; SMART suggested adequate reserveNot detected by the system after the rebootTech Report, page 1
Kingston HyperX 3K 240 GB (compressible data)Failed after 2.1 PB of writes, following a power outageWarning messages started after the life attribute flattened outHard-locked the machine on access, then not detected; the article says it cannot tell whether the outage caused the failureTech Report, page 2
Samsung 840 Pro 256 GBJust over 2.4 PB; over 7,000 reallocated sectors totalling 10.7 GB of flashNone: Samsung Magician gave a clean bill of health and the used-block counter showed ample reserveUnresponsive; the system reported the drive no longer connectedTech Report, page 2

Wear-out in Google's field data: program-erase cycles versus errors

Figures from the FAST 2016 paper (Table 1 and sections 4.2.1, 4.3 and 5.3). PE cycles are program-erase cycles; RBER is raw bit error rate; UE is uncorrectable error.
ItemValue (unit)What the paper statesWhy it matters hereSource
MLC models, rated PE cycle limit3,000 cycles (MLC-A to MLC-D)Average PE cycles in the data were 529 to 949 per modelField drives sat well below the rated limit on averageFAST '16 paper
SLC and eMLC rated limits100,000 cycles (SLC); 10,000 cycles (eMLC)Average PE cycles were 185 to 860 for SLC and 377 and 607 for eMLCRated limits span 3,000 to 100,000 cycles, so one cycle count means different things per flash typeFAST '16 paper
Shape of UE growthLinear in PE cyclesUE probability grows linearly with PE cycles rather than exponentially, with no sharp increase once the PE cycle limit is reached (model MLC-D, limit 3,000)A gauge reaching 100% is not a sudden cliff in this dataFAST '16 paper
Field RBER versus lab, model eMLC-A1e-05 (field median at about 600 PE cycles)Accelerated tests reached comparable rates only after more than 4,000 PE cyclesLab endurance runs can understate field errorsFAST '16 paper
Models that reached the PE limitRBER 3e-08 to 8e-08Three models in the study reached their limit; the paper says this is significantly lower than most RBER reported from lab testsReaching the limit did not coincide with extreme error rates in this dataFAST '16 paper

NVMe wear fields and what each is allowed to mean

Field names and wording from the libnvme manual pages for struct nvme_smart_log and enum nvme_smart_crit. Percentages are normalised values, not physical measurements.
Field or bitUnit and rangeDefinition (paraphrased)Limit stated or implied by the textSource
percent_usedpercent; 0 to 255 (values over 254 shown as 255)Vendor-specific estimate of NVM subsystem life used, from actual usage and the manufacturer's prediction of NVM life; updated once per power-on hour100 means the estimated endurance has been consumed but may not indicate failure; the value may exceed 100nvme_smart_log man page
avail_sparepercent; 0 to 100Normalised percentage of remaining spare capacityA single normalised value; the field gives no block-level detailnvme_smart_log man page
spare_threshpercent; 0 to 100 (101 to 255 reserved)When Available Spare falls below this value an asynchronous event completion may occurThe alert is allowed, not guaranteednvme_smart_log man page
Critical warning, spare bitbit flag (set or clear)Available spare capacity has fallen below the thresholdA flag only; no value or trendnvme_smart_crit man page
Critical warning, reliability bitbit flagReliability degraded by significant media errors or an internal error that degrades reliabilityThe bit does not identify the causenvme_smart_crit man page
Critical warning, read-only bitbit flagAll media placed in read-only mode; not set if read-only results from a namespace write-protection changeShows the state, not the time left before itnvme_smart_crit man page
media_errorscountOccurrences where the controller detected an unrecovered data integrity error, such as uncorrectable ECC or CRC failureCounts occurrences, not their causenvme_smart_log man page

What each wear signal can and cannot tell about a dying SSD

A reading of the sources above, not a measurement. Each row is a statement the opened documents support.
SignalCan showCannot show (per sources)UnitSource
Percentage UsedA vendor's own estimate of life consumedThat failure is near: the value may reach or pass 100 with failure not indicated, and it rests on the manufacturer's predictionpercentnvme_smart_log man page
Available Spare and spare bitThat spare capacity fell below a vendor thresholdSudden flash failure: in the endurance run two drives died while SMART reserve still looked healthypercent; flagTech Report, page 1
Reallocated sector countFlash blocks retired so far, for example over 7,000 sectors (10.7 GB) on the 840 ProRemaining life: that drive reported good health and then stopped respondingsectors; GBTech Report, page 2
Read-only stateThat the drive stopped accepting writes (Intel 335, by design)That data stays reachable: the 335 bricked on the next power cycleflagTech Report, page 1
PE cycles against the rated limitA population-level error risk that rises steadilyA sharp threshold at the rated limitcyclesFAST '16 paper

Reading the numbers

The headline from the endurance run is that wear-out did not announce itself the same way twice. The Intel 335 Series and the first Kingston HyperX followed their life indicators and warnings to the end, and the second HyperX and the Neutron GTX produced warning messages; the Samsung 840 and 840 Pro died while their SMART attributes showed plenty of reserve. The article says the 840 may have been brought down by a sudden surge of flash failures too severe to counteract, and says the same of the 840 Pro, so those are open possibilities, not established causes.

The quiet exits matter more than the loud ones. The 840 Pro had logged over 7,000 reallocated sectors, 10.7 GB of flash, and Samsung's own utility still reported good health before the drive vanished from the bus; the Corsair Neutron GTX passed 1.1 PB almost cleanly and then showed thousands of reallocated sectors and warnings before it failed to come back after a reboot. A rising reallocation count showed the drives wearing, but in this sample it did not give a countdown.

The failed drives did not stay usable. Intel designed the 335 Series to go read-only and then brick on a power cycle, and the sample did exactly that; Kingston said the HyperX will not boot once its NAND reserve is exhausted; the others were not detected at the end. Read literally, a read-only state or an end-of-life warning is a short window, and a reboot can end it. The second HyperX is a caution against over-reading any one drive: its compressible workload wrote 28% less to the flash than its twin's, it reached about 2.1 PB of host writes, and a power outage obscures its last days.

The Google data give the statistical counterweight. Across ten models the chance of an uncorrectable error rose roughly linearly with PE cycles and showed no sharp rise at the rated limit, and for model eMLC-A the field RBER at about 600 cycles was already at a level that lab tests reached only after more than 4,000 cycles. That fits the NVMe wording that Percentage Used at 100 'may not indicate' failure and may exceed 100: it is a manufacturer's projection, not a measured failure probability. No opened source ties a given Percentage Used value to a failure rate, so a reading of 100 should be taken as 'past the vendor's estimate', not 'about to fail', and a low reading does not rule out a sudden flash failure. The earlier notes on this site cover the rest of the log and the field studies.

Related: the NVMe health-log note, SSD thermal throttling, SSD power-loss protection, what fails most often on HDDs, SSDs and links and the SMART failure-signal note.

Method

Everything here was read on 2026-10-10 from documents opened that day: the two pages of The Tech Report's SSD Endurance Experiment final article by Geoff Gasior, dated 12 March 2015 (Internet Archive copies of 14 March 2015 and 19 July 2018); the FAST 2016 paper PDF by Schroeder, Lagisetty and Merchant (Table 1 and sections 4.2.1, 4.3 and 5.3); and the libnvme manual pages for struct nvme_smart_log and enum nvme_smart_crit as hosted by Ubuntu. Numbers are copied from the text of those documents. Values that exist only inside figures are not used.

Limits

The endurance run is one drive per model (two HyperX units), 2013 to 2015 consumer SATA models, written with a near-constant stream of writes, so it cannot give a durability figure for any model; the article's own account of the second HyperX says some NAND is simply more resistant to wear, and that unit outlasted its twin by roughly a petabyte. The SMART attributes and warnings in that run are vendor-specific attributes read with Windows tools, not the NVMe log. The Google data cover ten drive models in one operator's datacenters and do not examine the NVMe Percentage Used value. No opened source reports how many current NVMe drives raise the spare-below-threshold bit before failing, how many go read-only at 100%, or how many fail with no prior change in these fields. The manual pages describe fields; they are not field measurements, and vendors may implement the gauges differently. The cause of death of the 840, the 840 Pro and the second HyperX is not established: the article says a sudden burst of flash failures may have been responsible for the Samsung deaths, and that a power outage obscures the second HyperX case. The page covers public failure analysis only; nothing here concerns encryption, drive security features, or recovering anyone else's media.

Sources

  1. 01Gasior: The SSD Endurance Experiment: They're all dead, The Tech Report, 12 March 2015 (page 1, Internet Archive copy of 14 March 2015) · accessed 2026-10-10
  2. 02Gasior: The SSD Endurance Experiment: They're all dead, page 2 (Internet Archive copy of 19 July 2018) · accessed 2026-10-10
  3. 03Schroeder, Lagisetty, Merchant: Flash Reliability in Production: The Expected and the Unexpected, FAST 2016 (paper PDF) · accessed 2026-10-10
  4. 04struct nvme_smart_log: SMART / Health Information Log (Log Identifier 02h), Ubuntu manual page · accessed 2026-10-10
  5. 05enum nvme_smart_crit: Critical Warning, Ubuntu manual page · accessed 2026-10-10