How do worn-out SSDs actually die, and do Percentage Used and Available Spare warn you?
Published 2026-10-10
Worn-out SSDs did not die the same way in the one public endurance run that logged it: in The Tech Report's experiment (published 12 March 2015) six consumer SSDs wrote from about 700 TB to just over 2.4 PB before failing, four gave warning through SMART life indicators or software messages, two (Samsung 840 and 840 Pro) died without warning while their SMART reserves still looked healthy, and none of the six was usable afterwards: they bricked on a power cycle or were no longer detected. The NVMe health log offers two life gauges, Percentage Used and Available Spare, but its own definitions limit them: Percentage Used is a vendor estimate that may exceed 100 and at 100 'may not indicate an NVM subsystem failure', and at Available Spare below its threshold an alert 'may occur'. In Google's FAST 2016 field data the chance of an uncorrectable error grew roughly linearly with program-erase cycles, with no sharp rise at the rated limit, so a gauge reaching 100 is not a cliff edge.
- Endurance run
- Six SATA SSDs (five models, two HyperX 3K units), written continuously for about 18 months; final article dated 12 March 2015 (The Tech Report).
- Field data
- Ten drive models, SLC, MLC and eMLC, many millions of drive days, one operator's datacenters (Schroeder et al., FAST 2016).
- Public signals
- NVMe SMART log fields Available Spare, Available Spare Threshold and Percentage Used, plus critical-warning bits for spare, reliability and read-only media.
- Not covered
- Data retention after wear, QLC drives, failure rates by Percentage Used value, and vendor warranty terms.
Six consumer SSDs written to death: what each log and symptom showed
| Drive (flash type) | Writes at death (TB or PB) | Warning before death | State after the end | Source |
|---|---|---|---|---|
| Intel 335 Series 240 GB | Media wear indicator ran out shortly after 700 TB | Wear indicator ran out and the drive shifted to read-only; only 1 reallocated sector | Data accessible until a reboot, then the drive bricked itself on the power cycle (by design, per the article) | Tech Report, page 1 |
| Kingston HyperX 3K 240 GB (incompressible data) | Refused to write after 728 TB | Declining SMART life indicator and warning messages; many failures and reallocated sectors between 600 and 728 TB | Data accessible at first; no response after a reboot (Kingston: it will not boot once its NAND reserve is exhausted) | Tech Report, page 1 |
| Samsung 840 Series 250 GB (TLC) | Died before 1 PB | None before death; reallocated sectors from 200 TB and uncorrectable errors around 300 TB and again toward 900 TB, while SMART suggested plenty of reserve | Died without warning | Tech Report, page 1 |
| Corsair Neutron GTX 240 GB | Still working after 1.2 PB; failed after a later reboot | Thousands of reallocated sectors and many warning messages after 1.1 PB; SMART suggested adequate reserve | Not detected by the system after the reboot | Tech Report, page 1 |
| Kingston HyperX 3K 240 GB (compressible data) | Failed after 2.1 PB of writes, following a power outage | Warning messages started after the life attribute flattened out | Hard-locked the machine on access, then not detected; the article says it cannot tell whether the outage caused the failure | Tech Report, page 2 |
| Samsung 840 Pro 256 GB | Just over 2.4 PB; over 7,000 reallocated sectors totalling 10.7 GB of flash | None: Samsung Magician gave a clean bill of health and the used-block counter showed ample reserve | Unresponsive; the system reported the drive no longer connected | Tech Report, page 2 |
Wear-out in Google's field data: program-erase cycles versus errors
| Item | Value (unit) | What the paper states | Why it matters here | Source |
|---|---|---|---|---|
| MLC models, rated PE cycle limit | 3,000 cycles (MLC-A to MLC-D) | Average PE cycles in the data were 529 to 949 per model | Field drives sat well below the rated limit on average | FAST '16 paper |
| SLC and eMLC rated limits | 100,000 cycles (SLC); 10,000 cycles (eMLC) | Average PE cycles were 185 to 860 for SLC and 377 and 607 for eMLC | Rated limits span 3,000 to 100,000 cycles, so one cycle count means different things per flash type | FAST '16 paper |
| Shape of UE growth | Linear in PE cycles | UE probability grows linearly with PE cycles rather than exponentially, with no sharp increase once the PE cycle limit is reached (model MLC-D, limit 3,000) | A gauge reaching 100% is not a sudden cliff in this data | FAST '16 paper |
| Field RBER versus lab, model eMLC-A | 1e-05 (field median at about 600 PE cycles) | Accelerated tests reached comparable rates only after more than 4,000 PE cycles | Lab endurance runs can understate field errors | FAST '16 paper |
| Models that reached the PE limit | RBER 3e-08 to 8e-08 | Three models in the study reached their limit; the paper says this is significantly lower than most RBER reported from lab tests | Reaching the limit did not coincide with extreme error rates in this data | FAST '16 paper |
NVMe wear fields and what each is allowed to mean
| Field or bit | Unit and range | Definition (paraphrased) | Limit stated or implied by the text | Source |
|---|---|---|---|---|
| percent_used | percent; 0 to 255 (values over 254 shown as 255) | Vendor-specific estimate of NVM subsystem life used, from actual usage and the manufacturer's prediction of NVM life; updated once per power-on hour | 100 means the estimated endurance has been consumed but may not indicate failure; the value may exceed 100 | nvme_smart_log man page |
| avail_spare | percent; 0 to 100 | Normalised percentage of remaining spare capacity | A single normalised value; the field gives no block-level detail | nvme_smart_log man page |
| spare_thresh | percent; 0 to 100 (101 to 255 reserved) | When Available Spare falls below this value an asynchronous event completion may occur | The alert is allowed, not guaranteed | nvme_smart_log man page |
| Critical warning, spare bit | bit flag (set or clear) | Available spare capacity has fallen below the threshold | A flag only; no value or trend | nvme_smart_crit man page |
| Critical warning, reliability bit | bit flag | Reliability degraded by significant media errors or an internal error that degrades reliability | The bit does not identify the cause | nvme_smart_crit man page |
| Critical warning, read-only bit | bit flag | All media placed in read-only mode; not set if read-only results from a namespace write-protection change | Shows the state, not the time left before it | nvme_smart_crit man page |
| media_errors | count | Occurrences where the controller detected an unrecovered data integrity error, such as uncorrectable ECC or CRC failure | Counts occurrences, not their cause | nvme_smart_log man page |
What each wear signal can and cannot tell about a dying SSD
| Signal | Can show | Cannot show (per sources) | Unit | Source |
|---|---|---|---|---|
| Percentage Used | A vendor's own estimate of life consumed | That failure is near: the value may reach or pass 100 with failure not indicated, and it rests on the manufacturer's prediction | percent | nvme_smart_log man page |
| Available Spare and spare bit | That spare capacity fell below a vendor threshold | Sudden flash failure: in the endurance run two drives died while SMART reserve still looked healthy | percent; flag | Tech Report, page 1 |
| Reallocated sector count | Flash blocks retired so far, for example over 7,000 sectors (10.7 GB) on the 840 Pro | Remaining life: that drive reported good health and then stopped responding | sectors; GB | Tech Report, page 2 |
| Read-only state | That the drive stopped accepting writes (Intel 335, by design) | That data stays reachable: the 335 bricked on the next power cycle | flag | Tech Report, page 1 |
| PE cycles against the rated limit | A population-level error risk that rises steadily | A sharp threshold at the rated limit | cycles | FAST '16 paper |
Reading the numbers
The headline from the endurance run is that wear-out did not announce itself the same way twice. The Intel 335 Series and the first Kingston HyperX followed their life indicators and warnings to the end, and the second HyperX and the Neutron GTX produced warning messages; the Samsung 840 and 840 Pro died while their SMART attributes showed plenty of reserve. The article says the 840 may have been brought down by a sudden surge of flash failures too severe to counteract, and says the same of the 840 Pro, so those are open possibilities, not established causes.
The quiet exits matter more than the loud ones. The 840 Pro had logged over 7,000 reallocated sectors, 10.7 GB of flash, and Samsung's own utility still reported good health before the drive vanished from the bus; the Corsair Neutron GTX passed 1.1 PB almost cleanly and then showed thousands of reallocated sectors and warnings before it failed to come back after a reboot. A rising reallocation count showed the drives wearing, but in this sample it did not give a countdown.
The failed drives did not stay usable. Intel designed the 335 Series to go read-only and then brick on a power cycle, and the sample did exactly that; Kingston said the HyperX will not boot once its NAND reserve is exhausted; the others were not detected at the end. Read literally, a read-only state or an end-of-life warning is a short window, and a reboot can end it. The second HyperX is a caution against over-reading any one drive: its compressible workload wrote 28% less to the flash than its twin's, it reached about 2.1 PB of host writes, and a power outage obscures its last days.
The Google data give the statistical counterweight. Across ten models the chance of an uncorrectable error rose roughly linearly with PE cycles and showed no sharp rise at the rated limit, and for model eMLC-A the field RBER at about 600 cycles was already at a level that lab tests reached only after more than 4,000 cycles. That fits the NVMe wording that Percentage Used at 100 'may not indicate' failure and may exceed 100: it is a manufacturer's projection, not a measured failure probability. No opened source ties a given Percentage Used value to a failure rate, so a reading of 100 should be taken as 'past the vendor's estimate', not 'about to fail', and a low reading does not rule out a sudden flash failure. The earlier notes on this site cover the rest of the log and the field studies.
Related: the NVMe health-log note, SSD thermal throttling, SSD power-loss protection, what fails most often on HDDs, SSDs and links and the SMART failure-signal note.
Method
Everything here was read on 2026-10-10 from documents opened that day: the two pages of The Tech Report's SSD Endurance Experiment final article by Geoff Gasior, dated 12 March 2015 (Internet Archive copies of 14 March 2015 and 19 July 2018); the FAST 2016 paper PDF by Schroeder, Lagisetty and Merchant (Table 1 and sections 4.2.1, 4.3 and 5.3); and the libnvme manual pages for struct nvme_smart_log and enum nvme_smart_crit as hosted by Ubuntu. Numbers are copied from the text of those documents. Values that exist only inside figures are not used.
Limits
The endurance run is one drive per model (two HyperX units), 2013 to 2015 consumer SATA models, written with a near-constant stream of writes, so it cannot give a durability figure for any model; the article's own account of the second HyperX says some NAND is simply more resistant to wear, and that unit outlasted its twin by roughly a petabyte. The SMART attributes and warnings in that run are vendor-specific attributes read with Windows tools, not the NVMe log. The Google data cover ten drive models in one operator's datacenters and do not examine the NVMe Percentage Used value. No opened source reports how many current NVMe drives raise the spare-below-threshold bit before failing, how many go read-only at 100%, or how many fail with no prior change in these fields. The manual pages describe fields; they are not field measurements, and vendors may implement the gauges differently. The cause of death of the 840, the 840 Pro and the second HyperX is not established: the article says a sudden burst of flash failures may have been responsible for the Samsung deaths, and that a power outage obscures the second HyperX case. The page covers public failure analysis only; nothing here concerns encryption, drive security features, or recovering anyone else's media.
Sources
- 01Gasior: The SSD Endurance Experiment: They're all dead, The Tech Report, 12 March 2015 (page 1, Internet Archive copy of 14 March 2015) · accessed 2026-10-10
- 02Gasior: The SSD Endurance Experiment: They're all dead, page 2 (Internet Archive copy of 19 July 2018) · accessed 2026-10-10
- 03Schroeder, Lagisetty, Merchant: Flash Reliability in Production: The Expected and the Unexpected, FAST 2016 (paper PDF) · accessed 2026-10-10
- 04struct nvme_smart_log: SMART / Health Information Log (Log Identifier 02h), Ubuntu manual page · accessed 2026-10-10
- 05enum nvme_smart_crit: Critical Warning, Ubuntu manual page · accessed 2026-10-10