Skip to content

Stat analysis

What the NVMe health log can see: five warning bits, a few counters, and the failures that show up elsewhere

Published 2026-10-10

The NVMe SMART / Health Information log is a 512-byte snapshot of five current-state warning bits (spare low, temperature, reliability degraded, read-only, backup power failed) plus counters for media errors, unsafe shutdowns, error-log entries, temperature time and wear, and it has no field for link, boot, lost-device or slow-I/O faults. In an Alibaba fleet of over one million NVMe SSDs, 99.97% of valid critical-warning readings were zero, only 14.15% of failure tickets with a recorded symptom were drives whose SMART values crossed a threshold, and 40 groups of drives showed no clear correlation between SMART attributes and fail-slow behaviour.

Log page
SMART / Health Information, log identifier 02h, 512 bytes (Linux struct nvme_smart_log ends with 280 reserved bytes at offset 232)
Field evidence, SSD fleet
Over one million NVMe SSDs at Alibaba; SMART logs from 2019-11-04 to 2020-11-14; about 20,000 failure tickets, about 35% of drives with a recorded cause
Field evidence, flash errors
Ten Google flash drive models, millions of drive days, first four years per drive
Not covered
SATA SMART attributes, SCSI sense data, tape drive logs; no source on those was opened for this note

What the health log reports

Selected fields of the SMART / Health Information log with the byte position and unit given by the sources, and what the source itself says the value does not mean. Byte positions are from Kingston's table; the Microsoft reference lists the same fields in the same order.
Field (as the source names it)BytesUnitWhat the source says it isWhat the source says it is notSource
Critical Warning0bit flags (5 defined, 3 reserved)Current state of the controller; bits are not persistentNot a history: a bit clears when the condition clearsKingston SMART doc
Available Spare3percent (0-100)Normalised remaining spare capacitySays nothing about which blocks or dies are weakKingston SMART doc
Percentage Used5percent (0-255, may exceed 100)Vendor estimate of NVM life used from actual usage and the maker's life prediction100 'may not indicate an NVM subsystem failure'; values above 254 are shown as 255Microsoft Learn
Media and Data Integrity Errors175:160count of occurrencesUnrecovered data integrity errors, including uncorrectable ECC, CRC checksum failure and LBA tag mismatchCounts occurrences, not which LBAs or how many bytesMicrosoft Learn
Unsafe Shutdowns159:144countIncremented when no shutdown notification was received before power lossNot an error on its own; it records an event, not damageMicrosoft Learn
Number of Error Information Log Entries191:176count over controller lifeCount of Error Information log entries over the controller lifeA count only; the entries live in the separate Error Information log (log identifier 01h in the kernel header)Microsoft Learn
Warning / Critical Composite Temperature Time195:192 and 199:196minutesTime operational at or above the warning / critical thresholdAlways 0 if the threshold field in Identify Controller is 0Microsoft Learn

The five Critical Warning bits

Each defined bit of byte 0, as named in the Linux kernel header and described in the Microsoft reference. The warnings that raise an asynchronous event to the host are configured by the host with Set Features.
BitKernel nameMeaning when setFailure it coversUnitSource
0NVME_SMART_CRIT_SPAREAvailable spare space below the thresholdSpare capacity exhaustedflag (1 bit)Microsoft Learn
1NVME_SMART_CRIT_TEMPERATUREA temperature is above an over-temperature or below an under-temperature thresholdThermal excursionflag (1 bit)Microsoft Learn
2NVME_SMART_CRIT_RELIABILITYReliability degraded by significant media-related errors or any internal errorMedia or internal degradationflag (1 bit)Microsoft Learn
3NVME_SMART_CRIT_MEDIAMedia placed in read-only modeDrive has stopped accepting writesflag (1 bit)Microsoft Learn
4NVME_SMART_CRIT_VOLATILE_MEMORYVolatile-memory backup device failed; valid only if the controller has oneBackup-power failureflag (1 bit)Microsoft Learn

How NVMe drives were recorded as failed in one fleet

Failure symptoms in the Alibaba NVMe study, for drives whose ticket recorded a cause. Share is of those tickets; ARR is the annual replacement rate, device failures divided by device-years, for that symptom across all models.
Symptom (as the source names it)Source's definitionShare of tickets with a cause (%)ARR (% per device-year)Source
I/ODrive fails to perform a read or write request49.550.40ATC 2022
BootDrive fails to initiate (for example when mounting a file system)19.590.16ATC 2022
Thres.One or more SMART attributes reached a pre-defined threshold14.150.11ATC 2022
LinkConnection error during PCIe transmission, or abnormal bandwidth11.070.09ATC 2022
LostA functioning drive becomes unfound5.650.05ATC 2022

Drive-level error types seen in Google's flash fleet

Fraction of drives with at least one error of each type in their first four years, as the minimum and maximum across ten models (MLC-A to MLC-D, SLC-A to SLC-D, eMLC-A, eMLC-B), from Table 2 of the paper. These are driver-reported counters, not NVMe health-log fields.
Error type (as the source names it)Source's definitionLowest model (% of drives)Highest model (% of drives)WindowSource
Correctable errorDetected and corrected by the drive's ECC96.199.9first 4 yearsFAST 2016
Uncorrectable errorA read finds more corrupted bits than the ECC can correct20.390.5first 4 yearsFAST 2016
Final read errorA read error that persists after retries10.962.7first 4 yearsFAST 2016
Final write errorA write error that persists after retries0.575.20first 4 yearsFAST 2016
Meta errorError accessing drive-internal metadata0.003.68first 4 yearsFAST 2016
Timeout errorAn operation timed out after 3 seconds0.001.64first 4 yearsFAST 2016

Per-command status codes the Linux NVMe driver can name

Selected status values defined in the kernel header, shown as hexadecimal 'status code type + status code'. They describe one failed command, not a cumulative counter, and the kernel prints names for them only when CONFIG_NVME_VERBOSE_ERRORS is enabled (otherwise the helper returns the generic string 'I/O Error').
Status value (hex)Kernel nameStatus code typeUnitSource
0x280NVME_SC_WRITE_FAULTMedia and data integrity error (0x200)per-command statusLinux nvme.h
0x281NVME_SC_READ_ERRORMedia and data integrity error (0x200)per-command statusLinux nvme.h
0x282NVME_SC_GUARD_CHECKMedia and data integrity error (0x200)per-command statusLinux nvme.h
0x287NVME_SC_UNWRITTEN_BLOCKMedia and data integrity error (0x200)per-command statusLinux nvme.h
0x300NVME_SC_INTERNAL_PATH_ERRORPath-related error (0x300)per-command statusLinux nvme.h
0x371NVME_SC_HOST_ABORTED_CMDPath-related error (0x300)per-command statusLinux nvme.h

Reading the numbers

The log answers a narrow question well: has the drive itself declared that spare capacity, temperature, reliability, write access or backup power has crossed a line. Each of those five conditions has a defined bit, and the Microsoft reference and the Linux header agree on the bit order. The two numbers that look like life gauges are softer than they appear: Percentage Used is a vendor estimate that is allowed to exceed 100 and, at 100, 'may not indicate an NVM subsystem failure', and Available Spare is a normalised percentage with no information about where the weak blocks are.

In the one large public NVMe fleet study, the warning bits were rarely set: 99.97% of valid critical-warning readings were zero, and in the baseline table the median critical-warning value is 0 for each of the 22 model and capacity rows. Of the tickets that recorded a cause, 14.15% were drives whose SMART values crossed a threshold; 49.55% were I/O failures, 19.59% boot failures, 11.07% link failures and 5.65% lost drives. The source does not say how many of the I/O, boot, link or lost cases had an earlier health-log signal, so this page does not claim any.

Slow drives are the clearest documented miss. The same study found no clear correlation between fail-slow behaviour and SMART attributes (including Critical Warning, program/erase error and CRC error) in any of 40 groups, and concluded that the root causes or symptoms of fail-slow are not well captured by SMART. That is a statement about that fleet and those attributes, not about every drive.

Counters can run far ahead of any warning bit. In Google's flash fleet, between 96.1% and 99.9% of drives per model saw a correctable error, between 20.3% and 90.5% an uncorrectable error and between 0.57% and 5.20% a final write error in four years. Whether the standard media-error counter moves for each of these is vendor-specific and neither source says; what the NVMe log can promise is only that the media-error count covers uncorrectable ECC, CRC and tag-mismatch events. For a failure that matters at the time it happens, the kernel's per-command status (Table 5) and the controller's own error log are the finer-grained records, and link and path faults are named in a different status class from media errors.

Related: the SMART failure-signal noteand what fails most often on HDDs, SSDs and links.

Method

Every figure was copied from a source that was opened on 2026-10-10: the field list and meanings from two independent publications of the NVMe log definition (Microsoft's NVME_HEALTH_INFO_LOG reference and Kingston's SMART attribute document) and the Linux kernel's own nvme.h header; the field evidence from the Alibaba NVMe study (Lu et al., USENIX ATC 2022) and the Google flash study (Schroeder et al., FAST 2016). Each table keeps the study's own unit and window. The mapping from a fault to the log field that could see it is stated only where a source defines both sides; where no source does, the table says so.

Limits

The log definition is taken from two vendor/OS documents that restate the NVMe specification and from the Linux header; the NVMe base specification PDF itself was not opened. Vendors may add fields beyond the standard log (the Alibaba paper says vendors do not necessarily follow one counting mechanism and that its numbers were standardised from manufacturer manuals), so a given drive can expose more or less than the table. Table 3 covers only the roughly 35% of Alibaba drives whose failure ticket recorded a cause, from one operator and one period (SMART logs 2019-11-04 to 2020-11-14). Table 4 comes from Google's custom PCIe flash drives (ten models, first four years) and counts errors seen by the device driver, not NVMe log fields; neither source says how each driver-level error maps onto the media-error counter. The Google paper's summary says 20-63% of drives saw an uncorrectable error, while its Section 3.1 gives the same range for final read errors and its Table 2 gives 20.3-90.5% for uncorrectable errors; this page quotes Table 2 and names each row as the source does. Nothing here covers SATA SMART, SCSI sense data or tape, and nothing is a failure-prediction claim.

Sources

  1. 01NVME_HEALTH_INFO_LOG structure, Win32 apps, Microsoft Learn · accessed 2026-10-10
  2. 02SMART Attribute Details (NVMe), Kingston Technology · accessed 2026-10-10
  3. 03include/linux/nvme.h, Linux kernel, master branch · accessed 2026-10-10
  4. 04NVMe SSD Failures in the Field: the Fail-Stop and the Fail-Slow, Lu, Xu, Zhang, Zhu, Wang, Zhu, Xue, Li and Wu, USENIX ATC 2022 · accessed 2026-10-10
  5. 05Flash Reliability in Production: The Expected and the Unexpected, Schroeder, Lagisetty and Merchant, FAST 2016 · accessed 2026-10-10