What the NVMe health log can see: five warning bits, a few counters, and the failures that show up elsewhere
Published 2026-10-10
The NVMe SMART / Health Information log is a 512-byte snapshot of five current-state warning bits (spare low, temperature, reliability degraded, read-only, backup power failed) plus counters for media errors, unsafe shutdowns, error-log entries, temperature time and wear, and it has no field for link, boot, lost-device or slow-I/O faults. In an Alibaba fleet of over one million NVMe SSDs, 99.97% of valid critical-warning readings were zero, only 14.15% of failure tickets with a recorded symptom were drives whose SMART values crossed a threshold, and 40 groups of drives showed no clear correlation between SMART attributes and fail-slow behaviour.
- Log page
- SMART / Health Information, log identifier 02h, 512 bytes (Linux struct nvme_smart_log ends with 280 reserved bytes at offset 232)
- Field evidence, SSD fleet
- Over one million NVMe SSDs at Alibaba; SMART logs from 2019-11-04 to 2020-11-14; about 20,000 failure tickets, about 35% of drives with a recorded cause
- Field evidence, flash errors
- Ten Google flash drive models, millions of drive days, first four years per drive
- Not covered
- SATA SMART attributes, SCSI sense data, tape drive logs; no source on those was opened for this note
What the health log reports
| Field (as the source names it) | Bytes | Unit | What the source says it is | What the source says it is not | Source |
|---|---|---|---|---|---|
| Critical Warning | 0 | bit flags (5 defined, 3 reserved) | Current state of the controller; bits are not persistent | Not a history: a bit clears when the condition clears | Kingston SMART doc |
| Available Spare | 3 | percent (0-100) | Normalised remaining spare capacity | Says nothing about which blocks or dies are weak | Kingston SMART doc |
| Percentage Used | 5 | percent (0-255, may exceed 100) | Vendor estimate of NVM life used from actual usage and the maker's life prediction | 100 'may not indicate an NVM subsystem failure'; values above 254 are shown as 255 | Microsoft Learn |
| Media and Data Integrity Errors | 175:160 | count of occurrences | Unrecovered data integrity errors, including uncorrectable ECC, CRC checksum failure and LBA tag mismatch | Counts occurrences, not which LBAs or how many bytes | Microsoft Learn |
| Unsafe Shutdowns | 159:144 | count | Incremented when no shutdown notification was received before power loss | Not an error on its own; it records an event, not damage | Microsoft Learn |
| Number of Error Information Log Entries | 191:176 | count over controller life | Count of Error Information log entries over the controller life | A count only; the entries live in the separate Error Information log (log identifier 01h in the kernel header) | Microsoft Learn |
| Warning / Critical Composite Temperature Time | 195:192 and 199:196 | minutes | Time operational at or above the warning / critical threshold | Always 0 if the threshold field in Identify Controller is 0 | Microsoft Learn |
The five Critical Warning bits
| Bit | Kernel name | Meaning when set | Failure it covers | Unit | Source |
|---|---|---|---|---|---|
| 0 | NVME_SMART_CRIT_SPARE | Available spare space below the threshold | Spare capacity exhausted | flag (1 bit) | Microsoft Learn |
| 1 | NVME_SMART_CRIT_TEMPERATURE | A temperature is above an over-temperature or below an under-temperature threshold | Thermal excursion | flag (1 bit) | Microsoft Learn |
| 2 | NVME_SMART_CRIT_RELIABILITY | Reliability degraded by significant media-related errors or any internal error | Media or internal degradation | flag (1 bit) | Microsoft Learn |
| 3 | NVME_SMART_CRIT_MEDIA | Media placed in read-only mode | Drive has stopped accepting writes | flag (1 bit) | Microsoft Learn |
| 4 | NVME_SMART_CRIT_VOLATILE_MEMORY | Volatile-memory backup device failed; valid only if the controller has one | Backup-power failure | flag (1 bit) | Microsoft Learn |
How NVMe drives were recorded as failed in one fleet
| Symptom (as the source names it) | Source's definition | Share of tickets with a cause (%) | ARR (% per device-year) | Source |
|---|---|---|---|---|
| I/O | Drive fails to perform a read or write request | 49.55 | 0.40 | ATC 2022 |
| Boot | Drive fails to initiate (for example when mounting a file system) | 19.59 | 0.16 | ATC 2022 |
| Thres. | One or more SMART attributes reached a pre-defined threshold | 14.15 | 0.11 | ATC 2022 |
| Link | Connection error during PCIe transmission, or abnormal bandwidth | 11.07 | 0.09 | ATC 2022 |
| Lost | A functioning drive becomes unfound | 5.65 | 0.05 | ATC 2022 |
Drive-level error types seen in Google's flash fleet
| Error type (as the source names it) | Source's definition | Lowest model (% of drives) | Highest model (% of drives) | Window | Source |
|---|---|---|---|---|---|
| Correctable error | Detected and corrected by the drive's ECC | 96.1 | 99.9 | first 4 years | FAST 2016 |
| Uncorrectable error | A read finds more corrupted bits than the ECC can correct | 20.3 | 90.5 | first 4 years | FAST 2016 |
| Final read error | A read error that persists after retries | 10.9 | 62.7 | first 4 years | FAST 2016 |
| Final write error | A write error that persists after retries | 0.57 | 5.20 | first 4 years | FAST 2016 |
| Meta error | Error accessing drive-internal metadata | 0.00 | 3.68 | first 4 years | FAST 2016 |
| Timeout error | An operation timed out after 3 seconds | 0.00 | 1.64 | first 4 years | FAST 2016 |
Per-command status codes the Linux NVMe driver can name
| Status value (hex) | Kernel name | Status code type | Unit | Source |
|---|---|---|---|---|
| 0x280 | NVME_SC_WRITE_FAULT | Media and data integrity error (0x200) | per-command status | Linux nvme.h |
| 0x281 | NVME_SC_READ_ERROR | Media and data integrity error (0x200) | per-command status | Linux nvme.h |
| 0x282 | NVME_SC_GUARD_CHECK | Media and data integrity error (0x200) | per-command status | Linux nvme.h |
| 0x287 | NVME_SC_UNWRITTEN_BLOCK | Media and data integrity error (0x200) | per-command status | Linux nvme.h |
| 0x300 | NVME_SC_INTERNAL_PATH_ERROR | Path-related error (0x300) | per-command status | Linux nvme.h |
| 0x371 | NVME_SC_HOST_ABORTED_CMD | Path-related error (0x300) | per-command status | Linux nvme.h |
Reading the numbers
The log answers a narrow question well: has the drive itself declared that spare capacity, temperature, reliability, write access or backup power has crossed a line. Each of those five conditions has a defined bit, and the Microsoft reference and the Linux header agree on the bit order. The two numbers that look like life gauges are softer than they appear: Percentage Used is a vendor estimate that is allowed to exceed 100 and, at 100, 'may not indicate an NVM subsystem failure', and Available Spare is a normalised percentage with no information about where the weak blocks are.
In the one large public NVMe fleet study, the warning bits were rarely set: 99.97% of valid critical-warning readings were zero, and in the baseline table the median critical-warning value is 0 for each of the 22 model and capacity rows. Of the tickets that recorded a cause, 14.15% were drives whose SMART values crossed a threshold; 49.55% were I/O failures, 19.59% boot failures, 11.07% link failures and 5.65% lost drives. The source does not say how many of the I/O, boot, link or lost cases had an earlier health-log signal, so this page does not claim any.
Slow drives are the clearest documented miss. The same study found no clear correlation between fail-slow behaviour and SMART attributes (including Critical Warning, program/erase error and CRC error) in any of 40 groups, and concluded that the root causes or symptoms of fail-slow are not well captured by SMART. That is a statement about that fleet and those attributes, not about every drive.
Counters can run far ahead of any warning bit. In Google's flash fleet, between 96.1% and 99.9% of drives per model saw a correctable error, between 20.3% and 90.5% an uncorrectable error and between 0.57% and 5.20% a final write error in four years. Whether the standard media-error counter moves for each of these is vendor-specific and neither source says; what the NVMe log can promise is only that the media-error count covers uncorrectable ECC, CRC and tag-mismatch events. For a failure that matters at the time it happens, the kernel's per-command status (Table 5) and the controller's own error log are the finer-grained records, and link and path faults are named in a different status class from media errors.
Related: the SMART failure-signal noteand what fails most often on HDDs, SSDs and links.
Method
Every figure was copied from a source that was opened on 2026-10-10: the field list and meanings from two independent publications of the NVMe log definition (Microsoft's NVME_HEALTH_INFO_LOG reference and Kingston's SMART attribute document) and the Linux kernel's own nvme.h header; the field evidence from the Alibaba NVMe study (Lu et al., USENIX ATC 2022) and the Google flash study (Schroeder et al., FAST 2016). Each table keeps the study's own unit and window. The mapping from a fault to the log field that could see it is stated only where a source defines both sides; where no source does, the table says so.
Limits
The log definition is taken from two vendor/OS documents that restate the NVMe specification and from the Linux header; the NVMe base specification PDF itself was not opened. Vendors may add fields beyond the standard log (the Alibaba paper says vendors do not necessarily follow one counting mechanism and that its numbers were standardised from manufacturer manuals), so a given drive can expose more or less than the table. Table 3 covers only the roughly 35% of Alibaba drives whose failure ticket recorded a cause, from one operator and one period (SMART logs 2019-11-04 to 2020-11-14). Table 4 comes from Google's custom PCIe flash drives (ten models, first four years) and counts errors seen by the device driver, not NVMe log fields; neither source says how each driver-level error maps onto the media-error counter. The Google paper's summary says 20-63% of drives saw an uncorrectable error, while its Section 3.1 gives the same range for final read errors and its Table 2 gives 20.3-90.5% for uncorrectable errors; this page quotes Table 2 and names each row as the source does. Nothing here covers SATA SMART, SCSI sense data or tape, and nothing is a failure-prediction claim.
Sources
- 01NVME_HEALTH_INFO_LOG structure, Win32 apps, Microsoft Learn · accessed 2026-10-10
- 02SMART Attribute Details (NVMe), Kingston Technology · accessed 2026-10-10
- 03include/linux/nvme.h, Linux kernel, master branch · accessed 2026-10-10
- 04NVMe SSD Failures in the Field: the Fail-Stop and the Fail-Slow, Lu, Xu, Zhang, Zhu, Wang, Zhu, Xue, Li and Wu, USENIX ATC 2022 · accessed 2026-10-10
- 05Flash Reliability in Production: The Expected and the Unexpected, Schroeder, Lagisetty and Merchant, FAST 2016 · accessed 2026-10-10