Do hot SSDs fail more, and what do the thermal-throttle counters actually show?
Published 2026-10-10
In the one large public field study that measured SSD temperature (Meza et al., SIGMETRICS 2015, Facebook servers, snapshot of November 2014), heat raised failure rates only on the platforms whose SSDs rarely throttled: at 30 to 40 °C all platforms showed similar or slightly rising failure rates, and above 40 °C two platforms (A and B, almost no throttling) showed rising rates, two (C and E, aggressive throttling) were less temperature-sensitive, and two (D and F) showed falling rates that the authors link to young drives in the early-failure period; some controllers start acting from around 80 °C and in the extreme case shut the SSD down. For hard disks, Google's 2007 study of more than 100,000 drives found no consistent rise of failures with temperature at moderate temperatures. What an operator can read from outside is narrow: the NVMe health log keeps minutes above the warning and critical thresholds and, on drives that implement them, transition counts and seconds spent in thermal management, but a zero can mean either never or not implemented, and no opened source gives a failure rate for a given count.
- SSD field evidence
- Six Facebook server platforms, snapshot November 2014; temperature from sensors embedded on the SSD cards (Meza et al., SIGMETRICS 2015).
- Throttle start
- Some controllers limit activity at temperatures starting around 80 °C and may shut the SSD down in the extreme case.
- Hard-disk contrast
- More than 100,000 consumer ATA drives, 9-month window, December 2005 to August 2006: very little correlation of failure with temperature (Pinheiro et al., FAST 2007).
- Not covered
- Failure rates by throttle count, drives made after 2014, enclosure airflow and heatsink effects, and a second independent measurement of SSD temperature against failure.
SSD failure rate versus average SSD temperature, as described in the text
| Temperature band (°C) | Platforms | Trend of failure rate as stated | Throttling as stated | Source |
|---|---|---|---|---|
| 30 to 40 | All six | Similar failure rates or slight increases as temperature rises | Not separated by platform in the text | SIGMETRICS '15 paper |
| 40 and above | A and B | Temperature-sensitive: failure rate increases with temperature | No machines or few machines throttled | SIGMETRICS '15 paper |
| 40 and above | C and E | Less temperature-sensitive, relatively low failure rate | Throttled more aggressively across a range of temperatures | SIGMETRICS '15 paper |
| 40 and above | D and F | Failure rate decreases as temperature rises | Relatively low amount of throttling; SSDs mostly in early-failure period, which the authors say likely explains the trend | SIGMETRICS '15 paper |
| Starting around 80 | Some controllers | Controller tries to keep the SSD below temperature thresholds | Reduces access frequency or, in the extreme case, shuts the SSD down | SIGMETRICS '15 paper |
Thermal fields in the NVMe SMART / Health log
| Field | Unit | What it records | How a zero reads | Source |
|---|---|---|---|---|
| temperature (Composite Temperature) | kelvin | Current composite temperature of controller and namespaces; computed in an implementation-specific way and may not match any physical point | Not applicable | nvme_smart_log man page |
| warning_temp_time | minutes | Operational time with composite temperature at or above WCTEMP and below CCTEMP | Always 0 if WCTEMP or CCTEMP is 0, whatever the temperature | nvme_smart_log man page |
| critical_comp_time | minutes | Operational time with composite temperature at or above CCTEMP | Always 0 if CCTEMP is 0 | nvme_smart_log man page |
| thm_temp1_trans_count and thm_temp2_trans_count | count (does not wrap at FFFFFFFFh) | Times the controller moved to lower-power states or took vendor-specific thermal actions after the composite temperature rose above Thermal Management Temperature 1 or 2 | Never happened, or field not implemented | nvme_smart_log man page |
| thm_temp1_total_time and thm_temp2_total_time | seconds | Total time in those states; the nvme-cli WDC page calls 1 light throttle and 2 heavy throttle | Never happened, or field not implemented | nvme-cli WDC page |
| Thermal shutdown threshold | temperature, WDC vendor command only | Temperature of device shutdown, shown only on WDC devices that support the command | Results for other vendors are undefined | nvme-cli WDC page |
Hard disks: what the 2007 Google study says about temperature
| Item | Value (unit) | What the paper states | Cross-check | Source |
|---|---|---|---|---|
| Population | more than 100,000 drives | Serial and parallel ATA consumer drives, 5400 to 7200 rpm, 80 to 400 GB, put into production in or after 2001 | Data collected December 2005 to August 2006 | FAST '07 paper |
| Observation window | 9 months | SMART temperature every few minutes; averages, maxima, time above a value and threshold-crossing counts showed similar trends | Repairs database spans about 5 years | FAST '07 paper |
| Average temperature versus failure | 1 °C bins | Failures do not increase with average temperature; lower temperatures are associated with higher failure rates, with a slight reversal only at very high temperatures | Differs from earlier vendor-reported models | FAST '07 paper |
| Older drives | 3 and 4 years old | The trend of higher failures at higher temperature is more constant and more pronounced | Earlier effects confirmed only for the high end of the range and older drives | FAST '07 paper |
| Earlier claims cited | 15 °C change; 30 to 40 °C | Earlier work said a 15 °C delta can nearly double failure rates; a vendor report showed MTBF dropping by as much as 50% from 30 to 40 °C | Not confirmed in this population | FAST '07 paper |
What each public signal can and cannot tell about heat
| Signal | Can show | Cannot show (per sources) | Unit | Source |
|---|---|---|---|---|
| Composite temperature, one reading | Current state of controller and namespaces | Temperature at a physical point such as a flash die; history | kelvin | nvme_smart_log man page |
| Warning and critical time | Cumulative operation above the vendor thresholds | Anything when a threshold is 0; whether the time harmed the drive | minutes | nvme_smart_log man page |
| Thermal management counts and times | That the drive reduced power or took cooling action, and for how long | Whether the field is implemented (0 is ambiguous); the performance cost, which the Facebook paper says it did not examine | count; seconds | nvme_smart_log man page |
| Fleet average temperature | Population trend by platform and throttle behaviour | A per-drive failure forecast; the Facebook data are a snapshot, not a time series | °C | SIGMETRICS '15 paper |
| SMART temperature on disks | A drive's operating temperature over time | A temperature that predicts failure at moderate temperatures in the Google population | °C | FAST '07 paper |
Reading the numbers
Heat did not act alone in the one SSD fleet that measured it. In the Facebook data, the platforms whose SSDs were rarely throttled (A and B) showed failure rates rising with temperature above 40 °C, and the platforms that throttled aggressively (C and E) did not. The authors attribute the low rates on C and E to the aggressive throttling, but they state that they cannot directly measure what the controllers did and that throttling could cost performance, which they could not examine. So the evidence supports 'in this fleet, aggressive throttling went with lower failure rates at high temperature', not a measured causal effect.
Not every platform fits the simple story. D and F, which throttled little, showed failure rates that fall as temperature rises; the authors say these SSDs were mostly in the early-failure period, so the trend likely reflects age rather than heat. A fleet's temperature curve mixes drive age, model and enclosure design, and a one-time snapshot cannot separate them.
On hard disks the best public contrast is Google's 2007 result: over more than 100,000 drives in a 9-month window, higher average temperature did not mean more failures at moderate temperatures; the effect appeared at the high end and mainly in 3- and 4-year-old drives. The Google authors say their data do not allow them to conclude there is no correlation, only that other effects looked more prominent in a professionally managed datacenter. That population is 2005 to 2006 consumer ATA disks, so it says nothing about modern drives.
For an operator reading a drive today, the NVMe log gives minutes above the warning and critical thresholds and, where implemented, how many times and for how many seconds the drive took thermal action. Two limits matter. A zero in the thermal-management fields is ambiguous (never, or not implemented), and a threshold of 0 clears the minute counters whatever the temperature. No opened source ties any of these counters to a failure probability, so they show that a drive ran hot or limited itself, not that it will fail. The NVMe health-log note on this site covers the other fields of the same log.
Related: the NVMe health-log note, SSD power-loss protection, what fails most often on HDDs, SSDs and links and the SMART failure-signal note.
Method
Everything here was read on 2026-10-10 from documents opened that day: the SIGMETRICS 2015 paper PDF by Meza, Wu, Kumar and Mutlu (section 5.1 and the captions of Figures 10 and 11); the FAST 2007 paper PDF by Pinheiro, Weber and Barroso (temperature section and related work); the FAST 2016 paper PDF by Schroeder, Lagisetty and Merchant (one passage comparing the two SSD studies); the struct nvme_smart_log manual page as hosted by Ubuntu; and the nvme-cli manual page for the WDC vendor temperature command. Numbers are copied from the text of those documents. Values that exist only inside figures are not used; where a figure carries the result, the page gives only what the paper states in words.
Limits
The SSD evidence is one operator's fleet (Facebook, six server platforms, PCIe SSDs of 2014 or earlier, snapshot November 2014), so it shows which patterns existed, not what a drive made today does. The paper says it cannot directly measure what the controllers did in response to heat and uses 'ever throttled' as a correlated event; it also says it cannot examine the performance cost of throttling. Platforms D and F trend the other way, which the authors attribute to young drives in the early-failure period rather than to cooling. The FAST 2016 Google SSD paper states that its data do not account for temperature, so only one opened source measures SSD temperature against failure; the NVMe field definitions are a second source for what the log records, not a second failure measurement. The disk study covers December 2005 to August 2006 on 80 to 400 GB consumer ATA drives and cannot be carried over to current drives. The NVMe composite temperature is, in the manual page text, implementation specific and may not represent any physical point. No opened source gives a failure probability for a given warning-time, critical-time or throttle-count value, or the share of NVMe drives that implement the counters. The page covers public failure analysis only; nothing here concerns encryption, drive security features, or recovering anyone else's media.
Sources
- 01Meza, Wu, Kumar, Mutlu: A Large-Scale Study of Flash Memory Failures in the Field, SIGMETRICS 2015 (paper PDF) · accessed 2026-10-10
- 02Pinheiro, Weber, Barroso: Failure Trends in a Large Disk Drive Population, FAST 2007 (paper PDF) · accessed 2026-10-10
- 03Schroeder, Lagisetty, Merchant: Flash Reliability in Production: The Expected and the Unexpected, FAST 2016 (paper PDF) · accessed 2026-10-10
- 04struct nvme_smart_log: SMART / Health Information Log (Log Identifier 02h), Ubuntu manual page · accessed 2026-10-10
- 05nvme-wdc-vs-temperature-stats(1), nvme-cli manual page · accessed 2026-10-10