Skip to content

Stat analysis

Do hot SSDs fail more, and what do the thermal-throttle counters actually show?

Published 2026-10-10

In the one large public field study that measured SSD temperature (Meza et al., SIGMETRICS 2015, Facebook servers, snapshot of November 2014), heat raised failure rates only on the platforms whose SSDs rarely throttled: at 30 to 40 °C all platforms showed similar or slightly rising failure rates, and above 40 °C two platforms (A and B, almost no throttling) showed rising rates, two (C and E, aggressive throttling) were less temperature-sensitive, and two (D and F) showed falling rates that the authors link to young drives in the early-failure period; some controllers start acting from around 80 °C and in the extreme case shut the SSD down. For hard disks, Google's 2007 study of more than 100,000 drives found no consistent rise of failures with temperature at moderate temperatures. What an operator can read from outside is narrow: the NVMe health log keeps minutes above the warning and critical thresholds and, on drives that implement them, transition counts and seconds spent in thermal management, but a zero can mean either never or not implemented, and no opened source gives a failure rate for a given count.

SSD field evidence
Six Facebook server platforms, snapshot November 2014; temperature from sensors embedded on the SSD cards (Meza et al., SIGMETRICS 2015).
Throttle start
Some controllers limit activity at temperatures starting around 80 °C and may shut the SSD down in the extreme case.
Hard-disk contrast
More than 100,000 consumer ATA drives, 9-month window, December 2005 to August 2006: very little correlation of failure with temperature (Pinheiro et al., FAST 2007).
Not covered
Failure rates by throttle count, drives made after 2014, enclosure airflow and heatsink effects, and a second independent measurement of SSD temperature against failure.

SSD failure rate versus average SSD temperature, as described in the text

Wording of the paper's section 5.1 for its Figure 10 (failure rate against average operating temperature, sensors on the SSD cards) and Figure 11 (fraction of SSDs ever throttled). A failure here means an uncorrectable error the SSD could not fix but the host could. Platform labels are the paper's; platforms B, D and F have two SSDs per machine.
Temperature band (°C)PlatformsTrend of failure rate as statedThrottling as statedSource
30 to 40All sixSimilar failure rates or slight increases as temperature risesNot separated by platform in the textSIGMETRICS '15 paper
40 and aboveA and BTemperature-sensitive: failure rate increases with temperatureNo machines or few machines throttledSIGMETRICS '15 paper
40 and aboveC and ELess temperature-sensitive, relatively low failure rateThrottled more aggressively across a range of temperaturesSIGMETRICS '15 paper
40 and aboveD and FFailure rate decreases as temperature risesRelatively low amount of throttling; SSDs mostly in early-failure period, which the authors say likely explains the trendSIGMETRICS '15 paper
Starting around 80Some controllersController tries to keep the SSD below temperature thresholdsReduces access frequency or, in the extreme case, shuts the SSD downSIGMETRICS '15 paper

Thermal fields in the NVMe SMART / Health log

Field names and units from the struct nvme_smart_log manual page (Log Identifier 02h) and the nvme-cli WDC vendor command page. Temperatures are in kelvins in the log; times are as stated per field.
FieldUnitWhat it recordsHow a zero readsSource
temperature (Composite Temperature)kelvinCurrent composite temperature of controller and namespaces; computed in an implementation-specific way and may not match any physical pointNot applicablenvme_smart_log man page
warning_temp_timeminutesOperational time with composite temperature at or above WCTEMP and below CCTEMPAlways 0 if WCTEMP or CCTEMP is 0, whatever the temperaturenvme_smart_log man page
critical_comp_timeminutesOperational time with composite temperature at or above CCTEMPAlways 0 if CCTEMP is 0nvme_smart_log man page
thm_temp1_trans_count and thm_temp2_trans_countcount (does not wrap at FFFFFFFFh)Times the controller moved to lower-power states or took vendor-specific thermal actions after the composite temperature rose above Thermal Management Temperature 1 or 2Never happened, or field not implementednvme_smart_log man page
thm_temp1_total_time and thm_temp2_total_timesecondsTotal time in those states; the nvme-cli WDC page calls 1 light throttle and 2 heavy throttleNever happened, or field not implementednvme-cli WDC page
Thermal shutdown thresholdtemperature, WDC vendor command onlyTemperature of device shutdown, shown only on WDC devices that support the commandResults for other vendors are undefinednvme-cli WDC page

Hard disks: what the 2007 Google study says about temperature

Findings from Pinheiro, Weber and Barroso, FAST 2007. Temperatures were read from SMART records every few minutes during the observation window.
ItemValue (unit)What the paper statesCross-checkSource
Populationmore than 100,000 drivesSerial and parallel ATA consumer drives, 5400 to 7200 rpm, 80 to 400 GB, put into production in or after 2001Data collected December 2005 to August 2006FAST '07 paper
Observation window9 monthsSMART temperature every few minutes; averages, maxima, time above a value and threshold-crossing counts showed similar trendsRepairs database spans about 5 yearsFAST '07 paper
Average temperature versus failure1 °C binsFailures do not increase with average temperature; lower temperatures are associated with higher failure rates, with a slight reversal only at very high temperaturesDiffers from earlier vendor-reported modelsFAST '07 paper
Older drives3 and 4 years oldThe trend of higher failures at higher temperature is more constant and more pronouncedEarlier effects confirmed only for the high end of the range and older drivesFAST '07 paper
Earlier claims cited15 °C change; 30 to 40 °CEarlier work said a 15 °C delta can nearly double failure rates; a vendor report showed MTBF dropping by as much as 50% from 30 to 40 °CNot confirmed in this populationFAST '07 paper

What each public signal can and cannot tell about heat

A reading of the sources above, not a measurement. Each row is a statement the opened documents support.
SignalCan showCannot show (per sources)UnitSource
Composite temperature, one readingCurrent state of controller and namespacesTemperature at a physical point such as a flash die; historykelvinnvme_smart_log man page
Warning and critical timeCumulative operation above the vendor thresholdsAnything when a threshold is 0; whether the time harmed the driveminutesnvme_smart_log man page
Thermal management counts and timesThat the drive reduced power or took cooling action, and for how longWhether the field is implemented (0 is ambiguous); the performance cost, which the Facebook paper says it did not examinecount; secondsnvme_smart_log man page
Fleet average temperaturePopulation trend by platform and throttle behaviourA per-drive failure forecast; the Facebook data are a snapshot, not a time series°CSIGMETRICS '15 paper
SMART temperature on disksA drive's operating temperature over timeA temperature that predicts failure at moderate temperatures in the Google population°CFAST '07 paper

Reading the numbers

Heat did not act alone in the one SSD fleet that measured it. In the Facebook data, the platforms whose SSDs were rarely throttled (A and B) showed failure rates rising with temperature above 40 °C, and the platforms that throttled aggressively (C and E) did not. The authors attribute the low rates on C and E to the aggressive throttling, but they state that they cannot directly measure what the controllers did and that throttling could cost performance, which they could not examine. So the evidence supports 'in this fleet, aggressive throttling went with lower failure rates at high temperature', not a measured causal effect.

Not every platform fits the simple story. D and F, which throttled little, showed failure rates that fall as temperature rises; the authors say these SSDs were mostly in the early-failure period, so the trend likely reflects age rather than heat. A fleet's temperature curve mixes drive age, model and enclosure design, and a one-time snapshot cannot separate them.

On hard disks the best public contrast is Google's 2007 result: over more than 100,000 drives in a 9-month window, higher average temperature did not mean more failures at moderate temperatures; the effect appeared at the high end and mainly in 3- and 4-year-old drives. The Google authors say their data do not allow them to conclude there is no correlation, only that other effects looked more prominent in a professionally managed datacenter. That population is 2005 to 2006 consumer ATA disks, so it says nothing about modern drives.

For an operator reading a drive today, the NVMe log gives minutes above the warning and critical thresholds and, where implemented, how many times and for how many seconds the drive took thermal action. Two limits matter. A zero in the thermal-management fields is ambiguous (never, or not implemented), and a threshold of 0 clears the minute counters whatever the temperature. No opened source ties any of these counters to a failure probability, so they show that a drive ran hot or limited itself, not that it will fail. The NVMe health-log note on this site covers the other fields of the same log.

Related: the NVMe health-log note, SSD power-loss protection, what fails most often on HDDs, SSDs and links and the SMART failure-signal note.

Method

Everything here was read on 2026-10-10 from documents opened that day: the SIGMETRICS 2015 paper PDF by Meza, Wu, Kumar and Mutlu (section 5.1 and the captions of Figures 10 and 11); the FAST 2007 paper PDF by Pinheiro, Weber and Barroso (temperature section and related work); the FAST 2016 paper PDF by Schroeder, Lagisetty and Merchant (one passage comparing the two SSD studies); the struct nvme_smart_log manual page as hosted by Ubuntu; and the nvme-cli manual page for the WDC vendor temperature command. Numbers are copied from the text of those documents. Values that exist only inside figures are not used; where a figure carries the result, the page gives only what the paper states in words.

Limits

The SSD evidence is one operator's fleet (Facebook, six server platforms, PCIe SSDs of 2014 or earlier, snapshot November 2014), so it shows which patterns existed, not what a drive made today does. The paper says it cannot directly measure what the controllers did in response to heat and uses 'ever throttled' as a correlated event; it also says it cannot examine the performance cost of throttling. Platforms D and F trend the other way, which the authors attribute to young drives in the early-failure period rather than to cooling. The FAST 2016 Google SSD paper states that its data do not account for temperature, so only one opened source measures SSD temperature against failure; the NVMe field definitions are a second source for what the log records, not a second failure measurement. The disk study covers December 2005 to August 2006 on 80 to 400 GB consumer ATA drives and cannot be carried over to current drives. The NVMe composite temperature is, in the manual page text, implementation specific and may not represent any physical point. No opened source gives a failure probability for a given warning-time, critical-time or throttle-count value, or the share of NVMe drives that implement the counters. The page covers public failure analysis only; nothing here concerns encryption, drive security features, or recovering anyone else's media.

Sources

  1. 01Meza, Wu, Kumar, Mutlu: A Large-Scale Study of Flash Memory Failures in the Field, SIGMETRICS 2015 (paper PDF) · accessed 2026-10-10
  2. 02Pinheiro, Weber, Barroso: Failure Trends in a Large Disk Drive Population, FAST 2007 (paper PDF) · accessed 2026-10-10
  3. 03Schroeder, Lagisetty, Merchant: Flash Reliability in Production: The Expected and the Unexpected, FAST 2016 (paper PDF) · accessed 2026-10-10
  4. 04struct nvme_smart_log: SMART / Health Information Log (Log Identifier 02h), Ubuntu manual page · accessed 2026-10-10
  5. 05nvme-wdc-vs-temperature-stats(1), nvme-cli manual page · accessed 2026-10-10