Skip to content

Stat analysis

What smartctl and smartd can and cannot tell you about a failing drive

Published 2026-10-11

A smartctl health check that says PASSED only means the drive's own boolean status is not failing now (the manual says a failing status means the drive has already failed or predicts failure within 24 hours), and the data back that up as a weak all-clear: Backblaze found that 23.3% of failed drives had none of five watched SMART raw values above zero, and Google found that over 56% of failed drives had no count in any of its four strongest SMART signals. The tools do see real warnings when they fire: in Backblaze's data 42.2% to 44.8% of failed drives (33.0% for attribute 198) had a non-zero raw value in each of five attributes one day before failure, against 0.3% to 4.8% of working drives. smartd by default polls every 1,800 s (30 minutes), so it is a trend and threshold watcher, not a real-time alarm, and it reports what the firmware chooses to expose through vendor-defined attributes.

Question
What does each smartctl check and each smartd directive actually look at, what thresholds does the tooling apply, and how much of real drive failure shows up there in two published fleets?
Evidence types
Two manual pages from the smartmontools package, one operator blog post and one conference slide deck from Backblaze, one peer-reviewed field study from Google
Not repeated here
Per-attribute failure ratios and the Google 60-day figures are on the SMART failure-signal note; the NVMe health log has its own note
Access date
All sources opened on 2026-10-11

What each smartctl check reads and what it can say

From the smartctl(8) manual page as opened on 2026-10-11. Normalized values and thresholds are unitless integers (normalized 1 to 254, threshold 0 to 255) computed by the drive's firmware, not by smartctl.
CheckWhat it readsWhat the manual says a bad result meansWhat it does not saySource
-H (ATA)The boolean result of the SMART RETURN STATUS commandThe drive has already failed, or predicts its own failure within the next 24 hoursA pass is only the absence of that flag; the return value may be unknown through some RAID controller or USB bridge firmware, in which case smartctl falls back to comparing pre-failure attributes with their thresholdssmartctl(8) manual
-H (NVMe)The Critical Warning byte of the SMART/Health Information logA warning bit is setHealth beyond that one byte; the manual gives no list of bits in this sectionsmartctl(8) manual
-H (SCSI)Additional Sense Code and Qualifier from the Informal Exceptions log page or from sense dataThe device reports a failure prediction through that codeAnything the device does not put in that log pagesmartctl(8) manual
-A attribute tablePer-attribute raw value, normalized value (1 to 254), worst value, threshold (0 to 255) and type, as stored on the deviceIf normalized is at or below the threshold the attribute has failed; for a pre-failure attribute, failure is imminent; WHEN_FAILED shows FAILING_NOW or In_the_pastBeing of type Pre-fail does not mean the disk is about to fail; smartctl does not calculate values, thresholds or types, and the raw-to-physical-unit conversion is not specified by the SMART standardsmartctl(8) manual
Old-age (usage) attribute at thresholdThe same table, attributes of type Old_ageAn advisory that the device has exceeded its intended design lifeNot imminent failure; smartd's -f directive reports these and the smartd page says the samesmartd.conf(5) manual
-t short / -t long self-test (ATA)Electrical, mechanical and read performance of the disk, logged in the self-test logA logged failure in the self-test log (read with -l selftest)Short is usually under 10 minutes, long takes tens of minutes to several hours; the manual does not state a detection rate for eithersmartctl(8) manual

smartctl exit status: eight separate signals in one number

From the EXIT STATUS section of the smartctl(8) manual page, ATA meanings. Value is 2 to the power of the bit number (decimal, unitless); a returned status can have several bits set at once. Values are derived from the bit numbers by the rule the manual states for bit 3.
BitDecimal valueMeaning in the manualKind of signalSource
01Command line did not parseTool usage error, not a drive statesmartctl(8) manual
12Device open failed, no IDENTIFY DEVICE structure returned, or device in a low-power modeAccess problem, not a drive statesmartctl(8) manual
24A SMART or other ATA command failed, or a SMART data structure had a checksum errorCould not read the data (includes unknown status through a bridge)smartctl(8) manual
38SMART status check returned DISK FAILINGDrive says it is failing nowsmartctl(8) manual
416Pre-fail attributes found at or below thresholdAttribute failing nowsmartctl(8) manual
532Status check returned DISK OK but some attributes have been at or below threshold in the pastHistory flagsmartctl(8) manual
664The device error log contains records of errorsError log, not a thresholdsmartctl(8) manual
7128The self-test log contains records of errors (failed tests superseded by a newer successful extended test are ignored)Self-test resultsmartctl(8) manual

What smartd watches by default and how often

From the smartd.conf(5) manual page as opened on 2026-10-11. Intervals in seconds (s) or days; counts are unitless. Directives marked NEW EXPERIMENTAL in the manual are not counted as defaults.
SettingValueWhat it means for monitoringSource
Polling interval1,800 s default (30 minutes); minimum allowed 10 sUp to 48 checks per day (derived: 86,400 / 1,800); an event that appears and clears between two checks is only seen if the drive logged itsmartd.conf(5) manual
Directive -a (the default if none is given)-H -f -t -l error -l selftest -l selfteststs, and for ATA also -C 197 -U 198Health status, usage-attribute failures, attribute tracking, error log, self-test log and self-test status; pending-sector and offline-uncorrectable counts via attributes 197 and 198smartd.conf(5) manual
Pending sector check (-C)Reports any non-zero count in attribute 197 (default ID); with '+' only when the count has increased between two checksA pending sector is one the drive could not read and wants to reallocate; writing to it can force reallocation but the 512 bytes stored there are lostsmartd.conf(5) manual
Warning e-mail remindersonce, daily or diminishing: diminishing waits 1 day, then 2, then 4 and so onDefault is once unless state persistence (-s) is on, in which case it is daily; the counter resets when the problem is no longer seensmartd.conf(5) manual
Scheduled self-test example-s (O/../.././(00|06|12|18)|S/../.././01|L/../../6/03)The manual's own example: offline test every 6 hours, short test at 01:00 daily, long test Saturdays at 03:00; tests run after a polling cycle, so a polling interval above 60 minutes can miss the slotsmartd.conf(5) manual

How many failed drives raised each flag, compared with working drives

Backblaze: percent of drives with a raw value above zero, failed drives measured one day before failure (MSST 2017 slide, 'Behind the Curtain of Backblaze Hard Drive Stats'); ratio is derived (failed % / working %). The group figures come from the 6 October 2016 blog post (67,814 drives). Google: percent of failed drives, FAST 2007, with its own definition of failure (drive replaced during a repair).
MeasureFailed drives (%)Working drives (%)Ratio failed / working (times, derived)Source
SMART 5, Reallocated Sectors Count42.21.138.4Backblaze MSST 2017 slides
SMART 187, Reported Uncorrectable Errors (Seagate only)43.50.587.0Backblaze MSST 2017 slides
SMART 188, Command Timeout (Seagate only)44.84.89.3Backblaze MSST 2017 slides
SMART 197, Current Pending Sector Count43.10.761.6Backblaze MSST 2017 slides
SMART 198, Uncorrectable Sector Count (Seagate only)33.00.3110.0Backblaze MSST 2017 slides
Any one or more of those five above zero (blog post, 2016)76.74.218.3Backblaze, 6 Oct 2016
SMART 189, High Fly Writes (candidate stat, 2016)47.016.42.9Backblaze, 6 Oct 2016
Google: failed drives with no count in any of four strong signals (scan errors, reallocations, offline reallocations, probational count)over 56not statednot statedPinheiro et al., FAST 2007
Google: failed drives with zero counts on all SMART variables except temperatureover 36not statednot statedPinheiro et al., FAST 2007

Reading the numbers

The answer. smartctl and smartd read what the drive's firmware chooses to publish, and a clean reading is weak evidence of health. A PASSED status is the absence of a boolean failing flag, which the manual ties to a drive that has already failed or predicts failure within 24 hours; it is not a statement that nothing is wrong. In the field data a large minority of failures came with nothing to read: 23.3% of Backblaze's failed drives (from 76.7% having at least one flag) and over 56% of Google's failed drives against its four strongest signals.

What the tools do well. When a flag is raised it carries a lot of information: Backblaze's reallocated-sector, pending-sector and uncorrectable-error attributes were non-zero on 42.2% to 43.5% of failed drives one day before failure but on 0.5% to 1.1% of working drives, ratios of about 38 to 87 times (derived), and 4.2% of working drives had any of the five against 76.7% of failed ones. smartd turns this into automation without extra code: a one-line DEVICESCAN configuration expands to the health check, usage-attribute check, error log, self-test log and the 197 and 198 counters, polled every 30 minutes, with e-mail reminders that thin out over days. smartctl's exit status separates eight signals, so a script can tell 'drive says failing' (8) from 'could not read the drive' (2 or 4) from 'an attribute failed in the past' (32).

Thresholds are the weakest part. A threshold is a number the manufacturer set, applied to a normalized value the firmware computed, and the smartctl manual states that smartctl only reports them. A pre-fail attribute at its threshold means failure is imminent, but a pre-fail attribute with a bad raw count and a normalized value still above the threshold stays quiet on the health check. This is why operators using raw counts above zero (Backblaze) catch drives that a threshold-only check would still call fine, and also why they get false positives: 4.2% of working drives tripped one of the five flags, and the Backblaze post says a single non-zero value, such as two remapped sectors, can mean little on its own.

Blind spots seen in the sources. First, the ceiling: Google concluded that models built on SMART alone are unlikely to predict individual failures, because over 56% of failures had none of the four strong signals and over 36% had none on any variable except temperature. Second, coverage by vendor: the Backblaze slide says attributes 187, 188 and 198 come from Seagate drives only, so a fleet that mixes vendors cannot apply the same five-flag rule to all of them. Third, path: the smartctl manual warns that a RAID controller or USB bridge can leave the health status unknown, and the smartd manual notes that a polling interval above 60 minutes can skip a scheduled self-test slot. Fourth, timing: a 1,800 s poll cannot see a short event that left no log entry. Fifth, rates versus counts: Backblaze describes SMART 189 as 47.0% against 16.4% (failed against working) and says the signal may lie in clusters of errors over a short time, not in the cumulative count, which a simple greater-than-zero rule would miss.

What to do with this, within the evidence. Use smartd for what it is good at, a cheap standing watch for health status, pending and uncorrectable sectors, error and self-test logs. Treat a non-zero 197 or 198 as a data-integrity warning (the manual says it means some data on the disk is currently unreadable), and treat a clean report as only that. Do not read either fleet as a detection rate for current drives: neither study measures how often an alert preceded a failure by enough time to act, the two fleets used different definitions of failure, and none of the sources opened here covers SSD or NVMe monitoring in numbers.

Related: the SMART failure-signal note, NVMe health log coverage and USB bridge and UAS failure signatures.

Method

Everything here was read on 2026-10-11 from documents opened that day: the Debian unstable manual pages for smartctl(8) and smartd.conf(5) (the pages list package versions 7.4-3 in trixie and 7.5-2 in testing; features the pages call new or experimental are labelled as such below); the Backblaze blog post 'What SMART Stats Tell Us About Hard Drives' (6 October 2016); Backblaze's MSST 2017 slides 'Behind the Curtain of Backblaze Hard Drive Stats' (Klein), read as a PDF; and the FAST 2007 paper 'Failure Trends in a Large Disk Drive Population' (Pinheiro, Weber, Barroso), USENIX HTML edition. Numbers marked derived are my own arithmetic on figures printed in those sources: ratios are the failed-drive percentage divided by the working-drive percentage for the same attribute, and 48 polls per day is 86,400 s divided by the 1,800 s default interval. Exit-status values are 2 to the power of the bit number, which is the rule the smartctl manual itself states for bit 3 (8 = 2^3). The Backblaze and Google figures come from two different fleets, different years and different definitions of failure, so they are placed side by side, not combined.

Limits

The two manual pages are documentation of tool behaviour, not tests: I ran no smartctl or smartd against any drive and did not measure how any drive responds to any command. They describe the ATA, SCSI and NVMe paths unevenly; I did not open the smartmontools wiki, the ATA or NVMe specifications, or any vendor's attribute documentation, so what an individual attribute counts on an individual model is not established here (the smartctl manual itself says vendors use unusual conventions, for example power-on hours stored in minutes). The Backblaze figures are drive counts from one operator's fleet in one data centre; the per-attribute table is one slide from 2017 whose period is not stated on that slide, and the same slide says attributes 187, 188 and 198 are reported by Seagate drives only, so a zero for another vendor may mean the attribute does not exist, not that the count is zero. The blog post does not say how many failed drives had each combination of values; it also says that a single non-zero value can mean little by itself and that predicting failure takes human and artificial judgement beyond the raw numbers. The 76.7% and 23.3% figures are not a prediction accuracy: they count failed drives that had a flag, with no false-positive rate stated beyond the 4.2% of working drives with a flag. The Google study covers more than 100,000 consumer-grade ATA drives from December 2005 to August 2006, with models withheld, so it says nothing about current helium, SMR or SSD products. Nothing opened here measures how often smartd's e-mail alerts arrived before an outage, how often '-H' flipped before a failure, or the behaviour of SMART through USB bridges and hardware RAID except for the manual's own warning that the status return value may be unknown through some of them. NVMe appears only as the manual's statement that health is read from the Critical Warning byte; the NVMe health log is covered on a separate page. Nothing here covers recovering data from other people's drives.

Sources

  1. 01smartctl(8), smartmontools, Debian unstable manual page · accessed 2026-10-11
  2. 02smartd.conf(5), smartmontools, Debian unstable manual page · accessed 2026-10-11
  3. 03Backblaze: What SMART Stats Tell Us About Hard Drives, 6 October 2016 · accessed 2026-10-11
  4. 04Klein (Backblaze): Behind the Curtain of Backblaze Hard Drive Stats, MSST 2017 slides · accessed 2026-10-11
  5. 05Pinheiro, Weber, Barroso: Failure Trends in a Large Disk Drive Population, FAST 2007 · accessed 2026-10-11