What smartctl and smartd can and cannot tell you about a failing drive
Published 2026-10-11
A smartctl health check that says PASSED only means the drive's own boolean status is not failing now (the manual says a failing status means the drive has already failed or predicts failure within 24 hours), and the data back that up as a weak all-clear: Backblaze found that 23.3% of failed drives had none of five watched SMART raw values above zero, and Google found that over 56% of failed drives had no count in any of its four strongest SMART signals. The tools do see real warnings when they fire: in Backblaze's data 42.2% to 44.8% of failed drives (33.0% for attribute 198) had a non-zero raw value in each of five attributes one day before failure, against 0.3% to 4.8% of working drives. smartd by default polls every 1,800 s (30 minutes), so it is a trend and threshold watcher, not a real-time alarm, and it reports what the firmware chooses to expose through vendor-defined attributes.
- Question
- What does each smartctl check and each smartd directive actually look at, what thresholds does the tooling apply, and how much of real drive failure shows up there in two published fleets?
- Evidence types
- Two manual pages from the smartmontools package, one operator blog post and one conference slide deck from Backblaze, one peer-reviewed field study from Google
- Not repeated here
- Per-attribute failure ratios and the Google 60-day figures are on the SMART failure-signal note; the NVMe health log has its own note
- Access date
- All sources opened on 2026-10-11
What each smartctl check reads and what it can say
| Check | What it reads | What the manual says a bad result means | What it does not say | Source |
|---|---|---|---|---|
| -H (ATA) | The boolean result of the SMART RETURN STATUS command | The drive has already failed, or predicts its own failure within the next 24 hours | A pass is only the absence of that flag; the return value may be unknown through some RAID controller or USB bridge firmware, in which case smartctl falls back to comparing pre-failure attributes with their thresholds | smartctl(8) manual |
| -H (NVMe) | The Critical Warning byte of the SMART/Health Information log | A warning bit is set | Health beyond that one byte; the manual gives no list of bits in this section | smartctl(8) manual |
| -H (SCSI) | Additional Sense Code and Qualifier from the Informal Exceptions log page or from sense data | The device reports a failure prediction through that code | Anything the device does not put in that log page | smartctl(8) manual |
| -A attribute table | Per-attribute raw value, normalized value (1 to 254), worst value, threshold (0 to 255) and type, as stored on the device | If normalized is at or below the threshold the attribute has failed; for a pre-failure attribute, failure is imminent; WHEN_FAILED shows FAILING_NOW or In_the_past | Being of type Pre-fail does not mean the disk is about to fail; smartctl does not calculate values, thresholds or types, and the raw-to-physical-unit conversion is not specified by the SMART standard | smartctl(8) manual |
| Old-age (usage) attribute at threshold | The same table, attributes of type Old_age | An advisory that the device has exceeded its intended design life | Not imminent failure; smartd's -f directive reports these and the smartd page says the same | smartd.conf(5) manual |
| -t short / -t long self-test (ATA) | Electrical, mechanical and read performance of the disk, logged in the self-test log | A logged failure in the self-test log (read with -l selftest) | Short is usually under 10 minutes, long takes tens of minutes to several hours; the manual does not state a detection rate for either | smartctl(8) manual |
smartctl exit status: eight separate signals in one number
| Bit | Decimal value | Meaning in the manual | Kind of signal | Source |
|---|---|---|---|---|
| 0 | 1 | Command line did not parse | Tool usage error, not a drive state | smartctl(8) manual |
| 1 | 2 | Device open failed, no IDENTIFY DEVICE structure returned, or device in a low-power mode | Access problem, not a drive state | smartctl(8) manual |
| 2 | 4 | A SMART or other ATA command failed, or a SMART data structure had a checksum error | Could not read the data (includes unknown status through a bridge) | smartctl(8) manual |
| 3 | 8 | SMART status check returned DISK FAILING | Drive says it is failing now | smartctl(8) manual |
| 4 | 16 | Pre-fail attributes found at or below threshold | Attribute failing now | smartctl(8) manual |
| 5 | 32 | Status check returned DISK OK but some attributes have been at or below threshold in the past | History flag | smartctl(8) manual |
| 6 | 64 | The device error log contains records of errors | Error log, not a threshold | smartctl(8) manual |
| 7 | 128 | The self-test log contains records of errors (failed tests superseded by a newer successful extended test are ignored) | Self-test result | smartctl(8) manual |
What smartd watches by default and how often
| Setting | Value | What it means for monitoring | Source |
|---|---|---|---|
| Polling interval | 1,800 s default (30 minutes); minimum allowed 10 s | Up to 48 checks per day (derived: 86,400 / 1,800); an event that appears and clears between two checks is only seen if the drive logged it | smartd.conf(5) manual |
| Directive -a (the default if none is given) | -H -f -t -l error -l selftest -l selfteststs, and for ATA also -C 197 -U 198 | Health status, usage-attribute failures, attribute tracking, error log, self-test log and self-test status; pending-sector and offline-uncorrectable counts via attributes 197 and 198 | smartd.conf(5) manual |
| Pending sector check (-C) | Reports any non-zero count in attribute 197 (default ID); with '+' only when the count has increased between two checks | A pending sector is one the drive could not read and wants to reallocate; writing to it can force reallocation but the 512 bytes stored there are lost | smartd.conf(5) manual |
| Warning e-mail reminders | once, daily or diminishing: diminishing waits 1 day, then 2, then 4 and so on | Default is once unless state persistence (-s) is on, in which case it is daily; the counter resets when the problem is no longer seen | smartd.conf(5) manual |
| Scheduled self-test example | -s (O/../.././(00|06|12|18)|S/../.././01|L/../../6/03) | The manual's own example: offline test every 6 hours, short test at 01:00 daily, long test Saturdays at 03:00; tests run after a polling cycle, so a polling interval above 60 minutes can miss the slot | smartd.conf(5) manual |
How many failed drives raised each flag, compared with working drives
| Measure | Failed drives (%) | Working drives (%) | Ratio failed / working (times, derived) | Source |
|---|---|---|---|---|
| SMART 5, Reallocated Sectors Count | 42.2 | 1.1 | 38.4 | Backblaze MSST 2017 slides |
| SMART 187, Reported Uncorrectable Errors (Seagate only) | 43.5 | 0.5 | 87.0 | Backblaze MSST 2017 slides |
| SMART 188, Command Timeout (Seagate only) | 44.8 | 4.8 | 9.3 | Backblaze MSST 2017 slides |
| SMART 197, Current Pending Sector Count | 43.1 | 0.7 | 61.6 | Backblaze MSST 2017 slides |
| SMART 198, Uncorrectable Sector Count (Seagate only) | 33.0 | 0.3 | 110.0 | Backblaze MSST 2017 slides |
| Any one or more of those five above zero (blog post, 2016) | 76.7 | 4.2 | 18.3 | Backblaze, 6 Oct 2016 |
| SMART 189, High Fly Writes (candidate stat, 2016) | 47.0 | 16.4 | 2.9 | Backblaze, 6 Oct 2016 |
| Google: failed drives with no count in any of four strong signals (scan errors, reallocations, offline reallocations, probational count) | over 56 | not stated | not stated | Pinheiro et al., FAST 2007 |
| Google: failed drives with zero counts on all SMART variables except temperature | over 36 | not stated | not stated | Pinheiro et al., FAST 2007 |
Reading the numbers
The answer. smartctl and smartd read what the drive's firmware chooses to publish, and a clean reading is weak evidence of health. A PASSED status is the absence of a boolean failing flag, which the manual ties to a drive that has already failed or predicts failure within 24 hours; it is not a statement that nothing is wrong. In the field data a large minority of failures came with nothing to read: 23.3% of Backblaze's failed drives (from 76.7% having at least one flag) and over 56% of Google's failed drives against its four strongest signals.
What the tools do well. When a flag is raised it carries a lot of information: Backblaze's reallocated-sector, pending-sector and uncorrectable-error attributes were non-zero on 42.2% to 43.5% of failed drives one day before failure but on 0.5% to 1.1% of working drives, ratios of about 38 to 87 times (derived), and 4.2% of working drives had any of the five against 76.7% of failed ones. smartd turns this into automation without extra code: a one-line DEVICESCAN configuration expands to the health check, usage-attribute check, error log, self-test log and the 197 and 198 counters, polled every 30 minutes, with e-mail reminders that thin out over days. smartctl's exit status separates eight signals, so a script can tell 'drive says failing' (8) from 'could not read the drive' (2 or 4) from 'an attribute failed in the past' (32).
Thresholds are the weakest part. A threshold is a number the manufacturer set, applied to a normalized value the firmware computed, and the smartctl manual states that smartctl only reports them. A pre-fail attribute at its threshold means failure is imminent, but a pre-fail attribute with a bad raw count and a normalized value still above the threshold stays quiet on the health check. This is why operators using raw counts above zero (Backblaze) catch drives that a threshold-only check would still call fine, and also why they get false positives: 4.2% of working drives tripped one of the five flags, and the Backblaze post says a single non-zero value, such as two remapped sectors, can mean little on its own.
Blind spots seen in the sources. First, the ceiling: Google concluded that models built on SMART alone are unlikely to predict individual failures, because over 56% of failures had none of the four strong signals and over 36% had none on any variable except temperature. Second, coverage by vendor: the Backblaze slide says attributes 187, 188 and 198 come from Seagate drives only, so a fleet that mixes vendors cannot apply the same five-flag rule to all of them. Third, path: the smartctl manual warns that a RAID controller or USB bridge can leave the health status unknown, and the smartd manual notes that a polling interval above 60 minutes can skip a scheduled self-test slot. Fourth, timing: a 1,800 s poll cannot see a short event that left no log entry. Fifth, rates versus counts: Backblaze describes SMART 189 as 47.0% against 16.4% (failed against working) and says the signal may lie in clusters of errors over a short time, not in the cumulative count, which a simple greater-than-zero rule would miss.
What to do with this, within the evidence. Use smartd for what it is good at, a cheap standing watch for health status, pending and uncorrectable sectors, error and self-test logs. Treat a non-zero 197 or 198 as a data-integrity warning (the manual says it means some data on the disk is currently unreadable), and treat a clean report as only that. Do not read either fleet as a detection rate for current drives: neither study measures how often an alert preceded a failure by enough time to act, the two fleets used different definitions of failure, and none of the sources opened here covers SSD or NVMe monitoring in numbers.
Related: the SMART failure-signal note, NVMe health log coverage and USB bridge and UAS failure signatures.
Method
Everything here was read on 2026-10-11 from documents opened that day: the Debian unstable manual pages for smartctl(8) and smartd.conf(5) (the pages list package versions 7.4-3 in trixie and 7.5-2 in testing; features the pages call new or experimental are labelled as such below); the Backblaze blog post 'What SMART Stats Tell Us About Hard Drives' (6 October 2016); Backblaze's MSST 2017 slides 'Behind the Curtain of Backblaze Hard Drive Stats' (Klein), read as a PDF; and the FAST 2007 paper 'Failure Trends in a Large Disk Drive Population' (Pinheiro, Weber, Barroso), USENIX HTML edition. Numbers marked derived are my own arithmetic on figures printed in those sources: ratios are the failed-drive percentage divided by the working-drive percentage for the same attribute, and 48 polls per day is 86,400 s divided by the 1,800 s default interval. Exit-status values are 2 to the power of the bit number, which is the rule the smartctl manual itself states for bit 3 (8 = 2^3). The Backblaze and Google figures come from two different fleets, different years and different definitions of failure, so they are placed side by side, not combined.
Limits
The two manual pages are documentation of tool behaviour, not tests: I ran no smartctl or smartd against any drive and did not measure how any drive responds to any command. They describe the ATA, SCSI and NVMe paths unevenly; I did not open the smartmontools wiki, the ATA or NVMe specifications, or any vendor's attribute documentation, so what an individual attribute counts on an individual model is not established here (the smartctl manual itself says vendors use unusual conventions, for example power-on hours stored in minutes). The Backblaze figures are drive counts from one operator's fleet in one data centre; the per-attribute table is one slide from 2017 whose period is not stated on that slide, and the same slide says attributes 187, 188 and 198 are reported by Seagate drives only, so a zero for another vendor may mean the attribute does not exist, not that the count is zero. The blog post does not say how many failed drives had each combination of values; it also says that a single non-zero value can mean little by itself and that predicting failure takes human and artificial judgement beyond the raw numbers. The 76.7% and 23.3% figures are not a prediction accuracy: they count failed drives that had a flag, with no false-positive rate stated beyond the 4.2% of working drives with a flag. The Google study covers more than 100,000 consumer-grade ATA drives from December 2005 to August 2006, with models withheld, so it says nothing about current helium, SMR or SSD products. Nothing opened here measures how often smartd's e-mail alerts arrived before an outage, how often '-H' flipped before a failure, or the behaviour of SMART through USB bridges and hardware RAID except for the manual's own warning that the status return value may be unknown through some of them. NVMe appears only as the manual's statement that health is read from the Critical Warning byte; the NVMe health log is covered on a separate page. Nothing here covers recovering data from other people's drives.
Sources
- 01smartctl(8), smartmontools, Debian unstable manual page · accessed 2026-10-11
- 02smartd.conf(5), smartmontools, Debian unstable manual page · accessed 2026-10-11
- 03Backblaze: What SMART Stats Tell Us About Hard Drives, 6 October 2016 · accessed 2026-10-11
- 04Klein (Backblaze): Behind the Curtain of Backblaze Hard Drive Stats, MSST 2017 slides · accessed 2026-10-11
- 05Pinheiro, Weber, Barroso: Failure Trends in a Large Disk Drive Population, FAST 2007 · accessed 2026-10-11