What do ZFS and btrfs checksum counters catch that SMART cannot see, and where do they stop?
Published 2026-10-11
A filesystem checksum catches the one fault that SMART and the drive's own error log cannot, a block that the drive returned as a successful read but with wrong contents, and it does so per block with a counter you can read (ZFS CKSUM, btrfs corruption_errs). It does not catch everything and it does not replace drive-level signals. In the one published failing drive log with all layers shown (OpenZFS issue 10214, a WD40EFAX during a RAIDZ1 resilver; zpool status pasted at 27.61%, smartctl pasted the next day), SMART overall health said PASSED, the reallocated, pending, uncorrectable and CRC attributes were all 0, the ATA error log held 7,487 errors, and ZFS counted 47.7 K write errors with a CKSUM count of 0: the drive's refusal to write showed up in the write counter and the kernel log, not in the checksum counter. Checksums also stop at the page cache: in a University of Wisconsin fault-injection study of ZFS, every injected on-disk bit flip was detected, but a single memory bit flip made ZFS read corrupt data in 0.6% to 7.1% of runs, depending on the workload. btrfs documents a second gap: files marked NOCOW lose data checksums, so scrub cannot validate their contents.
- Question
- Which faults does a filesystem-level checksum counter (ZFS, btrfs) detect that SMART and the drive's error log do not, and which faults does it miss or misattribute?
- Counters covered
- ZFS READ, WRITE and CKSUM per vdev plus the slow I/O count; btrfs write_io_errs, read_io_errs, flush_io_errs, corruption_errs and generation_errs; scrub results in both
- Cross-layer evidence
- One published failing-drive log with SMART, ATA error log, kernel log and ZFS counters side by side (OpenZFS issue 10214)
- Experimental evidence
- Fault injection into ZFS on disk (9 block classes) and in memory (4 filebench workloads, 100 trials each), Zhang et al., FAST 2010
- Related pages
- Fleet-wide silent corruption rates are on the checksum-mismatch paper note; SMART coverage is on the SMART failure-signal note
What each filesystem counter counts
| Counter | What it counts | Scope | Unit or threshold | Source |
|---|---|---|---|---|
| ZFS READ, WRITE, CKSUM (JSON: read_errors, write_errors, checksum_errors) | I/O read errors, I/O write errors, and checksum errors, one count each per vdev | Each vdev from leaf disk up to pool | events (count) | OpenZFS zpool-status(8) |
| ZFS checksum error | A disk returned data that was expected to be correct but was not. Reported in zpool status and zpool events. If a block cannot be reconstructed, errors are reported for all disks the block is stored on, because the damaged disk cannot be identified | Per block read or scrubbed | events (count) | OpenZFS zpoolconcepts(7) |
| ZFS slow I/O (zpool status -s) | Operations that did not complete within zio_slow_io_ms; the page says this does not necessarily mean they failed | Leaf vdevs | 30,000 ms by default | OpenZFS zpool-status(8) |
| btrfs write_io_errs, read_io_errs, flush_io_errs | Failed writes, failed reads, and failed writes carrying the FLUSH flag, that the layers beneath the filesystem could not satisfy | Per device, persistent across mounts | events (count) | btrfs-device(8) |
| btrfs corruption_errs, generation_errs | A block checksum mismatch or a corrupted metadata header; a block generation that does not match the value expected (for example stored in the parent node) | Per device, updated during use or by scrub | events (count) | btrfs-device(8) |
| btrfs scrub error summary (csum, super, verify, read; Corrected, Uncorrectable, Unverified) | Checksum mismatches, superblock errors, metadata header errors and unreadable blocks; Corrected means repaired from another copy, Unverified means a first read failed but a retry succeeded | Per scrub run; exit status 3 means uncorrectable errors found | events (count) | btrfs-scrub(8) |
One failing drive seen from five layers at once
| Layer and signal | Value as pasted | Reads as a fault? | Unit | Source |
|---|---|---|---|---|
| SMART overall-health self-assessment | PASSED | No | verdict | OpenZFS issue 10214, 2020 |
| SMART attributes 5, 197, 198, 199 (Reallocated, Pending, Offline_Uncorrectable, UDMA_CRC) | 0, 0, 0, 0 | No | raw count | OpenZFS issue 10214, 2020 |
| ATA Extended Comprehensive error log, Device Error Count | 7,487 (log keeps only the most recent 24; the 8 entries pasted are all IDNF) | Yes | errors | OpenZFS issue 10214, 2020 |
| ZFS WRITE counter on the new disk, with READ and CKSUM | WRITE 47.7 K, READ 0, CKSUM 0 | Yes in WRITE only | events | OpenZFS issue 10214, 2020 |
| ZFS pool-level line | state ONLINE; errors: No known data errors | No | status text | OpenZFS issue 10214, 2020 |
What one flipped memory bit does in ZFS (per-workload probability, derived by the authors)
| Workload (filebench) | P1(R), read corrupt data | P1(W), write corrupt data | P1(C), crash or hang | Source |
|---|---|---|---|---|
| varmail | 0.6% [0.2, 1.2] | 0% [0, 0.2] | 0.3% [0.1, 0.8] | Zhang et al., FAST 2010 |
| oltp | 1.9% [1.2, 2.8] | 0.1% [0, 0.5] | 1.1% [0.6, 1.8] | Zhang et al., FAST 2010 |
| webserver | 0.7% [0.3, 1.3] | 1.4% [0.8, 2.2] | 1.3% [0.7, 2.1] | Zhang et al., FAST 2010 |
| fileserver | 7.1% [5.4, 9.0] | 3.6% [2.5, 4.8] | 1.6% [1.0, 2.5] | Zhang et al., FAST 2010 |
Where checksum coverage ends
| Condition | What the checksum layer does | Quantity | Unit | Source |
|---|---|---|---|---|
| Block already in the ZFS page cache | Checksum verified only when the block is read from disk; cached copies are not re-checked, so a memory flip is returned to the application | Window described as unbounded | time in cache | Zhang et al., FAST 2010 |
| Dirty block flushed to disk | A new checksum is computed over the already corrupted block, so the corruption becomes permanent and verifies cleanly later | Dirty blocks stay up to 30 s before flush | s | Zhang et al., FAST 2010 |
| User data with one copy (ZFS default, per the paper) | Detects the corruption and reports an error; cannot repair. Metadata with ditto copies (2 or 3 copies) was repaired from a good copy | 1 copy of user data; 2 of file system metadata; 3 of global metadata | copies | Zhang et al., FAST 2010 |
| Non-redundant pool | Manual page: a single case of bit corruption can render some or all data unavailable | Not stated | n/a | OpenZFS zpoolconcepts(7) |
| btrfs file with NOCOW attribute (chattr +C) | Implicitly sets NODATASUM: metadata still validated by scrub, file data is not; btrfs can return bad contents because it cannot tell which mirror is good | Not stated | n/a | btrfs-scrub(8) |
| btrfs scrub schedule | Scrub only validates what it is run over; the manual page says to run it manually or from a periodic service | Recommended period about 1 month; about 80% device bandwidth on an idle filesystem | months; % of bandwidth | btrfs-scrub(8) |
| ZFS scrub versus resilver | A scrub examines all data and verifies every block's checksum; a resilver examines only data known to be out of date | Not stated | n/a | OpenZFS zpool-scrub(8) |
Reading the numbers
The answer. A checksum counter is the only signal in these sources that catches data the drive handed back as a good read. The Wisconsin authors define disk corruption as exactly that case, distinct from latent sector errors and other conditions where the drive gives an explicit notification, and OpenZFS defines a checksum error the same way: data expected to be correct that was not, which it calls silent data corruption. SMART attributes and the ATA error log only record what the drive itself noticed, so by construction they cannot count this case.
The reverse also holds, and the published log shows it. When the drive refuses writes instead of returning wrong data, the filesystem checksum counter can stay at zero. In issue 10214 the WRITE counter reached 47.7 K while CKSUM stayed 0, SMART said PASSED, and only the ATA error log (7,487 errors) and the kernel log carried the fault. So a monitoring set built on CKSUM alone would have missed this failure, and one built on SMART overall health alone would have missed it too. This is one drive and one user, and a later comment in the same thread shows nonzero CKSUM next to WRITE errors, so read it as an example of two independent signals, not a rule.
Where a checksum is attached matters more than whether one exists. ZFS stores each block's checksum in its parent block pointer, so a lost or misdirected write is caught on read, and in the paper's fault injection every on-disk corruption was detected and no bad data was returned. But in memory the same design lets bad data through: the checksum is not rechecked for cached blocks, and a flip before flush gets a fresh valid checksum. The measured chance that one flipped bit leads to a corrupt read ranged from 0.6% (varmail) to 7.1% (fileserver) in that setup. The paper also notes that, from the software side, a corrupt block read from disk may not be distinguishable from memory corruption.
Detection is not repair, and neither is attribution. With one copy, ZFS reports the error and cannot fix it; with redundancy it repairs from a good copy. OpenZFS says that if a block cannot be reconstructed, checksum errors are reported for all the disks it sits on, so the per-disk CKSUM count points at a set of suspects, not the culprit. btrfs adds a scrub status that separates Corrected from Uncorrectable errors, and its scrub exit status 3 means uncorrectable errors were found.
Coverage is also bounded by what gets read. Both scrub manual pages describe scrub as the pass that verifies checksums over the data, and the btrfs page says to run it periodically, with about one month recommended. Data nobody reads and nobody scrubs is not verified. btrfs has one explicit hole: NOCOW files carry no data checksums, and the manual page notes that systemd journals and libvirt storage directories commonly set that attribute. For triage, treat filesystem checksum counters as one of three layers next to the drive's SMART and error logs and the SCSI sense and kernel log, because each sees faults the others do not.
Related: the SMART failure-signal note, SCSI sense and kernel log signatures, SMR drive failure signatures, the silent-corruption paper note (on its own branch, not yet merged) and NVMe health log coverage.
Method
Everything here was read on 2026-10-11 from pages opened that day: the OpenZFS manual pages zpool-status(8), zpoolconcepts(7) and zpool-scrub(8) (master branch documentation); the btrfs-progs manual pages btrfs-device(8) and btrfs-scrub(8) on btrfs.readthedocs.io (latest); the USENIX FAST 2010 paper 'End-to-end Data Integrity for File Systems: A ZFS Case Study' by Zhang, Rajimwale, Arpaci-Dusseau and Arpaci-Dusseau; and OpenZFS GitHub issue 10214 (opened 2020-04-16, closed 2020-12-22) with its pasted zpool status, smartctl and kernel log output. Numbers and counter names are copied from those texts. Where a row says 'derived' or 'as pasted', it is my reading of a log excerpt, not a figure the source states. The manual pages describe current software and the paper describes a 2009-era Solaris build, so the page separates what each source says about itself. This page does not repeat the fleet-wide checksum-mismatch rates from the Bairavasundaram et al. field study, which the silent-corruption paper note already covers; it only notes where Zhang et al. quote that study.
Limits
The manual pages say what the counters mean, not how often they fire: no source opened here gives a field rate of ZFS CKSUM or btrfs corruption_errs events, so this page cannot say how many fleet failures a checksum counter catches first. The SMART-blindness numbers come from one log set (one WD40EFAX, one user, one resilver, firmware 82.00A82, 540 power-on hours) and a few later comments in the same thread, and the issue was closed with a maintainer writing that, from the discussion, it sounds like the reported issues were caused by hardware, so the cause is not settled. In that thread one comment (a commenter quoting an earlier Blocks and Files comment) says that after a forced resilver of a WD40EFAX a scrub kept producing checksum errors from that drive, and another reports CKSUM counts of 3 and 13 next to WRITE counts of 14 to 18 on CMR drives, so a nonzero CKSUM count can occur beside write failures; counts from a single mirror are not a rule. The Zhang et al. results come from a Solaris Express build 108 virtual machine with 2 GB of non-ECC memory and ZFS pool version 14, one disk, injected faults, and 100 trials per workload with 16 simultaneous bit flips scaled down by the paper's own independence assumption to one flip; they are not field rates and may not describe current OpenZFS or btrfs. The paper's statement that more than 400,000 blocks had checksum mismatches, 8% of them found during RAID reconstruction, is the authors' quotation of a different study and was not re-checked here. Counter names differ by version: the ZFS JSON field names come from one example in the manual page, and btrfs scrub behaviour that depends on kernel version (for example interruption after kernel 6.19) is not covered. Nothing here measures a checksum's false-negative rate, and the sources do not give one. Cryptographic or recovery aspects of these filesystems are out of scope.
Sources
- 01OpenZFS documentation: zpool-status(8) (master) · accessed 2026-10-11
- 02OpenZFS documentation: zpoolconcepts(7) (master), Device Failure and Recovery · accessed 2026-10-11
- 03OpenZFS documentation: zpool-scrub(8) (master) · accessed 2026-10-11
- 04btrfs-progs documentation: btrfs-device(8), Device stats (latest) · accessed 2026-10-11
- 05btrfs-progs documentation: btrfs-scrub(8) (latest) · accessed 2026-10-11
- 06Zhang, Rajimwale, Arpaci-Dusseau, Arpaci-Dusseau: End-to-end Data Integrity for File Systems: A ZFS Case Study (USENIX FAST 2010) · accessed 2026-10-11
- 07OpenZFS GitHub issue 10214: WD WDx0EFAX drive unable to resilver (opened 2020-04-16, closed 2020-12-22) · accessed 2026-10-11