Skip to content

Research notes / Storage integrity

Disk scrubbing: detection time is not a durability guarantee

Published 11 October 2026 · Literature synthesis, not a hardware benchmark

Scrubbing can expose damage before a rebuild needs the affected data. That does not make a clean scan a prediction of future survival, or every detected error an instance of irrecoverable data loss.

Three boundaries to keep separate

A latent sector error becomes visible when the affected location is accessed. Bairavasundaram and colleagues distinguish this from a drive's internally recovered error. Their NetApp study also separates media verification inside the drive from data scrubs that read and compare higher-level checksums. Recovery can require reconstruction from another member before rewriting the sector. [1, sections 2-3]

Hesela's interpretation checklist. These outcomes are not interchangeable.
Observed outcomeQuestion still openEvidence to retain
An error was detectedCan another valid copy reconstruct the data?Address, error type, redundancy state
Redundancy disagreesWhich content is correct?Checksum coverage and repair policy
A scan finished cleanlyWhat did it actually verify?Scope, completion and unreadable regions

The historical field observations are conditioned on that operator's logging, remapping and removal policies. Section 4 documents missing write-error reports for nearline drives and censoring after removal. They are not an unfiltered sample of every physical defect. [1]

The overlooked variable: scan order

Oprea and Juels proposed staggered scrubbing: divide the address space into regions and visit a segment of each before advancing to the next segment. Finding a problem triggers closer inspection of its region. This aims to discover a damaged region earlier, not merely to scan adjacent sectors faster after discovery. [3, section 4.2]

Their evaluation considers seek overhead and a modeled single-drive latent-error exposure metric, MLET. It is not array mean time to data loss. The paper explicitly leaves system-level redundancy analysis and transfer to flash as future work. Treat scan-order design as an interesting hypothesis for a controlled evaluation, not a universal SSD tuning recipe. [3, sections 2 and 6]

A discovery timestamp is not a fault timestamp

Schroeder, Damouras and Gill reused a subset of the earlier NetApp observations; this is not an independent fleet replication. Their traces identify error-discovery time, while actual onset is hidden within the preceding scrub interval. They compare alternative onset assignments when evaluating scan policies. [2, sections 2 and 5.1]

In their trace-driven evaluation, staggered policies reduced mean detection time by 10-20% at 7-14-day scan intervals. Under their independent, randomly assigned onset-time scenario, gains fell to 2-5%. These are model-conditional comparisons, not measured reductions in lost files. [2, sections 5.2.3-5.2.4]

Our inference: an apparent change in error frequency can reflect a change in observation. Preserve the last successful verification time and the detection time separately. Treat the interval between them as uncertainty, rather than assigning all damage to the day a maintenance job ran.

What do MD and ZFS check?

Linux MD's check requests a full redundancy check; the documentation warns that some RAID levels may also repair. Do not describe it as a universally read-only operation. Its mismatch_cnt counts sectors rewritten, or that would be rewritten, and page-sized processing can inflate that count relative to individual errors. It is not a count of damaged files. [4]

OpenZFS documents ordinary scrubbing as verification of pool data against block checksums, with repair supported by redundant devices. Resilvering instead targets data known to be out of date. Scrub completion, repair and unrepairable errors are reported separately; no redundancy can recover arbitrary damage once all valid copies are gone. [5]

These are different verification boundaries. Neither label alone establishes application-level correctness. A valid checksum over the wrong application output is still possible. Consult documentation for the deployed version before acting; this review does not execute maintenance commands or prescribe an aggressive scan schedule.

A useful reliability record

Our reporting checklist, not a measured experiment:

  1. Name the verification layer: media readability, redundancy consistency, block checksum or application validation.
  2. Record scan scope and completion, including pauses, skipped regions and outstanding errors.
  3. Separate detected, repaired and unrecoverable outcomes. Keep their units explicit.
  4. Record missing members and the validation available for reconstructed data.
  5. Retain timestamps and policy changes before comparing periods or fitting an onset model.
  6. Test restoration separately. A scrub is not an independent backup or a restore rehearsal.

This review does not reproduce either FAST study, estimate contemporary failure rates or identify an optimal interval. See the separate prediction-assisted scrubbing review and RAID write-hole analysis for different failure mechanisms.

Structured definitions: latent sector error, data scrubbing, staggered scrubbing and RAID 5.

Primary sources

  1. Bairavasundaram, Goodson, Pasupathy and Schindler. An Analysis of Latent Sector Errors in Disk Drives. SIGMETRICS 2007, sections 2-4.
  2. Schroeder, Damouras and Gill. Understanding latent sector errors and how to protect against them. FAST 2010, sections 2 and 5.
  3. Oprea and Juels. A Clean-Slate Look at Disk Scrubbing. FAST 2010, sections 2, 4.2 and 6.
  4. Linux kernel documentation: RAID arrays, sync_action and mismatch_cnt (retrieved 11 October 2026).
  5. OpenZFS master documentation: zpool-scrub(8), description and operation boundaries (retrieved 11 October 2026).

Reviewed 11 October 2026, Europe/Warsaw. No third-party figures or datasets reproduced.