Skip to content

Research notes / Persistence

Why retrying a failed fsync does not prove durability

Published 10 October 2026 · Historical evidence, current documentation, explicit model

An error report and a repair are different events. After a failed writeback, a successful retry of fsync() is not, by itself, proof that the earlier data reached persistent storage.

The counterexample: clean does not mean durable

Rebello and colleagues injected single transient sector or block write faults into ext4, XFS and Btrfs on Linux 5.2.11. Their two micro-workloads exercised an overwrite and repeated appends. Failed writeback could leave pages marked clean; another sync then had nothing to write. The study also used CuttleFS to test five applications: Redis 5.0.7, LMDB 0.9.24, LevelDB 1.22, SQLite 3.30.1 and PostgreSQL 12.0. These are historical configurations, not a verdict on current releases. [1, sections 3-4]

The distinction matters during incident response: an apparently correct read can come from memory. A process restart is not necessarily a cold read from storage, because the operating-system cache can outlive that process. Testing the same value twice through the same cached path is weak evidence of persistence. [1, sections 4-5]

Four cases you can reproduce

This is a small, deterministic illustrative model, not measured filesystem behavior. Every case starts with old data on storage and new data in cache. One writeback fails. The first sync reports the error; subsequent I/O is assumed fault-free. The model varies dirty-state retention, clean-page eviction and an explicit rewrite from an independent good copy.

Model outputs: the retry succeeds in every row, but the persisted value differs.
Assumption after failureRetryCached readRead after cache loss
Clean page retained; retrysuccessnewold
Clean page evicted; retrysuccessoldold
Dirty page retained; retrysuccessnewnew
Clean page; rewrite then syncsuccessnewnew

The final two rows are sensitivity checks, not recovery recipes. A real rewrite may fail again, omit part of a transaction or overwrite the wrong state. This model has no directories, partial writes, metadata, concurrent writers or device caches. No probability or field failure rate can be inferred from four deliberately chosen cases.

Run the dependency-free JavaScript model with Node.js. It prints the machine-readable assumptions and event traces. Its four regression tests check the stale-data case, eviction, changed assumptions and determinism.

The less obvious part: an error cursor is not a repair log

Linux's generic writeback-error infrastructure reports a failure to file descriptions that were open when it occurred. After reporting it, later syncs on the same descriptor can return zero unless another error occurs. Attribution is deliberately coarse: a descriptor can receive an error even if its own writes succeeded. [3]

The underlying errseq_t mechanism combines an error value with sequence tracking. Observers keep a cursor and check what changed since their previous observation. It is not a count of all failures and does not identify every damaged byte. Advancing the cursor records that an error was observed; it does not rewrite anything. [4]

A second boundary: a consistent filesystem is not a committed application

Pillai and colleagues separated persistence ordering from atomicity. Their BOB tool examined six Linux filesystems; ALICE analyzed update protocols in eleven applications and found 60 static crash vulnerabilities. ALICE explored states permitted by abstract persistence models, so that count is neither an incident rate nor 60 bugs guaranteed to appear on every filesystem. [2, sections 2-4]

SQLite makes the ordering issue concrete: its rollback-journal protocol persists recovery information before modifying the database. Its documentation also describes syncing the directory containing a super-journal. Durable file content and a durable name are distinct requirements. Those details belong to the application's commit protocol, not to an SSD health counter. [5, sections 3 and 5]

What to test in your own system

The following is Hesela's engineering synthesis, not an additional experimental result. Define the promise first: which acknowledged transactions must survive, which unacknowledged states are allowed, and what recovery source remains trustworthy?

  • Keep the first I/O error, affected files, transaction boundary and software versions in the incident record. A later zero return must not erase the earlier uncertainty.
  • Test recovery against persistent state, not only a successful in-process read. Distinguish process failure, operating-system failure and device power loss.
  • Check the complete protocol: data, recovery log, names and commit acknowledgement. A checksum can detect a mismatch without supplying the missing version.
  • Use a disposable, isolated test environment for fault injection. Never inject storage faults into the production host serving this site.

The authors publish CuttleFS with a private page cache and selectable failure behaviors. It is a candidate for a separately pinned reproduction environment; we did not install it or reproduce the paper in this session. [6]

Evidence and limits

We inspected both full papers, their methods and result sections, plus the kernel and SQLite documentation on 10 October 2026. No new hardware measurements were collected. Our executable contribution demonstrates a logical counterexample under stated assumptions, not the behavior of any particular current kernel, database or drive. The papers' tested software and fault models bound their conclusions.

Related: SSD power-loss protection. Corpus records: fsync, crash consistency, and writeback error.

Primary sources

  1. Rebello et al., Can Applications Recover from fsync Failures? USENIX ATC 2020, pp. 753-767
  2. Pillai et al., All File Systems Are Not Created Equal. OSDI 2014, pp. 433-448
  3. Linux VFS: Handling errors during writeback
  4. Linux kernel: The errseq_t datatype
  5. SQLite: Atomic Commit
  6. CuttleFS research artifact (MIT license)

Accessed 10 October 2026. No DOI is asserted where none was verified in the publisher record.