Skip to content

Research notes / Filesystem persistence

Atomic rename is not a durable commit

11 October 2026 / Literature and API review

Replacing a filename atomically answers a visibility question. It does not, on its own, establish that the new bytes and their name will survive a system crash.

One path, two kinds of state

Linux rename() can replace an existing destination without a moment when another process finds that destination absent. Existing open descriptors remain attached to their files. This is a namespace operation, not an instruction to flush the new file's contents. [1]

The fsync() manual distinguishes file data and metadata from the directory entry that makes the file reachable by name. Synchronizing the file does not necessarily persist that entry; Linux documents an additional sync on the directory descriptor. [2]

For a configuration update, ask two separate questions: can a reader observe a partial replacement while the system is running, and which version must be reachable after recovery? A passing test of the first does not settle the second.

A scoped replacement sequence

The following is Hesela's engineering synthesis of those API boundaries, not a universal library implementation. Scope: one writer replaces a regular file inside the same, already durable directory on a local Linux filesystem with working file and directory synchronization. The storage stack must honor flush completion.

  1. Create a fresh temporary file in that directory. Write the complete replacement, handle short writes and errors, and set the required file metadata.
  2. Successfully synchronize the temporary file before publishing its name.
  3. Rename that temporary file over the destination and check the result.
  4. Successfully synchronize the containing directory, then acknowledge a durable update to the caller.

This is an ordering argument: prepare the object, persist it, publish its name, persist the name, then acknowledge. Merely placing calls in source-code order is not the same as waiting for their successful completion.

Review checkpoints under the stated assumptions, not an exhaustive catalogue of crash outcomes.
Last completed stageWhat it establishesWhat remains unresolved
Write temporary fileReplacement bytes supplied to the filesystemRequired persistence has not been established
Sync temporary fileFile synchronization completed successfullyThe final destination name has not been committed
Rename over targetThe new target is visible through the namespaceDirectory persistence still needs confirmation
Sync directoryThe scoped protocol's name-persistence boundary completedThe caller may still not have received an acknowledgement

A failed directory sync after rename is not proof that the replacement was rolled back: the new name is already visible. Record an indeterminate durability outcome and reconcile through the application's recovery protocol. Do not silently equate a retry with repair; see why successful fsync retries can mislead.

Cross-directory moves, newly created ancestor directories, concurrent writers, network filesystems and multi-file transactions are outside this sequence. They add state and failure cases that need their own protocol. Replacing an inode also requires an explicit policy for permissions, ownership and other metadata.

Why an unsafe shortcut can appear reliable

Ext4 documents auto_da_alloc: for recognized replacement patterns in data=ordered mode, it arranges for delayed-allocation data to reach disk before the rename is committed at the next journal commit. That can prevent a zero-length replacement after a crash. It is a particular allocation and ordering behavior, not a promise that returning from rename means the next commit has already completed. [4]

Pillai and colleagues' BOB study examined six Linux filesystems in sixteen configurations. ALICE separately explored application protocols using abstract persistence models. Section 4.4 distinguishes safe-rename heuristics from safe-file-flush behavior and missing directory synchronization. Their results warn against treating favorable ordering on one configuration as an application guarantee everywhere. They are historical evidence, not a test of today's releases. [3]

A crash test needs a recovery oracle

Mohan and colleagues' CrashMonkey and Ace work is useful precisely because it states its bounds. The study reproduced 24 of 26 historical unique bugs and found ten new ones. The small-workload observation excludes prerequisite operations from its core-operation count. Testing focuses on explicit persistence points and checks recovered data and metadata, rather than treating successful filesystem repair as sufficient. [5, sections 3-6]

Section 4.4 excludes crashes inside operations and alternative I/O reorderings; resource-exhaustion or long-history bugs can also escape the selected workload space. These results do not prove that a short test validates every crash outcome. [5]

For the replacement above, define the oracle before testing. After an acknowledged update, require the exact new bytes at the intended name with required metadata. Before acknowledgement, specify allowed old/new states and recovery of abandoned temporary files. A crash after persistence but before a reply creates uncertainty for the caller even when the stored version is correct.

  • Record the filesystem, kernel, mount options, cache and flush assumptions, initial directory state and every return value.
  • Test interruption around each protocol boundary in an isolated disposable environment. Include error paths separately from power-loss cases.
  • Inspect recovered storage state and application invariants; an ordinary read after killing only the process may still use the surviving operating-system cache.

Evidence boundary

This article reviews full-text methods and limitations from two papers plus Linux documentation. No power interruption, fault injection, current-kernel certification or reproduction of either paper was performed. The checkpoint table is an explanatory synthesis, not an empirical result or failure-rate estimate. API support and filesystem behavior must be verified for the deployed environment.

Corpus: directory fsync, persistence ordering, crash consistency. Related: RAID write-hole protection.

Primary sources

  1. Linux man-pages 6.19, rename(2): DESCRIPTION and BUGS.
  2. Linux man-pages 6.19, fsync(2): DESCRIPTION and error handling.
  3. Pillai, Chidambaram, Alagappan, Al-Kiswany, Arpaci-Dusseau and Arpaci-Dusseau. All File Systems Are Not Created Equal: On the Complexity of Crafting Crash-Consistent Applications. OSDI 2014, pp. 433-448, sections 2 and 4.4.
  4. Linux kernel documentation, ext4 General Information: auto_da_alloc and Data Mode, retrieved 11 October 2026.
  5. Mohan, Martinez, Ponnapalli, Raju and Chidambaram. Finding Crash-Consistency Bugs with Bounded Black-Box Crash Testing. OSDI 2018, pp. 33-50, sections 3-6.

Reviewed 11 October 2026. Publisher records verify the paper titles, authors, venues and page ranges; no unverified DOI is asserted.