Skip to content

Research notes / Zoned storage

ZNS resource limits: why closing a zone may not unblock writes

11 October 2026 / Literature review and worked examples

A ZNS SSD can reject a write to an empty zone while plenty of capacity remains. Closing another partially written zone may not help: open-zone slots and active-zone slots are different resources.

Count states, not just writers

Open-zone usage counts implicitly and explicitly open zones. Active-zone usage also counts closed zones. A partially written zone that is closed therefore still occupies an active slot. These are namespace resource budgets, not counts of application threads. [1, section 2.1.1.4]

The synthetic snapshot below has four zones, an open limit of two and an active limit of three. Every command shown has completed successfully; there are no concurrent state changes. The final column counts empty zones that could be opened using the remaining slots, not guaranteed successful I/O.

Hesela worked example. All counts are zones, not measurements from an SSD.
SnapshotZones A / B / C / DOpen / 2Active / 3Empty zones openable
Initial snapshotexplicit-open / explicit-open / closed / empty230
Close the first partially written zoneclosed / explicit-open / closed / empty130
Finish that closed zonefull / explicit-open / closed / empty121

After closing A, the open count falls from two to one but the active count stays at three. D still needs an active slot. B can continue appending within its remaining capacity; reopening C would consume an open slot but no additional active slot. Finishing A frees an active slot, at a capacity cost.

Do not reset a live zone to resolve a resource alarm. Reset makes its old data inaccessible. Finish instead makes the zone full, preventing further writes until reset; neither operation is interchangeable with Close. [6]

The address span is not the writable capacity

ZNS distinguishes zone size from zone capacity. A smaller capacity leaves an unwritable tail inside the zone's address span. Applications must respect the capacity reported for each zone, not assume that the next zone boundary is their write limit. [2]

For a synthetic zone with 65,536 addressable LBAs, capacity 61,440 LBAs, 16,384 LBAs already written and 4,096 data bytes per LBA:

Binary units: one MiB is 1,048,576 bytes. LBA metadata bytes are excluded.
QuantityCalculationMiB
Address span65,536 x 4,096 bytes256
Writable capacity61,440 x 4,096 bytes240
Unwritable tail(65,536 - 61,440) x 4,096 bytes16
Remaining before Finish(61,440 - 16,384) x 4,096 bytes176

Finishing this partially written zone would strand 176 MiB of otherwise writable capacity until reset. That is separate from the 16 MiB tail, which was never writable in this geometry. Reserve policy must account for both the zone-state budget and usable bytes.

Zone Append solves placement, not commit ordering

Regular writes target the current write pointer and need appropriate ordering. Zone Append names the zone; the controller chooses the actual write location and returns it on successful completion. Concurrent append requests can land in a different order from submission. Preserve the returned addresses in the application's mapping. [6]

An assigned address does not make a group of requests a transaction. The pinned ZNS 1.1 specification separately defines atomicity parameters and the append command's FUA persistence behavior, with no implied ordering of other commands. A recovery protocol must still persist and reconcile data and its mapping. [1, sections 2.1 and 3.4.1]

What the research actually establishes

ATC 2021: Bjorling and colleagues compared zoned and conventional interfaces on the same SSD hardware, with adapted f2fs and RocksDB/ZenFS. Table 3 reports a 2,048 MiB zone size, 1,077 MiB zone capacity and fourteen active zones for that device. These are device characteristics, not ZNS-wide constants. Section 5 used Linux 5.9 and RocksDB 6.12. Application layout, preconditioning and spare-space choices belong to the comparison, not just the command set. [4]

Moving reclamation to the host also does not abolish application write amplification: the same paper's Table 2 reports RocksDB write amplification near twelve for its tested SST sizes. Device-side and database-level amplification have different numerators and denominators. Do not relabel a device-side result as a whole-system guarantee. [4, Table 2 and section 5.2]

OSDI 2023: Min and colleagues characterized one commodity small-zone SSD using fio/SPDK, then evaluated eZNS with RocksDB and mixed workloads. Shared dies and write-cache contention still produced interference between zones. Their allocator and scheduling layer improved specific experiments, not a universal latency bound. Its hardware contract assumes particular zone placement behavior; section 4.2 explicitly cautions that firmware may use another policy. Section 4.4.4 also acknowledges copying overhead when reclaiming leased resources. [5]

These findings are compatible: removing one source of background work can improve performance without providing isolation from every competing workload. A zone boundary is not evidence of a dedicated die or a latency service-level guarantee.

A practical diagnosis

Start with the command's actual status and a refreshed zone report. Distinguish a full zone, a capacity boundary and exhausted resource budgets. Then inspect the decoded open and active limits exposed by the software layer. Raw NVMe MAR/MOR are zero-based, with all bits set meaning no limit; a raw zero does not mean unlimited. [1, Figure 48]

Linux zonefs exposes separate maximum and current counts for write-open and active sequential files. Its maximum-count attributes use zero for no limit. With explicit-open, resource acquisition moves to opening a file for writing; closing a partially written file still leaves an active zone. This lets software notice scarcity earlier but does not create extra resources. [3]

Hesela's design implication: budget for application metadata and reclamation progress, not only foreground writers. Before retiring a zone, decide where its surviving records and mappings will remain recoverable. Test the actual firmware, host stack and workload on disposable media.

Scope and reproducibility

No SSD benchmark, power-loss experiment or reproduction of either paper was performed. The tables are deterministic synthetic examples generated at build time from the published calculation module, covered by unit tests. It is not a controller emulator and does not model optional excursions, races, failures or persistence.

The specification reference is pinned to revision 1.1 to make the cited semantics traceable, not to claim it is the newest edition. Check the applicable device revision and supported features before implementation.

Corpus: ZNS, active-zone limit, open-zone limit, zone capacity, Zone Append. Related: discard is not zeroing and visibility versus durable commit.

Primary sources

  1. NVM Express, Zoned Namespace Command Set Specification, revision 1.1 (cover: 18 May 2021), sections 2.1.1.4, 3.4.1, 3.4.3 and Annex A. Copyright 2007-2021 NVM Express, Inc. ALL RIGHTS RESERVED.
  2. Western Digital, Zoned Storage: NVMe Zoned Namespaces devices, capacity and resource limits.
  3. Linux kernel documentation, zonefs: sequential files, explicit-open and runtime sysfs attributes.
  4. Bjorling, Aghayev, Holmberg, Ramesh, Le Moal, Ganger and Amvrosiadis. ZNS: Avoiding the Block Interface Tax for Flash-based SSDs. USENIX ATC 2021, pp. 689-703, sections 3-5 and Table 3.
  5. Min, Zhao, Liu and Krishnamurthy. eZNS: An Elastic Zoned Namespace for Commodity ZNS SSDs. OSDI 2023, pp. 461-477, sections 3-5.
  6. Western Digital, Zoned Storage: device models, zone types, management commands and Zone Append.

Reviewed 11 October 2026. Publisher records and full papers checked; no unverified DOI is asserted.