ZNStore: Decomposing the B+Tree Write Path for Zoned Namespace SSDs
Abstract
Zoned Namespace (ZNS) SSDs expose a sequential-write zone interface that reduces flash-level write amplification and offers cost and capacity advantages over conventional SSDs. Yet this interface creates a semantic mismatch for B+Tree-based storage engines, whose updates are fine-grained and logically in-place. Our write-path analysis reveals that state-of-the-art designs resolve this mismatch by coupling buffering with physical-address binding too early, inflating write latency and underutilizing the device. This paper presents ZNStore, a write-optimized B+Tree key-value store for raw ZNS SSDs that decouples buffering from address binding and restores device parallelism through three mechanisms. Leveled Record Chain buffers fine-grained updates in memory and consolidates them into page-sized writes. Address Determination on Eviction defers physical-address binding until eviction, eliminating premature serialization under high concurrency. PipeBuf coordinates IO dispatch through two complementary techniques: a host-side polling scheduler that separates latency-critical reads from writebacks before requests reach the raw device, and deterministic Physical Page Address reservation that aggregates page evictions into zone-aligned sequential writes. On a real ZNS SSD, ZNStore improves benchmark load throughput by 3.91 × and 2.98 × over RocksDB-on-ZenFS and the state-of-the-art ZKV, respectively, and reduces p99.99 write latency by 95.6% relative to ZKV.