Skip to content

D-NOVA: In-Storage Retrieval Accelerator via Dual-Bound 3D NAND-Optimized Similarity Search with Vector Adaptation

Jul 2026 · arXiv.org · Vol abs/2607.17538 · 0 citations · 85 references
Computer Science

TL;DR

D-NOVA is a hardware-software co-designed in-storage retrieval accelerator that executes an inverted file (IVF)-based hierarchical retrieval pipeline by deeply embedding the search functionality directly into the NAND memory array, and introduces a lightweight contrastive adapter that maps embedding vectors into a DTS-friendly domain, recovering near-software recall while improving performance and energy efficiency.

Abstract

Retrieval-Augmented Generation (RAG) enhances the factual grounding of large language model (LLM) inference by retrieving relevant information from external knowledge bases. However, its dense vector retrieval introduces significant latency and energy overhead, becoming the primary performance bottleneck. Although recent in-storage accelerators aim to reduce data movement, they still rely on host or embedded processors outside the memory, where nearly 70% of the total retrieval time is spent. As a result, they cannot fully overcome the bandwidth limitations, leading to yet another memory bottleneck. To tackle these limitations, we present D-NOVA, a hardware-software co-designed in-storage retrieval accelerator. D-NOVA executes an inverted file (IVF)-based hierarchical retrieval pipeline by deeply embedding the search functionality directly into the NAND memory array. This is achieved by incorporating a new distance metric, Dual-Bound Tight Similarity Sensing (DTS), which is specifically tailored for searching within the NAND string. In addition, we introduce a lightweight contrastive adapter that maps embedding vectors into a DTS-friendly domain, recovering near-software recall while improving performance and energy efficiency. D-NOVA is up to 41.7x faster and 71x more energy-efficient than a CPU baseline, and achieves 12.13x higher throughput while being up to 1.26x more energy-efficient than state-of-the-art in-storage RAG accelerators, demonstrating the potential of fully in-storage vector search for scalable RAG acceleration.

View source

Similar papers

Open access Sep 2026

DCSR-GCN: A High-Performance GCN Accelerator Based on Dynamic Compression and Sparsity Reordering

DCSR-GCN, a high-performance GCN accelerator based on dynamic compression and sparsity reordering based on dynamic compression and sparsity reordering, is presented and a two-phase reordering algorithm that combines conflict-aware row scheduling with reuse-aware column grouping is proposed, thereby mitigating RAW confl...

Jun-Sheng Chang, Yi-Min Zhao, Yu-Xin Huang et al. · 0 citations
Preprint Aug 2026

PLoRA: An NDP-Enhanced Pooled-Memory System for Cost-Efficient Multi-LoRA Serving

Multi-LoRA serving is how one base model becomes thousands of specialized variants, one adapter per user, task, or agent, and the deployments can hold 1000-plus adapters. Serving them is hard because the workload inverts what GPUs provide: terabytes of memory against only tens of TFLOPS, and because every published sys...

Zhong-Kai Yu, O. Venkatachalam, Zheng Wang et al. · 0 citations
Preprint Aug 2026

Chimera: Efficient Multi-Vector Retrieval via GPU-CPU Co-Processing

Chimera is proposed, a GPU-CPU co-processing system for multi-vector retrieval that eliminates this transfer bottleneck, and significantly outperforms existing approaches, achieving up to 59.5x higher QPS at the same recall level.

Yanqi Chen, Jue-Lin Liu, Alexandra Meliou et al. · 0 citations
Book Open access Aug 2026

Replacing NVMe Staging in LLM Inference with a High-Bandwidth CXL Memory Expander with an On-Device DMA Controller

This work evaluates a CXL memory expander equipped with an on-device DMA controller as a high-performance staging tier in GDS-style data paths, sustaining high bandwidth for fine-grained LLM inference workloads such as KV cache updates.

Veerasenareddy Burru, Pradeep Kumar Nalla, Alok Prasad · 0 citations
Preprint Aug 2026

Direct-Operable SIMD Bit-Slicing: A Framework for Memory-Efficient Predicate Evaluation

A novel framework that utilizes the Project Panama Vector API to perform predicate evaluation directly over bit-sliced, compressed data streams by transposing standard row-oriented data into parallel bit-planes to demonstrate a mechanism to evaluate complex filters using SIMD instructions without requiring prior decomp...

A. Mathiyazhagan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.