Chimera is proposed, a GPU-CPU co-processing system for multi-vector retrieval that eliminates this transfer bottleneck, and significantly outperforms existing approaches, achieving up to 59.5x higher QPS at the same recall level.
Abstract
Multi-vector retrieval has become an important primitive for fine-grained matching in information retrieval, with emerging applications in areas such as recommender systems and bioinformatics. However, its high computational complexity and memory costs make low-latency retrieval difficult. Prior systems have attempted to optimize query latency, but their designs remain CPU-centric. While GPUs offer substantial computational advantages, their limited memory capacity necessitates a heterogeneous architecture in which the dataset resides in host memory and the GPU serves as an accelerator. Existing GPU-based system, PLAID, is bottlenecked by CPU-GPU data movement, as vector data must be transferred from host memory to the GPU at query time. We propose Chimera, a GPU-CPU co-processing system for multi-vector retrieval that eliminates this transfer bottleneck. Chimera stores highly compressed, low-precision quantization codes on the GPU while maintaining high-precision data in CPU memory. At query time, it leverages GPU-resident data for efficient candidate generation and filtering, and further refines results through a GPU-CPU collaborative scoring scheme that completely avoids vector data transfer while enabling computation overlap. Experiments on real-world datasets demonstrate that Chimera significantly outperforms existing approaches, achieving up to 59.5x higher QPS at the same recall level.
Algorithms for processing large-scale spatial datasets are of significant interest in both scientific research and industrial applications. The efficient implementation of such algorithms is crucial for modern data-intensive systems, and GPU-based parallel processing has emerged as a particularly effective approach for...
Ioannis Pateras, Polychronis Velentzas, M. Vassilakopoulos et al.· ISPRS International Journal...· 0 citations
Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user intent, and broad coverage of a massive inventory under fixed latency and resource budgets. We characterize this as the...
Hao-Hao Fu, Ji-Chao Sun, Bai-Ting Zhu et al.· 0 citations
Tiny-Pipe comprises a holistic layer packing method that simultaneously reduces GPU memory footprint and improves training performance, an active CPU memory management that alleviates CPU memory pressure by eliminating redundant parameters, and a layer-wise runtime swapping strategy that further enhances overall perfor...
Yu-Quan Ding, Jie Shao· ACM Transactions on Architec...· 0 citations
Similarity search over high-dimensional sparse feature vectors is a fundamental problem in applications such as information retrieval, data mining, and machine learning. In this paper, we present PreFast, a scalable parallel prefix-filtering framework for top-k similarity search over non-negative weighted feature vecto...
S. Pokharel, A. Subedi, Elizabeth Oluwadamilola Durowoju et al.· Proceedings of the Internati...· 0 citations
This paper introduces GEM-KMeans, a spectrally normalized yet mathematically equivalent NLR formulation that fuses the gradient update, nonnegative projection, and sufficient statistics for normalization and iterate movement into a matrix-multiplication epilogue.
Peng Xu, Nihar Koganti, Volodymyr Kindratenko et al.· 0 citations
Encrypted AI using fully homomorphic encryption (FHE) enables inference directly over encrypted queries, providing strong privacy guarantees. However, its computational and memory overheads have limited practical deployment. Custom FHE accelerators improve performance, but rely on advanced manufacturing technologies th...
Siddharth Jayashankar, Joshua Kim, Michael B. Sullivan et al.· Proceedings of the ACM SIGOP...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.