Skip to content
Preprint

Chimera: Efficient Multi-Vector Retrieval via GPU-CPU Co-Processing

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

Chimera is proposed, a GPU-CPU co-processing system for multi-vector retrieval that eliminates this transfer bottleneck, and significantly outperforms existing approaches, achieving up to 59.5x higher QPS at the same recall level.

Abstract

Multi-vector retrieval has become an important primitive for fine-grained matching in information retrieval, with emerging applications in areas such as recommender systems and bioinformatics. However, its high computational complexity and memory costs make low-latency retrieval difficult. Prior systems have attempted to optimize query latency, but their designs remain CPU-centric. While GPUs offer substantial computational advantages, their limited memory capacity necessitates a heterogeneous architecture in which the dataset resides in host memory and the GPU serves as an accelerator. Existing GPU-based system, PLAID, is bottlenecked by CPU-GPU data movement, as vector data must be transferred from host memory to the GPU at query time. We propose Chimera, a GPU-CPU co-processing system for multi-vector retrieval that eliminates this transfer bottleneck. Chimera stores highly compressed, low-precision quantization codes on the GPU while maintaining high-precision data in CPU memory. At query time, it leverages GPU-resident data for efficient candidate generation and filtering, and further refines results through a GPU-CPU collaborative scoring scheme that completely avoids vector data transfer while enabling computation overlap. Experiments on real-world datasets demonstrate that Chimera significantly outperforms existing approaches, achieving up to 59.5x higher QPS at the same recall level.

View source

Similar papers

Open access Sep 2026

GPU-Based Algorithms for Processing the k-CP Query on Spatial Data

Algorithms for processing large-scale spatial datasets are of significant interest in both scientific research and industrial applications. The efficient implementation of such algorithms is crucial for modern data-intensive systems, and GPU-based parallel processing has emerged as a particularly effective approach for...

Ioannis Pateras, Polychronis Velentzas, M. Vassilakopoulos et al. · 0 citations
#machine learning Preprint Sep 2026

Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale

Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user intent, and broad coverage of a massive inventory under fixed latency and resource budgets. We characterize this as the...

Hao-Hao Fu, Ji-Chao Sun, Bai-Ting Zhu et al. · 0 citations
Open access Aug 2026

GPU and CPU Memory Co-Optimization in Heterogeneous Pipeline Parallelism for Efficient Large Language Model Fine-Tuning on Commodity Servers

Tiny-Pipe comprises a holistic layer packing method that simultaneously reduces GPU memory footprint and improves training performance, an active CPU memory management that alleviates CPU memory pressure by eliminating redundant parameters, and a layer-wise runtime swapping strategy that further enhances overall perfor...

Yu-Quan Ding, Jie Shao · 0 citations
Book Open access Sep 2026

Parallel Prefix Filter Search for Weighted Jaccard Similarity

Similarity search over high-dimensional sparse feature vectors is a fundamental problem in applications such as information retrieval, data mining, and machine learning. In this paper, we present PreFast, a scalable parallel prefix-filtering framework for top-k similarity search over non-negative weighted feature vecto...

S. Pokharel, A. Subedi, Elizabeth Oluwadamilola Durowoju et al. · 0 citations
#machine learning Preprint Sep 2026

GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization

This paper introduces GEM-KMeans, a spectrally normalized yet mathematically equivalent NLR formulation that fuses the gradient update, nonnegative projection, and sufficient statistics for normalization and iterate movement into a matrix-multiplication epilogue.

Peng Xu, Nihar Koganti, Volodymyr Kindratenko et al. · 0 citations
Book Open access Sep 2026

Cerium: A Multi-GPU Framework for Terabyte-Scale Encrypted Inference

Encrypted AI using fully homomorphic encryption (FHE) enables inference directly over encrypted queries, providing strong privacy guarantees. However, its computational and memory overheads have limited practical deployment. Custom FHE accelerators improve performance, but rely on advanced manufacturing technologies th...

Siddharth Jayashankar, Joshua Kim, Michael B. Sullivan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.