The Scalable Tensor-based Codebook Product Quantization for Multi-Label Image Retrieval.
Abstract
Scalable product quantization has recently attracted considerable attention in large-scale image retrieval, as it avoids the need to train multiple models for generating quantization codes of varying lengths. However, most existing approaches primarily concentrate on approximating the ground-truth similarity between image pairs, while overlooking the correlations among sub-codebooks and among codewords. In addition, limited work has addressed the challenge that increasing the number of subspaces or codewords substantially raises memory consumption. To address these limitations, we propose a novel scalable product quantization framework within an end-to-end network, termed Tensor-based Codebook Product Quantization (TCPQ). This work represents an innovative attempt to integrate tensor theory with product quantization methods. The framework leverages tensor-based methods to capture spatial correlations among sub-codebooks and among codewords, and adopts lightweight codebooks for efficiency. For optimization, a subspace-wise unbiased supervised contrastive loss is proposed to bring embeddings of the same class closer together, push embeddings of different classes farther apart within quantization subspaces, and precisely regulate the minimum distance between positive and negative samples. In addition, the ArcFace loss is incorporated to enhance the discriminative power of the learned features, and an orthogonality constraint is imposed on the factor matrix along the codebook dimension to avoid overfitting. Extensive experiments on three large-scale real-world benchmarks demonstrate that TCPQ achieves state-of-the-art retrieval performance.