Automated Caption-Guided Image Retrieval: A Multi-Positive Contrastive Learning Paradigm
Abstract
[Objective] This study aims to improve textual similarity measurement for image retrieval while mitigating the anisotropy of sentence embeddings. [Methods] We propose an image-caption semantic encoder trained with a multi-positive contrastive loss. The conventional contrastive objective is extended to accommodate multiple positive samples, and image captioning is used to generate training data automatically. On this basis, we construct a caption-based image-to-image retrieval framework. [Results] Experiments show that the proposed model outperforms baseline methods on semantic textual similarity (STS) benchmarks and improves the agreement between retrieval results and human semantic judgments. [Limitations] Short captions cannot fully represent the complex semantics, ambiguity, and fine-grained details of an image. [Conclusion] The proposed Multi-Positive Example Contrastive Learning (MPC) model provides more discriminative sentence embeddings and improves semantic similarity measurement in image retrieval.