Skip to content
Review Open access

Research on Cross-modal Image Retrieval and Image-Text Matching Based on Visual-Language Pre-trained Models

2026 · Academic Journal of Engineering and Technology Science · 0 citations · 5 references

Abstract

: Cross-modal image retrieval and image-text matching are core tasks connecting computer vision and natural language processing, aiming to eliminate the heterogeneous gap between visual and text modalities. Visual-language pre-trained models, through large-scale data learning and cross-modal alignment, have become the dominant technical paradigm for solving this task. This paper systematically reviews the development of visual-language pre-trained models in the field of cross-modal retrieval, classifies and analyzes existing methods from three dimensions: model architecture, pre-training objectives, and downstream adaptation, and focuses on discussing the architectural differences between dual encoders and fusion encoders, the design evolution of pre-training tasks, and adaptation techniques such as efficient parameter fine-tuning. Based on this, we summarize mainstream datasets and evaluation metrics, compare the performance of representative models, and deeply analyze three key challenges: fine-grained alignment, noise robustness, and inference efficiency. Finally, we look forward to future research directions such as few-shot generalization, unified multi-task framework, and interpretability, hoping to provide a reference for further research in this field.

Read PDF