Cross-Modal Comprehensive Multi-Level Granularity Alignment for Vision-Language Pre-Training
Abstract
Vision-Language Pre-training (VLP) aims to convert vision-language tasks into image-text matching challenges, necessitating robust interaction between visual and textual modalities. However, current research often faces a trade-off: methods pursuing fine-grained alignment typically rely on computationally expensive object detectors and pre-defined annotations, while coarse-grained alignment methods overlook subtle visual details. To alleviate these issues, we present OMEGA, a COmprehensive Multi-lEvel Granularity Alignment framework. OMEGA achieves sufficient cross-modal alignment without relying on external object detectors or annotations. Specifically, we first design a Text-Aware Patch Selection (TAPS) module, which dynamically constructs patch-level regions for fine-grained excavation based on informative token selection, effectively bypassing the need for bounding boxes. Secondly, to facilitate multi-grained information fusion, we introduce a Cross-Grained Alignment (CGA) module to learn modality-shared features across global and object levels. Experimental results demonstrate that OMEGA surpasses existing approaches in five pivotal vision-language tasks—image-text retrieval, visual question answering, visual reasoning, visual entailment, and image captioning, while maintaining superior inference efficiency compared to detector-based methods. We believe this study offers a promising object-free paradigm for scalable vision-language pre-training.