Open access
Aug 2026
RoFLIP: Robust and Fine-Grained Alignment for Vision-Language Compositional Reasoning
The Robust and Fine-grained training framework for CLIP-based vision-language models (RoFLIP) is proposed, enhancing both the robustness and granularity of vision-language alignment and underscore RoFLIP’s compositional reasoning and generalization abilities.
Yiwei Sun, Chuanbin Liu, Shancheng Fang et al.
· International Journal of Com... · 0 citations