Skip to content

Author

Libin Huang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access 2026

BiFGA: A Bidirectional Image–Text-Guided Fine-Grained Alignment Method for Remote Sensing Image–Text Retrieval

Remote sensing image–text retrieval matches remote sensing images with textual descriptions. However, dense objects, complex spatial layouts, and large-scale variations make fine-grained cross-modal alignment challenging. Existing methods mainly rely on global representations or direct patch-token similarities. Global matching may overlook small objects and spatial details, whereas local matching is susceptible to background clutter, scale variations, and redundant tokens. To address these limitations, this article proposes BiFGA, a bidirectional fine-grained alignment framework that uses cross-modal similarity responses as semantic anchors. Its bidirectionally guided fine-grained matching module reconstructs text-guided visual relations and vision-guided textual relations, enabling second-order interactions between image regions and textual units. BiFGA also introduces a hybrid window spatial-channel attention module to enhance visual regions and a channel-splitting phrase-level feature extraction module to model multiscale textual semantics. Global similarity and local fine-grained matching are integrated through coarse-to-fine scoring. Experiments on RSITMD, RSICD, UCM-Caption, and Sydney-Captions demonstrate the effectiveness and stability of BiFGA. Averaged over five independent runs, BiFGA achieves mR scores of 50.07%, 35.15%, 59.42%, and 53.19%, respectively. It obtains the best average mR among the compared methods on RSITMD, RSICD, and Sydney-Captions, while remaining comparable to the strongest method on UCM-Caption.

Yun Liao, Yong Liu, Junhui Liu et al. · 0 citations