Skip to content
Conference Open access

Object-Centric Alignment and Anchor Distillation for Weakly Supervised Referring Expression Comprehension

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · pp. 1704-1712 · 0 citations · 39 references

TL;DR

This work proposes OCAAD, an Object-Centric Alignment and Anchor Distillation framework, and introduces two complementary modules that bridges the semantic gap between anchors and objects by transferring object-level knowledge from attention heads to anchors via overlap-aware contrastive learning.

Abstract

Weakly supervised Referring Expression Comprehension (WREC) aims to localize referred objects using only image-text pairs without box-level annotations. Existing one-stage methods predominantly rely on anchor-level alignment, which suffers from two fundamental limitations: (1) anchors represent local visual patches rather than holistic objects, and (2) they lack the capability to model inter-object relations. To address these issues, we propose OCAAD, an Object-Centric Alignment and Anchor Distillation framework. Our key insight is that different self-attention heads in DINOv2 naturally attend to distinct semantic regions, effectively capturing object-level information. Building on this, OCAAD introduces two complementary modules: (1) Anchor-Object Distillation Module (AODM), which bridges the semantic gap between anchors and objects by transferring object-level knowledge from attention heads to anchors via overlap-aware contrastive learning; and (2) Intra-Modal Relation Consistency (IRC), which explicitly models inter-object relations by enforcing the relational structure among linguistic entities to match that of their visual counterparts. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg demonstrate that OCAAD achieves new state-of-the-art performance, validating the effectiveness of objectcentric alignment for WREC. Code is available at https://github.com/VILAN-Lab/OCAAD.

Read PDF

Similar papers

Preprint Aug 2026

LASA: Language-and-Source-Anchored Alignment for Domain Generalized Semantic Segmentation

The Language-and-Source-Anchored Alignment (LASA) framework is proposed, which comprises three synergistic components: Text-and-Source-Guided Style Transfer (TSGST), Domain-Aware Query Adapter (DAQA), and Domain-Aware Decoder Optimizer (DADO).

Jin-Hong Zhu, Wei-Qi Yan, Sheng-Chuan Zhang et al. · 0 citations
Open access Sep 2026

MODdapter: Spatially-aware text embeddings for zero-shot semantic segmentation.

Vision-language models show promise in zero-shot semantic segmentation, but a key challenge is the disconnect between text and visual features. While text embeddings can roughly localize unseen objects, they often lack the fine-grained detail necessary for accurate segmentation, leading to oversegmentation or undersegm...

Jia-Xiang Fang, Shi-Qiang Ma, Jing Wang et al. · 0 citations
Open access 2026

Decoupled global-local collaborative network for visual question answering

Visual Question Answering (VQA) aims to achieve cross-modal semantic understanding through joint modeling of visual content and natural language. Although existing attention-based approaches effectively align features, they struggle to simultaneously accommodate global semantic modeling and local fine-grained perceptio...

Gan-Long Zhou, Dezhi Han, Xiang Shen et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Cost-efficient Active Learning for Referring Image Segmentation and Grounding

This work formulates active learning for VG under the realistic setting where only raw images are available without accompanying text, and introduces Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates.

Junbeom Hong, Seonghoon Yu, Hyungsik Jung et al. · 0 citations
Preprint Sep 2026

FineHOI: Part-Aware Dense Representations for Zero-Shot Human-Object Interaction Detection

This work proposes FineHOI, a zero-shot HOI framework that explicitly models interactions from dense patch-level features, and introduces an Adaptive Part-Level Attention module that decomposes humans and objects into semantically coherent parts via unsupervised clustering, and re-weights them based on their interactio...

Francesco Tonini, Lorenzo Vaquero, Mohammad Mahdi Derakhshani et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.