Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· pp. 1704-1712· 0 citations· 39 references
TL;DR
This work proposes OCAAD, an Object-Centric Alignment and Anchor Distillation framework, and introduces two complementary modules that bridges the semantic gap between anchors and objects by transferring object-level knowledge from attention heads to anchors via overlap-aware contrastive learning.
Abstract
Weakly supervised Referring Expression Comprehension (WREC) aims to localize referred objects
using only image-text pairs without box-level annotations. Existing one-stage methods predominantly
rely on anchor-level alignment, which suffers from
two fundamental limitations: (1) anchors represent
local visual patches rather than holistic objects, and
(2) they lack the capability to model inter-object
relations. To address these issues, we propose
OCAAD, an Object-Centric Alignment and Anchor Distillation framework. Our key insight is that
different self-attention heads in DINOv2 naturally
attend to distinct semantic regions, effectively capturing object-level information. Building on this,
OCAAD introduces two complementary modules:
(1) Anchor-Object Distillation Module (AODM),
which bridges the semantic gap between anchors
and objects by transferring object-level knowledge
from attention heads to anchors via overlap-aware
contrastive learning; and (2) Intra-Modal Relation Consistency (IRC), which explicitly models
inter-object relations by enforcing the relational
structure among linguistic entities to match that of
their visual counterparts. Extensive experiments on
RefCOCO, RefCOCO+, and RefCOCOg demonstrate that OCAAD achieves new state-of-the-art
performance, validating the effectiveness of objectcentric alignment for WREC. Code is available at
https://github.com/VILAN-Lab/OCAAD.
The Language-and-Source-Anchored Alignment (LASA) framework is proposed, which comprises three synergistic components: Text-and-Source-Guided Style Transfer (TSGST), Domain-Aware Query Adapter (DAQA), and Domain-Aware Decoder Optimizer (DADO).
Jin-Hong Zhu, Wei-Qi Yan, Sheng-Chuan Zhang et al.· 0 citations
Vision-language models show promise in zero-shot semantic segmentation, but a key challenge is the disconnect between text and visual features. While text embeddings can roughly localize unseen objects, they often lack the fine-grained detail necessary for accurate segmentation, leading to oversegmentation or undersegm...
Jia-Xiang Fang, Shi-Qiang Ma, Jing Wang et al.· Neural Networks· 0 citations
Visual Question Answering (VQA) aims to achieve cross-modal semantic understanding through joint modeling of visual content and natural language. Although existing attention-based approaches effectively align features, they struggle to simultaneously accommodate global semantic modeling and local fine-grained perceptio...
Gan-Long Zhou, Dezhi Han, Xiang Shen et al.· Computer Science and Informa...· 0 citations
This work formulates active learning for VG under the realistic setting where only raw images are available without accompanying text, and introduces Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates.
Junbeom Hong, Seonghoon Yu, Hyungsik Jung et al.· 0 citations
Inspired by the linguistic phenomenon of code-switching, MMCS interleaves vision and language by replacing textual entities with their corresponding visual objects, enforcing local vision-language grounding.
Changhao Xiang, Shangyu Xing, Zhen Wu et al.· 0 citations
This work proposes FineHOI, a zero-shot HOI framework that explicitly models interactions from dense patch-level features, and introduces an Adaptive Part-Level Attention module that decomposes humans and objects into semantically coherent parts via unsupervised clustering, and re-weights them based on their interactio...
Francesco Tonini, Lorenzo Vaquero, Mohammad Mahdi Derakhshani et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.