Skip to content

GUIDED++: Enhancing Discrimination with Conjunctive Verification for Fine-Grained Open-Vocabulary Object Detection

Aug 2026 · International Journal of Computer Vision · Vol 134 · 0 citations · 35 references

TL;DR

This paper introduces a novel conjunctive multi-attribute verification mechanism to explicitly combat attribute under-representation and significantly outperforms existing methods and establishes a new state-of-the-art on challenging FG-OVD benchmarks, demonstrating a more robust approach to compositional visual reasoning.

View source

Similar papers

Sep 2026

Enhancing target identification and query discrimination for visual grounding

FSG-AID is presented, which integrates fine-grained semantic guidance with an attribute-aware iterative decoder and jointly exploits visual and language features to mine attribute semantics, initialize the target query, and iteratively refine the target representation.

Xiya Bu, Yu Liu, Jizhe Yu et al. · 0 citations
Open access Aug 2026

RoFLIP: Robust and Fine-Grained Alignment for Vision-Language Compositional Reasoning

The Robust and Fine-grained training framework for CLIP-based vision-language models (RoFLIP) is proposed, enhancing both the robustness and granularity of vision-language alignment and underscore RoFLIP’s compositional reasoning and generalization abilities.

Yiwei Sun, Chuan-Bin Liu, Shancheng Fang et al. · 0 citations
#small language model Preprint Aug 2026

OVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained Objects

OVIP-SG is presented, a unified framework for instance-preserving semantic mapping, functional scene partitioning, and language-guided small, fine-grained object retrieval that outperforms ConceptGraphs under a unified evaluation protocol on Replica.

Tianjing Hao, Hai-Yu Lan, Ang Li et al. · 0 citations
Preprint Aug 2026

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.

Hangrui Xu, Zheng-Xian Wu, Yu Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement

To disentangle visual evidence at both semantic and spatial levels, ProtoLIP is introduced, a lightweight prototype-mediated evidence layer that organizes reusable visual prototypes into text-derived semantic families and uses coarse-to-fine evidence routing, where semantic families constrain prototype eligibility and...

Yan Zhu, Yong-Bo Chen, Zheng-Ming Ding et al. · 0 citations
Book Open access Sep 2026

PLAIN: An Explainable Generative Search System Enhanced by Multi-granularity Semantic Alignment

Industrial search platforms must efficiently retrieve relevant items from billions of candidates while satisfying both query relevance and user preferences. Generative Search (GS) has emerged as a transformative paradigm that reformulates traditional indexing and matching as an autoregressive generation task. However,...

Guo-Liang Zhang, Wei-Fan Wang, Jun-Yao Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.