Skip to content

Detection-Guided Attention Steering for Vision Language Models

· 0 citations · 39 references

TL;DR

A detection-guided dynamic attention steering system that leverages the locality insight from CNNs to efficiently steer a VLM’s attention toward more relevant sections of an image and demonstrates the effectiveness of combining different model architectures to harness their respective strengths for advancing VLM capabilities.

View source

Similar papers

Preprint Sep 2026

ConvCue: Complementary Visual Inductive Biases for Vision-Language Models

Modern vision-language models (VLMs) achieve strong performance across a broad range of multimodal tasks, yet still struggle with visual questions that require fine-grained discrimination and spatial understanding. These limitations motivate investigating whether supplementary visual representations can improve existin...

Zi-Xuan Lan, Shi-Chu Sun · 0 citations
Preprint Open access Sep 2026

Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs

Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining distribution, yet fail on specialized domains where the required discriminative features were never learned, a regime we term distant out-of-distribution (OOD). Standard ad...

Hung-Jen Chen, Yuek F. Ho, Ting-Yao Huang et al. · 0 citations
Preprint Aug 2026

UC-VLM: Consistency-Driven Learning for AI-Generated Image Detection with Vision-Language Large Models

UC-VLM is a unified multi-stage binary-supervised framework that consistently reuses the same authenticity labels for visual adaptation and label-conditioned text generation, while leveraging automatically optimized instructions to reduce prompt sensitivity without requiring human-written rationales or hand-crafted pro...

Lei Tan, Shuwei Li, Mohan S. Kankanhalli et al. · 0 citations
#machine learning Preprint Sep 2026

Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis

It is found that post-pretraining QK-RMSNorm injection fails to reproduce the protection of native QK-RMSNorm, while several off-the-shelf weight-merging settings fail to recover the lost capability after VL training, highlighting the value of screening backbones with Sink Strength before VL training and narrow the int...

Minsik Choi, Geewook Kim, Young Geun Kim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.