Skip to content

FORUM: Frozen Outputs Reconciled Using Model Agreement for Visual Grounding

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

ForUM is presented, a training-free test-time fusion of frozen MLLMs guided by two fixed geometric rules: agreement-based selection keeps the region supported by the most distinct models, and medoid localization returns an actual member box instead of a coordinate average, so one loose prediction cannot shift the answer.

Abstract

Frozen multimodal large language models (MLLMs) now solve standard referring expression comprehension with a single prompted call, yet on adversarial benchmarks with same-category distractors and negation, even the largest models are confidently wrong, and resampling repeats the error. Models built from different data and architectures rarely fall for the same confounder, so their agreement is a strong label-free signal of the correct target. We present FORUM, a training-free test-time fusion of frozen MLLMs guided by two fixed geometric rules: agreement-based selection keeps the region supported by the most distinct models, and medoid localization returns an actual member box instead of a coordinate average, so one loose prediction cannot shift the answer. Fusing three open MLLMs, FORUM surpasses the 397B-parameter published reference by a relative 5% in mean accuracy on the adversarial Ref-Adv-s benchmark, and a plain averaging ensemble by 15%. The gains transfer to standard RefCOCO+, and a balanced lineup with no dominant member still surpasses the 397B model by 5%.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

A Shared Encoder Is Not a Shared Task: Conditional Comparison for Deep Expert Pools

It is shown that cross-evaluated heads on a frozen shared representation inherit the extrapolation confound of shallow exchange scores: pure input rotations with fixed labels inflate a deep exchange score from about 0 to 0.80, while representation-novelty scores are blind in the complementary direction.

Kentaro Oda · 0 citations
Open access Aug 2026

Mitigating Hallucination in Long Referring Expressions via Training-Free, Anchor-Preserved Visual Grounding

Long referring expressions create two coupled sources of hallucination in visual grounding. A detector can select an object that matches only part of the instruction, while a structured vision–language model (VLM) branch can hallucinate a target head or an attribute–object binding. We propose DeRecG, a training-free, a...

Hao-Xuan Song, Li-Huan Shao · 0 citations
Preprint Sep 2026

CrowdCue: Specialist-Cue Conditioning for Vision-Language Crowd Counting

Generative vision-language models (VLMs) offer a counting paradigm in which one model produces both a count and a natural-language account of the scene, yet their raw counting accuracy sits in the range of sub-million-parameter specialist regressors. The open question is whether auxiliary guidance from a pretrained spe...

M. Farazi, B. Ciftler, Abdulhalim Dandoush et al. · 0 citations
#machine learning Preprint Aug 2026

HalluPrism: When Multimodal Uncertainty Should Diagnose, Not Decide

HalluPrism, a behavioral diagnostic that re-runs an answer after visual degradation, blank-image replacement, and grounding or relation checks is proposed, a behavioral diagnostic that separates failure diagnosis from abstention scoring.

Aman Prakash, Sourish Dasgupta, Tanmoy Chakraborty · 0 citations
#artificial intelligence Preprint Aug 2026

Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores

This work introduces a diagnostic protocol using a minimal, target-label-free additive correction, showing that many apparent zero-shot reasoning deficits are expression failures masking intact internal logic, urging a narrower interpretation of benchmark evaluations.

Qi-Yao Yan, Chen-Peng Wang, Liang-Ming Pan · 0 citations
#natural language process... Preprint Sep 2026

Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict

Modality robustness under knowledge conflict is studied across 13 MLLMs and two datasets, and it is found that instability has practical consequences: it degrades performance in multimodal RAG and can be exploited by adversarial attacks.

Jungyeong Lee, Yejin Yoon, Taeuk Kim · 1 citation

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.