Skip to content
Preprint

FAU at ImageCLEF 2026 Task on Multimodal Reasoning Robust Candidate Scoring and Concise Multilingual Visual Answering

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

The central finding is that robust output control is as important as model choice for robust reasoning over multilingual educational and scientific images with dense text, diagrams, charts, tables, formulas, and units.

Abstract

We present our ImageCLEF 2026 Multimodal Reasoning system for the Visual Multiple Choice Question Answering (Visual MCQ) and Visual Open Question Answering (Visual OpenQA) subtasks. The challenge requires reliable reasoning over multilingual educational and scientific images with dense text, diagrams, charts, tables, formulas, and units, while enforcing strict answer formats. Our central finding is that robust output control is as important as model choice. For Visual MCQ, we replace fragile free-form generation with direct candidate label scoring from vision-language model logits, then combine complementary runs through score fusion and voting. For Visual OpenQA, we use image enhancement, concise final answer prompting, deterministic decoding, and targeted post-processing to remove reasoning traces and formatting artifacts. Without task-specific model training, our official submissions achieved third place in Visual MCQ with 0.7108 accuracy and first place in Visual OpenQA with 0.6488 COMET, 0.1391 BLEU, 0.2762 ROUGE L, and 0.2383 METEOR. The results highlight the practical value of inference engineering: careful scoring, ensembling, prompting, and cleanup can turn strong VLMs into reliable competition systems.

View source

Similar papers

Book Open access Aug 2026

SciChart: Visual Question Answering and Reasoning for Scientific Spectral Chart

This work designs two tasks, basic question answering (BasicQA) and reasoning-based question answering (ReaQA), to evaluate the models' ability to directly extract information from charts, and understand the textual and visual information for reasoning.

Tan Yue, Rui Mao, Xuzhao Shi et al. · 1 citation
#computer vision Preprint Aug 2026

ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

ReVA is proposed, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space.

A. Senthil · 0 citations
Open access Aug 2026

Dual Text-guided Cross-attention for Global–local Visual Fusion in Vietnamese Visual Question Answering

This study proposes a dual-stream architecture that exploits complementary Transformer-based and convolutional visual representations for Vietnamese VQA, and shows improvements over the corresponding single-stream convolutional baselines, supporting the complementary role of the two visual representations.

Huy Tran, V. Nguyen · 0 citations
#artificial intelligence Preprint Sep 2026

VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics

Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to compr...

Bo Lv, Mao Zheng, Zheng Li et al. · 0 citations
Aug 2026

Cross-modal alignment enhancement for lightweight large vision language models

A Low-Complexity Cross-Modal Alignment via Projection (LCAP) network is proposed, which introduces Projective Token Compression (PTC), which leverages Mish activation and adaptive average pooling to reduce feature redundancy while enhancing discriminative information, and Positional Spatial Enhancement (PSE), which exp...

Yu-Chen Sha, Lingli Wan, Ge Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.