Skip to content

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

Jul 2026 · arXiv.org · Vol abs/2607.14660 · 0 citations · 38 references
Computer Science

TL;DR

VIABench is introduced, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves, and proposes a rigorous benchmarking pipeline that supports both online (real-time) and offline settings.

Abstract

Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.

View source

Similar papers

Preprint Aug 2026

HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

HUG-VIS, a unified benchmark for Human-centered Understanding and Generation in Visual Intelligence, contains 8,400 seated half-body videos of 30 professional actors, each performing the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alp...

Fei Ma, Ze-Bang Cheng, Ming-Hui Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Touvigation: Embodied Adaptive Object Acquisition for Blind and Low-Vision Users in Unfamiliar Indoor Environments

Blind and low-vision users often face challenges when locating and physically acquiring objects in unfamiliar indoor environments. Existing vision-language-model-based assistants can provide semantic descriptions but may introduce latency, hallucinations, and guidance that is poorly aligned with embodied action. We pre...

George Xi Wang, Xiang-Yu Li, Shaoyue Wen et al. · 0 citations
Conference Aug 2026

VizAdapt: novel dataset development with a voice-interactive visual question answering system for blind and low vision individuals

Blind and Low Vision Individuals (BLV) encounter significant difficulties in comprehending complex visual environments, while current assistive technologies typically lack speech-interactive reasoning capabilities. To address this gap, this study fine-tunes the Qwen2.5-Omni framework using Llamafactory based on our enh...

Li-Ting Chen, Ping-Ting Lin, Wei-Wei Li et al. · 0 citations
Preprint Aug 2026

OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

The OmniEdit-Bench provides a comprehensive and reliable testbed for evaluating instruction-based video editing and offers insights into future research directions, including an accuracy-aware penalty mechanism that conditions other scores on accuracy, preventing visually plausible but incorrect edits from receiving in...

Chenxuan Miao, Yutong Feng, Yi Lu et al. · 1 citation
#computer vision Preprint Sep 2026

Small yet Assistive: Spatially-Aware Post-Training for Low Vision

An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate high-quality audio-description-compliant narrations but cannot run on mobile devices; small VLMs...

Rishabh C. Choudhary, S. Raj, Umesh Goyal et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Assisted Spatial Cognition Through Vision-Language Models

Multimodal AI, powered by Large Language Models (LLMs) and Vision-Language Models (VLMs), is transforming assistive technologies by enabling simultaneous processing of visual and textual data. This advancement holds significant promise for over 43 million visually impaired and neuro-divergent individuals worldwide who...

Hamza Riaz, Jaime B. Fernandez, Ian Mills et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.