Skip to content
Preprint

LLaVA-Assessor: Building the Foundation LMM For Visual Quality Assessment

Aug 2026 · 0 citations · 84 references
Computer Science

TL;DR

This work proposes LLaVA-Assessor, a unified data construction and model training system for LMM-based machine vision, and introduces a simple yet effective prompt disentanglement strategy to alleviate training-objective confusion in multi-task learning, thereby enabling stable and coherent joint training.

Abstract

Aligning with the human visual system~(HVS) in perceiving and evaluating the quality of visual signals is a central objective of machine-vision-based visual quality assessment systems. With the rapid progress of large multi-modal models~(LMMs), visual question answering provides a promising paradigm for building unified foundation models for visual quality assessment under multi-modal and multi-task scenarios. Inspired by the classical ``perception-decision"process in HVS-based quality evaluation, we formulate visual quality assessment for LMM-based machine vision as two complementary tasks: ``quality interpretation''and ``quality scoring". Centered on these objectives, we propose LLaVA-Assessor, a unified data construction and model training system. To support multi-modal inputs, we design an adaptive model architecture that enables efficient processing of both images and videos. For data construction, we develop rigorous human annotation protocols and a novel machine-synthesis-dominated data expansion pipeline to build a large-scale and high-quality datasets. Furthermore, we introduce a simple yet effective prompt disentanglement strategy to alleviate training-objective confusion in multi-task learning, thereby enabling stable and coherent joint training. The resulting all-in-one LMM LLaVA-Assessor-GIGA achieves superior performance on $11$ image/video quality scoring test sets and 4 visual quality interpretation benchmarks. Extensive results demonstrate the effectiveness of integrating structured data construction, adaptive model design, and multi-task joint training for automated visual quality assessment. Our work provides compelling insights for developing foundation LMMs for automatic visual quality assessment. Project page at https://github.com/jzhws/LLaVA-Assessor.

View source

Similar papers

Preprint Aug 2026

SciQNet: Two-Stage Multimodal Adaptation for Scientific Image Quality Assessment

SciQNet is presented, a two-stage multimodal adaptation framework for scientific image quality assessment and experiments show that using a 40% stratified subset of the domain-adaptive data gives the best performance among the evaluated pretraining fractions, suggesting that pretraining-data relevance may be as importa...

Yin-Loon Khor, Y. Wong, Jing-Jie Tan et al. · 0 citations
Open access Aug 2026

Structured Creative Evaluation for Text-to-Image Generative AI Models

In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...

W.-C. Ma, Q. Zhang · 0 citations
#artificial intelligence Preprint Sep 2026

Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition

Vision-Language Models (VLMs) have recently been proposed as promising tools for face recognition, as they can produce natural language explanations alongside similarity scores. This capability is considered appealing for face comparisons in forensic contexts, which require decisions to be transparent and auditable. Ho...

Laurent Colbois, Sébastien Marcel · 0 citations
#artificial intelligence Preprint Sep 2026

An Evolutionary Agentic Approach for Open-ended Image Quality Perception

Generative models are rapidly expanding image quality assessment (IQA) beyond traditional fidelity factors to emerging dimensions such as physical plausibility and text-rendering correctness. However, existing IQA models rely on fixed definitions and heavy supervision, making them difficult to extend to open-ended perc...

Zhen-Chen Tang, Bo Peng, Zi-Chuan Wang et al. · 1 citation
#computer vision Preprint Aug 2026

LookBack: Where and How to Score LVLM Responses via Visual Reference Usage

LookBack, a training-free LVLM response scoring method that augments token likelihood with visual lookback score, a lightweight measure of how strongly each response token refers to image tokens, consistently improves Best-of-$N$ selection over existing baselines with negligible additional overhead.

Beomsik Cho, Jin-Ha Kim, Dongseok Lee et al. · 0 citations
Preprint Aug 2026

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model, is proposed, endowing models with native fine-grained region description and flexible reasoning capabilities.

Chang-Jiang Jiang, Qiannian Zhao, Lei Xin et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.