Skip to content
Review

A Survey of Multimodal Models on Language and Vision: A Unified Modeling Perspective

· 2 citations · 265 references

TL;DR

This survey investigates the current research landscape of multimodality modeling from three perspectives: the first group of multimodal models adopts a heterogeneous architecture to bridge different modality data, the second leverages LLM for multimodality modeling via a unified language modeling objective, and the third represents multimodal data entirely within a single visual representation.

View source

Similar papers

Review Open access Jul 2026

Multimodal Video Understanding: A Capability-Based Survey of Alignment, Expression, and Reasoning

Multimodal video understanding (MVU) has emerged as a fast-growing research frontier, driven by major advances in video-language pre-training and large multimodal models over the past decade. MVU aims to synergistically integrate visual, audio and textual modalities to interpret complex video semantics, supporting widespread downstream tasks including cross-modal retrieval, dense captioning, video question answering, event analysis and intelligent assistance. Despite the rapid proliferation of specialized MVU models, the community still lacks a unified capability-centric framework to systematically clarify the hierarchical competency architecture and evolutionary trajectory of state-of-the-art approaches. To address this issue, this paper presents a structured, comprehensive survey of the latest MVU progress, establishing a novel three-tier taxonomy that categorizes existing studies into cross-modal alignment, multi-granularity semantic expression and multimodal reasoning. Along this pipeline, we further systematically synthesize core modality fusion strategies, mainstream benchmark datasets and standardized evaluation protocols. Through a fine-grained analysis of representative published results, we highlight the critical impact of inconsistent evaluation settings, cross-experiment comparability bottlenecks and inherent methodological trade-offs between performance and efficiency. Finally, we identify and dissect three key open challenges: ultra-long video scalability, performance degradation from modality noise and missing data, and factual reliability risks in generative MVU systems. This capability-oriented systematic reference clarifies the methodological evolution logic of MVU, and provides actionable guidance for developing next-generation robust, high-performance multimodal video understanding systems.

Rongyong Zhao, Da Pu, Cuiling Li et al. · 0 citations
Preprint Aug 2026

A Model-Internal Protocol for Assessing Multimodal Models as Integrated Systems

As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs). In this work, we propose Self-Generative-Understanding (SGU), a novel, annotation-free evaluation framework that probes the integrated capabilities of unified models through a semantic closed-loop challenge. Without requiring new annotations, SGU leverages the dual understanding-and-generation abilities of UMMs by asking them to first perceive an image and produce a textual description, subsequently reconstruct a visual context based on that description, and finally perform reasoning over the self-generated output. This pipeline provides a zero-cost testbed that yields an integrated performance score specifically tailored for evaluating UMMs as unified systems. Extensive experiments show that even high-performing UMMs often struggle to reason over their own generated contexts, revealing limitations that are not captured by separate evaluations of understanding or generation alone. Our work provides a complementary holistic evaluation framework and offers a foundation for benchmarking the development of next-generation unified multimodal models.

Hao Zhang, Jiaxin Qi, Zhijiang Tang et al. · 0 citations
Preprint Jul 2026

Vision as Unified Multimodal Generation

Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry.

Xiaoyang Han, Jianhua Li, Kewang Deng et al. · 1 citation
Review Open access Jul 2026

A Review of Multimodal Large Language Models: Fusion Mechanisms and Capability Evolution

This paper reviews the principal technical paradigms of multimodal fusion, including early fusion, intermediate fusion, late fusion, and hybrid fusion, and compares the structural characteristics and applicable scenarios of different fusion approaches and explores the development of MLLMs from the perspectives of vision-language understanding, multimodal content generation, multimodal interaction and agent-oriented tasks, as well as domain-specific applications.

Zibo Xu, Shuqiang Gao · 0 citations
Preprint Jul 2026

Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering

It is shown that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness, and establish cross-branch steering as a practical tool for probing multimodal representations.

Yu Wang, Sharon Li · 0 citations
Preprint Jul 2026

MIRROR: Learning from the Other View for Multi-Modal Reasoning

Modality-Informed Reciprocal Reasoning Optimization (MIRROR), a reinforcement learning approach for improving multimodal reasoning via self supervision, is developed and improves over standard RL and yields more accurate and consistent behavior across modalities.

Wen Ye, Yuxiao Qu, Aviral Kumar et al. · 0 citations