Skip to content

Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning

Sep 2026 · 0 citations · 151 references
Computer Science

TL;DR

Unified Latent Diffusion Reasoner improves multimodal reasoning by weaving it into a single thread, where the model predicts successive thoughts in a common representation space.

Abstract

Multimodal models increasingly think with different modalities such as images, 3D point clouds, and robot states, not just text. Yet each modality is still encoded into its own representation space, creating a modality-switching gap whenever reasoning moves from one modality to another. In this paper, we introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a framework that unifies different modalities into a shared latent space for multimodal reasoning. A unified encoder maps teacher reasoning steps from different modalities into latent thought tokens in a shared space, trained to extract the information needed for later reasoning steps and the final output. A diffusion reasoner, trained jointly with the encoder, generates these tokens at inference without teacher reasoning steps. Across eleven vision-language model (VLM) benchmarks and two vision-language-action (VLA) suites, Uni-LaDiR achieves relative gains over the strongest baselines of 7.3% on four mathematical and logical VLM benchmarks and 6.1% on RLBench manipulation tasks. Controlled comparisons show increasing gains as more teacher modalities are unified. These results suggest that unification improves multimodal reasoning by weaving it into a single thread, where the model predicts successive thoughts in a common representation space.

View source

Similar papers

#machine learning Preprint Sep 2026

Latent-Aligned Reasoning for Multimodal Recommendation

LARK (Latent-Aligned Reasoning frameworK), a two-stage latent reasoning framework with complementary alignment mechanisms within a single VLM, achieves state-of-the-art performance across multiple recommendation architectures, with controlled ablations confirming the distinct contribution of each component.

Jia-Rui Jin, An-Ya Ji · 0 citations
#computer vision Preprint Aug 2026

Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning

Latent-OPD is proposed, which augments OPD with trajectory-level latent distillation and introduces a progressive teacher-lookahead strategy, which aligns middle-to-late student layers with increasingly deeper teacher layers, establishing Latent-OPD as a highly effective approach to frame-efficient video reasoning.

Aoni Shen, Yongheng Zhang, Ying-Hui Li et al. · 1 citation
Preprint Sep 2026

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces

We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with continuous image generation. The former introduces a visual quantization bottleneck. The...

Hong-Yuan Tao, Xing-Gang Wang, Liang-Hui Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBe...

Hyungjin Chung, Byeongjun Park, Joonseok Lee et al. · 0 citations
Preprint Aug 2026

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

Space Tokens is introduced, a lightweight, architecture-agnostic framework that equips VLMs with explicit continuous spatial representations without requiring additional inference-time modules, and demonstrates that continuous spatial tokens provide an effective, interpretable, and computationally efficient mechanism f...

Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base LLM. While the base LLM retains usable reasoning after scalin...

Zi-Yi Wang, Li Li, Ao-Lin Zhou et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.