Skip to content
Preprint

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

Aug 2026 · 0 citations · 155 references
Computer Science

TL;DR

This study provides empirical clarity through a systematic exploration of multimodal pretraining and derives efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget.

Abstract

Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity through a systematic exploration of multimodal pretraining. Our controlled experiments on both synthetic and large-scale real-world datasets yield four key insights into the physics of multimodal pretraining: (i) Knowledge Flow: We disentangle how language, visual understanding, and visual generation transfer knowledge across modalities, revealing distinct patterns of influence and asymmetry; (ii) Synergy vs. Competition: We show that data"complexity"largely determines whether modalities are synergistic, identify architectural choices that promote synergy: such as shared attention and normalization with modality-specific feed-forward layers, and find that these behaviors generalize across different visual tokenizer designs; (iii) Early Unification: Unifying modalities from the very early stages and training them jointly is shown to be more effective than late alignment or sequential training. This process uncovers a vision laziness phenomenon, where delayed integration leads models to rely on language priors; (iv) Recipes: We derive efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget. These core findings are subsequently validated at scale by training multiple 13.5B MoE models on 2T tokens. We hope this study provides a principled foundation for understanding and scaling multimodal pretraining.

View source

Similar papers

Preprint Jul 2026

Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering

It is shown that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness, and establish cross-branch steering as a practical tool for probing multimodal representations.

Yu Wang, Sharon Li · 0 citations

Multimodal Concept Mapping

This work investigates how MLLMs learn novel concepts by introducing concepts in unimodal pre-training or multimodal fine-tuning, and evaluates the model’s ability to generalize between the two settings, and evaluates how deeply a model maps concepts across modalities.

S. Boppana, Tian Yun, Carina Curto et al. · 0 citations
Preprint Jul 2026

Scaling Native Multimodal Pre-Training From Scratch

This empirical research establishes the essential groundwork for predictably scaling multimodal foundation models by modeling the influence of data composition on compute laws and allocation exponents and derive an efficiency frontier specifying precise configurations of model size, token count, and data mixture.

Haoyuan Wu, Aoqi Wu, Hai Wang et al. · 1 citation
Preprint Jul 2026

Monkey King Bang: A Unified Scientific Multimodal Foundation Model

Experiments show that MKB achieves competitive scientific understanding across biological and molecular benchmarks, produces high-fidelity native outputs for weather forecasting, biological generation, and medical-image segmentation, and largely retains the general capabilities of its Qwen3-VL backbone.

Hesen Chen, Xinyue Su, Xiaomeng Yang et al. · 0 citations
Preprint Jul 2026

Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis

This work proposes a training framework, VVM-Tuning, to equip LMMs with these capabilities through modality synthesis and modality contexts, and introduces modality contexts in the prompt and use instruction tuning to assist the model in mapping these appearance variations back to modality-related attributes.

Shihao Yuan, Yuanze Li, Ruyi Zhang et al. · 0 citations

One Flow Fits All! A Scale-Aware Generative Framework for Diverse Data

The Inverse Heat Mean Flow is introduced, a general-purpose solver that is compatible with a wide range of model backbones that directly learns an average velocity field through an inverse heat formul ation, thereby simplifying trajectory learning and enabling adaptive topological alignment.

Hubin Cao, Jun Ma, Yusupu Ainiwaer et al. · 0 citations