Preprint
Jul 2026
Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis
This work proposes a training framework, VVM-Tuning, to equip LMMs with these capabilities through modality synthesis and modality contexts, and introduces modality contexts in the prompt and use instruction tuning to assist the model in mapping these appearance variations back to modality-related attributes.
Shihao Yuan, Yuanze Li, Ruyi Zhang et al.
· 0 citations