Preprint
Aug 2026
MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment
Inspired by the linguistic phenomenon of code-switching, MMCS interleaves vision and language by replacing textual entities with their corresponding visual objects, enforcing local vision-language grounding.
Changhao Xiang, Shangyu Xing, Zhen Wu et al.
· 0 citations