Review
A Survey of Multimodal Models on Language and Vision: A Unified Modeling Perspective
This survey investigates the current research landscape of multimodality modeling from three perspectives: the first group of multimodal models adopts a heterogeneous architecture to bridge different modality data, the second leverages LLM for multimodality modeling via a unified language modeling objective, and the third represents multimodal data entirely within a single visual representation.
Zhongfen Deng, Yibo Wang, Yueqing Liang et al.
· 2 citations