Seeing Helps Reasoning in Language Models
Cross-Modal Alignment Regularization (CMAR) is proposed, a method designed to improve LLMs by aligning their internal representations with those of vision models during training by bringing the internal representations of the language and vision models closer together.
Yulu Gan, Kaiya Ivy Zhao, Tomaso A. Poggio et al.
· 0 citations