Preprint
Aug 2026
Projector Is All You Train
It is found that training only the projector is sufficient to achieve strong multimodal performance relative to existing baseline models and the authors' jointly trained MLLMs with the same encoder and backbone, and that joint training leads to undesirable drift in existing capabilities of the language model.
Nyx Iskandar, Saathvik Selvan, Slater Victoroff
· 0 citations