Jul 2026
One Aligned LLM to Serve Them All: A Transfer Recipe for Training VLMs without Visual-Language Re-Alignment
This work demonstrates that the aligned LLM with a general-purpose vision encoder can effectively enhance downstream VQA performance with task-specific encoders, and investigates several alignment strategies between the aligned LLM and new task-specific encoders.
Jiazuo Yu, Yunzhi Zhuge, Lu Zhang et al.
· International Journal of Com... · 0 citations