Preprint
Sep 2026
Distilling Visual Reasoning into Text Space
V2T first trains a teacher LVLM using interleaved visual and textual chains of thought, and then uses knowledge distillation to train a student LVLM using the teacher's logits and cross-entropy supervision from ground-truth textual reasoning.
Wen-Han Yang, Nilay Naharas, Ali Payani et al.
· 0 citations