Preprint
Jul 2026
Attending to Multimodal Generation One Token at a Time
This work introduces multimodal tasks that require explicit switching between visual and textual context within a single response and proposes a simple test-time intervention to boost attention to the relevant modality at the right time, significantly improving multimodal task performance.
Varun Gupta, Vineet Gandhi, Makarand Tapaswi
· 0 citations