Improving Scene Text Recognition in Multimodal Large Language Models using Visual Text Grounding
This work introduces three tasks/objectives for reverse localization of text as an instruction-tuning mechanism, where the model is guided to extract textual content based on spatial localization cues, thereby enhancing its spatial grounding ability.
Shashank Krishna Vempati, Chetan Arora
· 0 citations