Preprint
Jul 2026
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI
MonkeyOCRv2, a visual-text pretrained model for document AI, is presented, and a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction is proposed: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details.
Yuliang Liu, Zhang Li, Ziyang Zhang et al.
· 1 citation