Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-c...
Xin Chen, An-An Du, Feng-Juan Feng et al.· 0 citations
As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environ...
Yu Liu, Zhi-Lin Liu, Zhi-Wei Yang et al.· 0 citations
This work proposes NTEP-R (NTEP Reward), a supervision mechanism ensuring that each tool invocation strictly advances the reasoning process toward the final solution, and introduces a non-repeated-goal regularizer to penalize redundant calls that revisit satisfied NTEP goals.
Xing-Ming Long, Yu Liu, Zhi-Wei Yang et al.· 0 citations
Switch-Reasoner is proposed, a GRPO-based framework that learns to adaptively select reasoning modes for MLLMs and introduces a dual-level regulation mechanism that balances the overall use of Thinking Mode and Direct Mode while providing sample-level supervision based on the relative benefit of the two choices.
Yiyang Fang, Pei Fu, Jinjie Li et al.· arXiv.org· 0 citations
DeltaV is proposed, a ULMM that replaces full-image generation with visual updates and introduces a temporal similarity (TSIM) Router, which stops allocating tokens once the marginal reconstruction gain falls below a threshold.
Pengjie Wang, Linger Deng, Zujian Zhang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.