Skip to content

Author

Shao-Jie Zhang

We have 5 of 9 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Xiaomi-OCR-0 Technical Report

Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-c...

Xin Chen, An-An Du, Feng-Juan Feng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environ...

Yu Liu, Zhi-Lin Liu, Zhi-Wei Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

This work proposes NTEP-R (NTEP Reward), a supervision mechanism ensuring that each tool invocation strictly advances the reasoning process toward the final solution, and introduces a non-repeated-goal regularizer to penalize redundant calls that revisit satisfied NTEP goals.

Xing-Ming Long, Yu Liu, Zhi-Wei Yang et al. · 0 citations
Jul 2026

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning

Switch-Reasoner is proposed, a GRPO-based framework that learns to adaptively select reasoning modes for MLLMs and introduces a dual-level regulation mechanism that balances the overall use of Thinking Mode and Direct Mode while providing sample-level supervision based on the relative benefit of the two choices.

Yiyang Fang, Pei Fu, Jinjie Li et al. · 0 citations
Preprint Jul 2026

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models

DeltaV is proposed, a ULMM that replaces full-image generation with visual updates and introduces a temporal similarity (TSIM) Router, which stops allocating tokens once the marginal reconstruction gain falls below a threshold.

Pengjie Wang, Linger Deng, Zujian Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.