Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans

While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largely uncharacterized. Existing benchmarks focus on isolated text or formulas, overlook handwritten tables and real-world degradation, and report aggregate accuracy without explaining why models fail. We present WildHandBench, a benchmark containing 500 handwritten documents across three structures (free text, tables, formulas), four languages, and nine real-world scenarios. We introduce a Prior-Driven Error (PDE) metric that quantifies whether errors originate from language priors rather than visual evidence. Evaluating 18 state-of-the-art models together with calibrated human baselines, we find: (1) the best model achieves only 71.85% overall; (2) humans outperform all models yet the gap is narrow (77.09% vs. 71.85%); and (3) model errors are qualitatively different from human errors -- 63-91% of model errors are prior-driven versus only 49% for humans, exposing systematic reliance on language priors that conventional accuracy metrics cannot capture.

Jun Zhang, Qiao Zhao, Cheng Cui et al. · 0 citations
Preprint Aug 2026

BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests

BulkPR-Bench is introduced, an executable benchmark in which an agent must recover consequential PR relations and return a large safe subset in executable order in executable order under a rolling-release protocol.

Zetong Xiong, Qiao Zhao, Jun Zhang et al. · 0 citations