Skip to content

5 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

PathoArgus: Advancing Evidence-Grounded Long-Context Visual Reasoning across Gigapixel Whole-Slide and Multi-Slide Case Contexts

Whole-slide pathology reasoning requires models to integrate gigapixel-scale visual evidence across complete case-linked slides, yet current question-answering benchmarks primarily measure final answer accuracy--a metric vulnerable to linguistic priors and benchmark regularities, and insufficient to establish that pred...

Bowen Liu, Qixiang Zhang, Xiaomeng Li · 0 citations
#artificial intelligence Preprint Sep 2026

TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models

Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reas...

Shu-Ning Wang, Zhi-Heng Wu, Xun-Lan Zhou et al. · 0 citations
Preprint Aug 2026

OVIBench: Benchmarking Online Video Question Answering under Interruption

Recent vision language models (VLMs) have achieved strong progress in video understanding. However, most existing video QA research and benchmarks still follow an offline, single-round paradigm, overlooking realistic interactions where users may interrupt the model during answer generation. To address this gap, we form...

Naiming Liu, Zhiheng Wu, Shuning Wang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores

Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether their conclusions are scientifically supported. We formulate evidence-grounded multimodal scientific analysis, requiring agents to produce executable analyses and claims supp...

Bowen Liu, Shuo Nie, Bo-Dong Du et al. · 0 citations
Jul 2026

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding

Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus on retaining relevant visual content as explicit evidence. Yet making evidence available does not ensure that complementary cues across moments are integrated for answe...

Bowen Liu, Shuning Wang, Xinpeng Ding et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.