Skip to content

Author

Sirui Han

We have 5 of 49 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations

Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understanding. We study whether this representation can serve as operational context for agentic coding, where an agent must navigate repositories, edit source files, and verify executable...

Weijie Liang, Yuanfeng Song, Xing Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring operational safety across multi-step workflows has become a critical challenge. While recent work has moved beyond single-turn evaluation toward multi-turn paradigms, key limitations persist: step-level methods treat ac...

Jia-Peng Sun, Yu-Jin Zhou, Han Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents

As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal benchmarks evaluate only single-turn QA with outcome-level metrics, while agentic...

Yu-Jin Zhou, Min Zheng, Chuxue Cao et al. · 0 citations
Preprint Aug 2026

SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent

SkillProx is introduced, a proximal-gradient-inspired forward--backward framework that couples closed-loop diagnostic evolution with utility-aware proximal refinement and demonstrates the complementary effects of closed-loop diagnosis and proximal refinement.

Mingxuan Zheng, Yu-Jin Zhou, Chuxue Cao et al. · 2 citations
Conference Open access 2026

Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across Modalities

Omni-RewardBench is introduced, the first benchmark for comprehensive evaluation of ORMs across modalities and demonstrates that current OLLMs fall short as reward models, revealing several common failure modes such as perception failure, modality dominance failure, and cross-modal fusion failure.

Chi-Min Chan, Yujin Zhou, Pengcheng Wen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.