Skip to content

Author

Peng-Cheng Wen

We have 2 of 12 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring operational safety across multi-step workflows has become a critical challenge. While recent work has moved beyond single-turn evaluation toward multi-turn paradigms, key limitations persist: step-level methods treat ac...

Jia-Peng Sun, Yu-Jin Zhou, Han Zhu et al. · 0 citations
Conference Open access 2026

Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across Modalities

Omni-RewardBench is introduced, the first benchmark for comprehensive evaluation of ORMs across modalities and demonstrates that current OLLMs fall short as reward models, revealing several common failure modes such as perception failure, modality dominance failure, and cross-modal fusion failure.

Chi-Min Chan, Yujin Zhou, Pengcheng Wen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.