Skip to content

Author

Dingwei Zhu

We have 5 of 15 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address thi...

Junjie Ye, Zhuohui Sheng, Shao-Hua Liu et al. · 0 citations

Prefix-Adaptive Block Diffusion for Efficient Document Recognition

The Prefix-Adaptive Block Diffusion Model (PA-BDM) is proposed, which replaces intra-block bidirectional denoising with causal denoising from prefix to suffix and treats the block size as a maximum candidate range rather than a fixed commitment unit.

Ming-Xu Chai, Zi-Yu Shen, Chen-Yu Liu et al. · 0 citations

JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees

A novel textual representation of fault trees is proposed, and a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments is constructed, evaluating a model's ability to assist in malfunction localization.

Yuhui Wang, Zhi-Xiong Yang, Ming Zhang et al. · 0 citations
Preprint Aug 2026

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles, is introduced, suggesting that a self-improving search agent needs feedback that co-evolves with the policy it guides.

Bo-Yang Liu, Senjie Jin, Pei-Xin Wang et al. · 1 citation
Conference Open access Jul 2026

AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

AgentGym2 is presented, a new evaluation framework with task instances grounded in real-world end-to-end working demands that measures agents'ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information.

Zhiheng Xi, Dingwen Yang, Jiaqi Liu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.