Skip to content

Author

Xiyang Wu

We have 5 of 23 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

ASSEMBLE: Atomic Skills for Evidence-Grounded Video Reasoning

Complex video reasoning often depends on evidence scattered across distant moments, entities, and events, yet a correct answer alone does not reveal whether a model relied on the right parts of the video. We introduce ASSEMBLE, a framework that makes supporting evidence explicit throughout long-video reasoning. ASSEMBL...

Xi-Yang Wu, Zong-Xia Li, Sheng Zhang et al. · 0 citations
Preprint Sep 2026

Causeway: Restoring Task Accessibility for Instruction Switching in VLA Policies

Vision-language-action (VLA) policies can execute many tasks from standard initial states, yet a new instruction may fail after another task has altered the robot's physical state. We study instruction switching, where a new task is issued during or after the execution of a different one. We observe that a target task...

Qing-Zi Wang, Kai Feng, Guang-Yao Shi et al. · 0 citations
Preprint Aug 2026

Noise Floor Audit for Agent Benchmarks

We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At temperature 0, reruns are nearly deterministic across Groq endpoints and a thinking-enabled Gemini setting: ever-flip fractions are 0.7%, 2.0%, and...

Yi-Hang Chen, Pinyan Qian, Su Wang et al. · 0 citations
Jul 2026

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

This work introduces Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing, and analyzes failure modes and error patterns to support future progress on...

Zongxia Li, Zhongzhi Li, Yucheng Shi et al. · 11 citations · ⚡2
Preprint Aug 2026

Recursive Agentic Reasoning

Analysis shows that BRANCH's advantage arises not only from exploring multiple reasoning paths, but also from recovering from truncation: its gains strongly correlate with the baseline rate of empty, budget-exhausted outputs, weakening the hypothesis that different problems require routing among test-time reasoning ope...

Sheng Zhang, Xiao-Min Wu, Xi-Yang Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.