Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the local optimization of pe...
Zi-Yang Cheng, Yu-Hao Wang, Hong-Cheng Liu et al.· 0 citations
This work finds that finetuning a reward model to guide the policy model is more robust than directly finetuning the policy model, and proposes AgentRM, a generalizable reward model, to guide the policy model for effective test-time search.
Yu Xia, Jing-Ru Fan, Weize Chen et al.· Annual Meeting of the Associ...· 26 citations· ⚡3
DirsWorld turns cross-device collaborative operation into an executable, reproducible, and diagnostically useful evaluation problem for research on reliable cross-device agents, and evaluates five frontier LLM-agent systems on a fixed evaluation set.
Hua-Tao Li, Xinwei Geng, Yu-Heng Wang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.