Recent advances in multimodal large reasoning models (MLRMs) have demonstrated impressive capabilities on complex multimodal tasks, yet their reliance on long Chain-of-Thoughts (CoTs) often leads to redundant reasoning and high computational cost. Existing chain-based distillation and refinement approaches alleviate re...
Yizhi Wang, Li-Nan Yue, Deng-Bao Wang et al.· Proceedings of the 32nd ACM...· 0 citations
LightTIR, a dual-penalty reward framework, is proposed to achieve efficient TIR and can reduce redundancy and trajectory expansion while maintaining answer correctness, achieving more efficient RL-based TIR.
Yichen Xiao, Siyu Gong, Li-Nan Yue· Proceedings of the 32nd ACM...· 0 citations
Recent methods using Reinforcement Learning (RL) have improved Tool-Integrated Reasoning (TIR) by training large language models to learn end-to-end policies for multi-step tool usage, enabling them to solve complex tasks more effectively. Despite these advances, existing methods often suffer from overthinking at both...
Yichen Xiao, Siyu Gong, Linan Yue· Proceedings of the 32nd ACM...· 0 citations
Recent advances in multimodal large reasoning models (MLRMs) have demonstrated impressive capabilities on complex multimodal tasks, yet their reliance on long Chain-of-Thoughts (CoTs) often leads to redundant reasoning and high computational cost. Existing chain-based distillation and refinement approaches alleviate re...
Yizhi Wang, Linan Yue, Deng-Bao Wang et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.