Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage training, whereas RL enables stable coexi...
Kejian Zhu, Zhuo-Ran Jin, Shangqing Tu et al.· 0 citations
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. While conventional \textit{agent-side warming} up via supervised fine-tuning (SFT) can alleviate th...
SwarmBench is proposed, a benchmark that evaluates model performance from multiple perspectives, including accuracy, efficiency, cost, and process quality, and SwarmExp is proposed, a simple yet effective method based on experience extraction and experience replay, which consistently improves the orchestration performa...
Jin Gao, Zhuoran Jin, Tianyi Men et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.