Looped language models (LoopLMs) increase computational depth through parameter sharing, offering a path to scale inference computation without adding parameters. However, it remains unclear when additional recurrence is beneficial and how architectural choices affect its effectiveness. Through controlled experiments,...
Xin-Lin Zhuang, Si-Yuan Wang, Imran Razzak et al.· 0 citations
Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond improvements in the distilled model's initial accuracy. Under shared RL settings, students...
Shuai Dong, Yong-Fu Zhu, Yu-Qi Xu et al.· 0 citations
Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research ideas lack uniquely verifiable outcomes. We introduce Research Attention Prediction (RAP), a rolling benchmark covering 278 AI/ML fields and 1,3...
Ying-Qian Wu, Jingcong Liang, Si-Yuan Wang et al.· 0 citations
Controlled fine-tuning, distillation, and agentic consistency-checking support the same conclusions, and the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task is described.
Qiming Bao, N. Tan, Si-Yuan Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.