On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need not improve future b...
Yu-Hao Sun, Bin-Rui Wu, Zhuo-Er Xu et al.· 0 citations
Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multimodality and Multi-Teacher On-Policy Dis...
Lu-Jia Bao, Qian Chen, Luyao Cheng et al.· 1 citation
Refusal-Enhanced INhibitory Steering is proposed, which suppresses harmful continuation features and enhances safe refusal features in the same SAE feature space, and substantially reduces harmful responses, markedly improves safe refusals and largely preserves general capabilities.
Kai-Xuan Ding, Haoyang Xu, Jihua Peng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.