On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can assign insufficient probability to other plausible continuations. Teacher entropy alone does not reveal whether uncertainty is concentrated among a few plausible next tokens or dispers...
Zi-Kun Qu, Min Zhang, Ming-Ze Kong et al.· 4 citations
WDL-OPD is introduced, a mixture-constrained co-training method with two trainable policies that shows that freezing the auxiliary recovers an anchor-plus-contrast proxy target closely related to OPD$^2$ and W2S-OPD, whereas joint training creates branch-level degrees of freedom that a static delta cannot express.
Zehao Chen, Gong-Xun Li, Tianxiang Ai et al.· 0 citations
We propose LLM-PeerReview, an unsupervised LLM Ensemble method that selects the most ideal response from multiple LLM-generated candidates for each query, harnessing the collective wisdom of multiple models with diverse strengths. LLM-PeerReview is built on a novel, peer-review-inspired framework that offers a transpar...
Zhijun Chen, Zeyu Ji, Qianren Mao et al.· arXiv.org· 5 citations
LLMODE is proposed, a token-efficient framework for irregular spatio-temporal forecasting with a frozen LLM backbone that shows competitive overall performance, with clearer advantages under sparse or dynamically complex irregular sampling.
Di Zhang, Jing-Yang Zhang, Zi-Qian Wang et al.· 0 citations
Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities, and Autonomous-Driving Multiple-Choice Question (AD-MCQ) provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.
Zi-Xuan Huang, Yang Zhou, Kai-Xuan Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.