Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with pr...
Zhenrui Yue, Hui-Min Zeng, Yue-Qi Wang et al.· 0 citations
Generative recommendation reformulates sequential recommendation as an autoregressive generation task, yet a critical issue in this paradigm remains overlooked: topology distortion in item tokenization. In particular, we observe that the intrinsic adjacency relationships of items in the pretrained semantic embedding sp...
This work proposes a two-stage framework to assess the severity of false claims during disasters, and investigates false claim severity assessment as a human-AI alignment problem, evaluating whether models can reproduce human judgments under a shared evaluation rubric rather than merely predicting severity labels.
Ruichen Yao, Tejna Dasari, G. Baispay et al.· Proceedings of the 2026 ACM...· 0 citations
PropUQ-MAS is proposed, an error propagation-aware UQ framework that represents MAS execution as a communication-structured graph and estimates each step's reliability by combining local uncertainty with uncertainty inherited from upstream messages.
Yaokun Liu, Yifan Liu, D. Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.