AdaTutoRank is proposed, a setwise reranker trained with Adaptive Tutoring Optimization under a three-level hierarchy of nine rubric dimensions, which supplies silver labels for the cold start, rewards for reinforcement learning, and hints for distillation.
Kai-Lin Jiang, Lei Liu, Jian-Fei Xi et al.· 0 citations
Rubric4Setwise is proposed, a training-free method that converts rubric-based evaluation criteria into document set selection signals, achieving the best downstream generation performance with fewer documents and search rounds, validating the effectiveness of closing the loop from evaluation to optimization.
Kai-Lin Jiang, Lei Liu, Jian-Fei Xi et al.· arXiv.org· 3 citations
Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges. Outcome rewards scale readily but cannot distinguish g...
Ruoxi Cheng, Haoxuan Ma, Hongyi Zhang et al.· 0 citations
On-Policy Omni Distillation (OPOD), which consolidates text, image, and audio teachers into one omni model, and surpasses the base model and pooled RL training on all twelve benchmarks, and ranks first or second on eleven even when the teachers are included.
Tong Zhao, Yuyang Hu, Reed Li et al.· arXiv.org· 0 citations
This work introduces a generalizable evaluation framework that maps native MAS traces into a shared space of unified collaboration graphs, enabling different methods to be evaluated under the same representation, reference set, and metric panel.
Guo Chen, Ziwen Li, Reed Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.