AdaTutoRank is proposed, a setwise reranker trained with Adaptive Tutoring Optimization under a three-level hierarchy of nine rubric dimensions, which supplies silver labels for the cold start, rewards for reinforcement learning, and hints for distillation.
Kai-Lin Jiang, Lei Liu, Jian-Fei Xi et al.· 0 citations
Rubric4Setwise is proposed, a training-free method that converts rubric-based evaluation criteria into document set selection signals, achieving the best downstream generation performance with fewer documents and search rounds, validating the effectiveness of closing the loop from evaluation to optimization.
Kai-Lin Jiang, Lei Liu, Jian-Fei Xi et al.· arXiv.org· 3 citations
OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.
Baochen Fu, Wenzhi Deng, Baihao Jin et al.· arXiv.org· 3 citations
Multi-Agent Protocol Distillation (MAPD), a joint distillation and RL framework uses a structured, style-normalized protocol as an intermediate representation that generalizes robustly across diverse proprietary teachers while effectively mitigating the student policy from style drift and verbosity degeneration.
Junlin Liu, Jiangwang Chen, Zixin Song et al.· arXiv.org· 6 citations
By unifying four hallucination dimensions with paired question design, KnowHal addresses an important gap in existing evaluation frameworks and enables a more comprehensive assessment of hallucinations in MLLMs.
Ruihan Li, Ji-Yang Tan, Kai-Lin Jiang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.