LLM judges increasingly evaluate responses against fine-grained rubric checklists. When a sample requires multiple rubrics, current methods typically assess each in a separate inference call. Evaluating all rubrics in a single pass is a natural alternative with greater efficiency, but we find that it introduces rubric...
Dingyao Yu, Tong Zhang, Yutao Mou et al.· 0 citations
This work proposes SHIFT, a retrieval training framework based on LLMs that transfers LLMs into reasoning-efficient retrievers with residual projection and task-oriented bidirectional attention aggregation in the latent space, and alleviates the mismatch between contrastive learning and implicit reasoning using fine-gr...
Yuxiao Luo, Da Li, Mingjie Zhang et al.· arXiv.org· 0 citations
It is found that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful.
Zhaoyi Li, Deyang Kong, Yuan Wei et al.· 1 citation
Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness, and ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.
Yutao Mou, Pengfei Yang, Zhenfei Yin et al.· 0 citations
A novel Le arnable Lo w-R ank A daptation (LeLoRA) framework that utilizes dynamically learned fine-tuning strategies to facilitate the effective adaptation of LLMs and provides compelling evidence that LeLoRA consistently outperforms existing baselines in adapting LLMs.
Xiaoling Zhou, Mingjie Zhang, Zhemg Lee et al.· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.