Existing Supervised Fine-Tuning paradigms, particularly Full Parameter Fine-Tuning are often plagued by parameter redundancy, inconsistent data quality, and catastrophic forgetting, which current methods typically address in isolation and lack a unified optimization signal to bridge data selection, parameter updates, a...
Ze-Yu Wu, Junchao Wu, Shu-Dong Liu et al.· 0 citations
Reinforcement learning (RL) excels on tasks with verifiable rewards, but in open-ended tasks, the reliability of reward models remains a key challenge. Existing solutions either depend on costly proprietary LLM-as-a-Judge systems or opaque scalar reward models that lack interpretability. Recent works on generative rewa...
Peng Lai, Yi-Chao Du, Junchao Wu et al.· 1 citation
CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context.
Bo Zeng, Lin-Feng Gao, Pei-Qing Lin et al.· 0 citations
Credit-Aware Hierarchical Memory Evolution (CHIME), a self-evolving memory framework that maintains a separate planning bank and execution bank and follows an attribute-before-memorize principle, which shows that CHIME consistently outperforms state-of-the-art training-based and self-evolving memory baselines.
Yongshi Ye, Tian Lan, Feihu Jiang et al.· 0 citations
This work proposes STAR-masked Preference Optimization (StarPO), a framework that ranks document-level hypotheses by structural quality and utilizes a dynamic alignment mask to focus optimization on misaligned segments and demonstrates that StarPO significantly enhances translation quality and structural integrity.
Yichen Dong, Hao Wang, Junhui Li et al.· 1 citation
WnW (Waxing-and-Waning KV cache), which classifies KV-heads into anchor, tidal, and fixed roles via offline calibration, and preserves near-Full-Cache accuracy while keeping only 20% of audio tokens on GPU, where prefill-only baselines fail to terminate.
Yi-Ming Yao, Chenyang Lyu, Xuan-Fan Ni et al.· 0 citations
This work proposes M ulti-Model C ontrastive D ecoding (MCD), which integrates a pretrained language model with an evil model and a truthful model for contrastive decoding and effectively reduces hallucinations in LLMs and outperforms state-of-the-art methods across various benchmarks.
Chenyu Zhu, Yefeng Liu, Hao Zhang et al.· Neural Information Processin...· 7 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.