Fire (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning, is proposed, which provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where s...
Seohyun Lee, Dong-Jun Han, Seyyedali Hosseinalipour et al.· 0 citations
Her HermesHFL, a hierarchical federated learning framework that supports selective unlearning, dynamic client participation, and client reintegration for scalable LLM fine-tuning via parameter-efficient fine-tuning (PEFT) with LoRA, is proposed and developed.
Chenxi Sun, Minghui Liwang, Wu-Si He et al.· arXiv.org· 0 citations
This work proposes a two-stage generator-in-the-loop alignment framework that consistently outperforms rank-order, random, and REPLUG-style likelihood baselines under various alignment losses and pool size settings, suggesting that answer-level generator feedback is an effective supervision signal for preference alignm...
Zhang-Yu Chang, Dong-Jun Han, Seyyedali Hosseinalipour et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.