Preprint
Aug 2026
Preference Data Selection for Mitigating the Alignment Tax in Large Language Models
This paper proposes BALIGN, a balanced data selection strategy that explicitly mitigates catastrophic forgetting while optimizing alignment efficacy, and identifies three key data-centric features that dictate parameter drift: the reference model's log-probability margin, the token length between chosen and rejected responses, and the TF-IDF similarity to general capability corpora.
Minsu Kim, Jianxun Lian, Xing Xie et al.
· 0 citations