Skip to content

AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection

Sep 2026 · 1 citation · 62 references
Computer Science

TL;DR

AlignDiff, a preference data filtering framework driven by intrinsic model signals, first identifies samples with clear preferences using both positive and inverse signals, then prioritizes the more challenging samples based on the average negative log-likelihood gap, encouraging the model to learn richer information from them.

Abstract

Aligning large language models with human preferences remains a challenge, primarily due to the critical role of preference data quality in effective alignment. Existing datasets are frequently plagued by inherent noise and distribution shifts, which inherently limit model performance. To bridge this gap, we propose AlignDiff, a preference data filtering framework driven by intrinsic model signals. AlignDiff first identifies samples with clear preferences using both positive and inverse signals, then prioritizes the more challenging samples based on the average negative log-likelihood gap, encouraging the model to learn richer information from them. AlignDiff is evaluated on two widely used model families (LLaMA and Qwen) and three benchmarks widely adopted in the alignment community (AlpacaEval 2.0, Arena-Hard, and MT-Bench). Across all settings, it consistently outperforms seven strong baselines. We conduct comprehensive ablation studies to validate the effectiveness of AlignDiff, and further show that difficulty-based curriculum learning improves model performance.

View source

Similar papers

Preprint Aug 2026

Preference Data Selection for Mitigating the Alignment Tax in Large Language Models

This paper proposes BALIGN, a balanced data selection strategy that explicitly mitigates catastrophic forgetting while optimizing alignment efficacy, and identifies three key data-centric features that dictate parameter drift: the reference model's log-probability margin, the token length between chosen and rejected re...

Minsu Kim, Jian-Xun Lian, Xing Xie et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Zeroth-Order Paradigm for LLM Preference Alignment

Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margins. In this paper, we...

Peter Chen, Xi Chen, Wo-Tao Yin et al. · 0 citations
Preprint Aug 2026

Language Chain in Alignment: Cross-lingual Ranking Preference Optimization

This paper proposes Cross-lingual Ranking Preference Optimization~ (CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language, thereby enhancing language adaptation and output quality.

Seungyoon Lee, Minhyuk Kim, Jungseob Lee et al. · 0 citations
Preprint Aug 2026

Private Direct Preference Optimization for LLM Alignment

This paper formalizes preference privacy, a label-DP-style privacy notion for DPO that protects only the relative preference between candidate responses, assuming an adversary who already knows the prompt and responses, and designs PrivDPO, a DPO variant that enforces preference privacy while remaining compatible with...

Yangfan Jiang, Fei Wei, Ergute Bao et al. · 0 citations
Jun 2026

Data-Efficient Online Training for Direct Alignment in LLMs

In recent years, online Direct Alignment from Preferences (DAP) has emerged as a popular alternative for Reinforcement Learning from Human Feedback (RLHF) due to its training stability and simplicity. In online DAP, training relies on preference data, each composed of a question and a pair of large language model (LLM)...

Chi Zhang, Jia-Chen T. Wang, Kun He et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.