Skip to content
Book Open access

Distribution-Value Coevolution for Adaptive RLHF Data Scheduling

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 6104-6115 · 0 citations · 39 references

TL;DR

This work identifies and formalizes the Distribution-Value Coevolution principle: the training value of data is not intrinsic, but emerges dynamically from the interaction between data characteristics and the model's evolving capability boundary, and operationalizes this principle through a unified framework.

Abstract

Reinforcement learning from human feedback (RLHF) has become the cornerstone of aligning large language models (LLMs) with human intent. Yet a fundamental question remains unaddressed: how should training data be scheduled when both the model's capabilities and the utility of data are constantly evolving? Current pipelines rely on fixed or uniform sampling, treating data value as static, an assumption we demonstrate to be fundamentally flawed. We identify and formalize the Distribution-Value Coevolution principle: the training value of data is not intrinsic, but emerges dynamically from the interaction between data characteristics and the model's evolving capability boundary. What is highly informative at one stage may become redundant, or even detrimental, at another. This insight demands a paradigm shift from static to adaptive curriculum design. We operationalize this principle through a unified framework with three components: (1) distribution-level organization that groups training data into coherent distributions; (2) sliding-window influence estimation that continuously tracks each distribution's evolving training value; and (3) bandit-guided scheduling that adaptively allocates resources with provable exploration-exploitation guarantees. Experiments show that this approach yields measurable improvements, with up to a 57.1% relative (or 8.9% absolute) improvement on AIME24 for Llama3.2-3B, and gains also observed for models ranging from 1B to 7B parameters.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Frontier Learning: Training LLM Reasoners at the Edge of Capability

Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as learning signal aris...

Robin Faro, S. Ramesh, Ilija Bogunovic et al. · 0 citations
#machine learning Preprint Sep 2026

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. Our primary exp...

Hao Liang, Mingrui Chen, Hengyi Feng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

Scaling laws hold that language models grow more capable with more parameters and more training data. Mixture-of-Experts (MoE) architectures are a remarkable demonstration of these laws, activating only a fraction of an enormous parameter bank for each token. But this success is built on static pretraining data --- the...

Jin-Lin Hu, Ross M. Clarke, Yi-Chuan Zhang et al. · 0 citations
Preprint Aug 2026

Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization

Modern systems are increasingly expected to transfer across tasks not specified during training. What data facilitates generalization in these new, unanticipated settings? One hypothesis is that data with more structural information could contain shared circuits and subprograms that could be recycled in a wider array o...

Ellen Su, Andres Potapczynski, Shikai Qiu et al. · 0 citations
#natural language process... Preprint Sep 2026

Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) has shown remarkable success in improving the mathematical reasoning of large language models. Yet problems beyond the model's current capability, where rollouts uniformly fail and no learning signal is produced, are structurally wasted despite marking the most info...

Yu-Kang Zhu, Zhen-Mao Han · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.