Skip to content

When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain

Sep 2026 · 0 citations · 59 references
Computer Science

TL;DR

A holistic framework based on learnable information gain, which measures how much novel, parameterizable information a round provides relative to the previous round, and proposes ATRI (Adaptive Training Regulation via Information-gain), which reweights samples within a round and halts training across rounds when information gain remains low.

Abstract

Self-evolution lets large language models (LLMs) improve iteratively using their own generated data, but often suffers from self-evolution degeneration: performance improves, plateaus, then declines. Existing methods address this issue at the component level, targeting either the Questioner or the Solver, and overlook that self-evolution is a tightly coupled system. We propose a holistic framework based on learnable information gain, which measures how much novel, parameterizable information a round provides relative to the previous round. Theoretically, this gain equals the Kullback-Leibler divergence between the two rounds'data distributions plus their entropy change. Practically, it is estimated by fitting a small language model to the previous round and scoring new data via negative log-likelihood. Based on this diagnostic, we propose ATRI (Adaptive Training Regulation via Information-gain), which reweights samples within a round and halts training across rounds when information gain remains low. Experiments on popular datasets demonstrate the superiority of our proposal.

View source

Similar papers

#machine learning Preprint Sep 2026

UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?

Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this issue, existing unlearning methods typically rely on training-based parameter updates, such as gradient ascent and its variants, to delete targeted content while preservin...

Pu-Ning Yang, Qi-Zhou Wang, Jun-Chi Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Continual Learning via Self-Probe Gradients

Adapting pretrained models to new data can cause catastrophic forgetting of previously learned behavior. When only a few past samples remain, they give continual learning methods sparse and narrow evidence about what to preserve. We show that language models can expand this evidence through self-probing, in which the f...

Dongkyu Cho, R. Chunara, Sung-Min Cha · 0 citations
#artificial intelligence Preprint Sep 2026

COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement

COEVO is introduced, a framework that updates model parameters from on-policy experience while adapting contextual guidance according to the state of the evolving policy, and suggests that external context should be viewed not merely as a fixed interface to a large language model, but as an adaptive component of recurs...

Si-Wei Chen, Xin-Ping Bao, Xin-Yu Cai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training

Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}vo...

Yu-Yang Deng, Yu Wang, Jia-Yun Wang · 0 citations
#natural language process... Preprint Sep 2026

Time-Incremental Continued Pretraining of LLMs: Knowledge Updates Without Catastrophic Forgetting

Large language models (LLMs) drift out of date the moment their pretraining ends, yet retraining from scratch is prohibitively expensive. Continued pretraining (CPT) is the natural remedy, but it is typically evaluated through a continual learning lens that assumes disjoint data streams. This is a poor fit for time-inc...

Firat Öncel, S. Ali, M. Ravanelli et al. · 0 citations
Book Open access Aug 2026

ANCHOR: Taming Entropy Dynamics for Stable and Efficient Reasoning of Large Language Models

It is shown that policy entropy bounds both the policy gradient and probability update norms; consequently, entropy collapse effectively stops reward signal backpropagation, preventing further policy learning regardless of data quality.

Cong Qin, Jiaye Lin, Xiaoliang Fu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.