Reinforcement learning (RL) trains representations on data selected by the agent's policy, which then uses the resulting returns to guide its next choices. We show that this loop can sustain a lower-return policy even when representation fitting is globally optimal on those data. In a self-confirming superposition trap...
Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident than their corresponding pretrained language models (PLMs). Because task-specific labeled calibration data can be costly or unavailable, the corresponding pretrained PLM...
Linhan Luo, Lequan Lin, Dai Shi et al.· 0 citations
Scaling laws hold that language models grow more capable with more parameters and more training data. Mixture-of-Experts (MoE) architectures are a remarkable demonstration of these laws, activating only a fraction of an enormous parameter bank for each token. But this success is built on static pretraining data --- the...
Jin-Lin Hu, Ross M. Clarke, Yi-Chuan Zhang et al.· 0 citations
Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump, has recently attracted considerable debate. A prominent position holds that LLMs are structurally incapable of such jumps, while recent studies challenge both its mechanism and em...
Dai Shi, Xiao-Yu Li, José Miguel Hernández-Lobato· 0 citations
Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons and motivates methods for recovering interpretable features from neural activations. Theoretical models typically start with a given set of input features and assumptio...
A formal account of the jump is developed in four steps and measured, proving that jump instances are well-posed and establish a family theorem that certifies instances of unbounded difficulty without enumeration and further formalize when a jump is correct and how successive jumps compound.
Dai Shi, Xiao-Yu Li, José Miguel Hernández-Lobato· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.