Skip to content

RLVR is a Kernel, Not a Function: Statistical Inference for pass@$k$ Crossovers

Sep 2026 · 0 citations · 29 references
Mathematics Computer Science

TL;DR

Reinforcement learning with verifiable rewards with verifiable rewards often improves pass@1 while falling behind its base model at larger sampling budgets, a crossover read as evidence that RLVR only sharpens existing capability is identified.

Abstract

Reinforcement learning with verifiable rewards (RLVR) often improves pass@1 while falling behind its base model at larger sampling budgets $k$, a crossover read as evidence that RLVR only sharpens existing capability. We identify two limits to this reading. First, a visible crossing need not be statistically established: comparing models on the same prompts, we build confidence bands across sampling budgets $k$ that require evidence of both an early gain and a later loss. Across five public RLVR pairs no crossing is statistically established in the initial evaluations, while a 32k-token evaluation on fresh prompts locates a reversal with first loss between 11 and 61 samples; power analysis shows why failure to detect a crossing need not mean no crossing, and why more prompts can help more than more answers per prompt. Second, base success alone does not determine what RLVR does to a prompt: prompts with the same base success rate have different post-RL success rates, and these differences repeat across independent generation halves. The relationship is a conditional distribution---a Markov kernel---rather than a single curve, and fitting it predicts crossings in independent generations for the same prompts and corrects the simple model's power estimates. Theory further shows how losses on a minority of the hardest prompts can overturn an early lead even when training improves other prompts, separating evidence that a crossover exists from claims about what it means for capability.

View source

Similar papers

#machine learning Preprint Sep 2026

Learning from the Gap Between Pass@K and Pass@1

GapFT is introduced, which selects training evidence by the source checkpoint's single-sample outcome and fine-tunes on the Pass@K-Pass@1 gap: problems the policy fails on one sample but solves within K samples, and its analysis relates available gains to transferable failure support.

Xuan-Wen Liu, Jing Qian, Hao-Sheng Chen · 0 citations
#artificial intelligence Preprint Sep 2026

Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR

Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find tha...

Chandak Chakma, Syed Nazmus Sakib, Nafiul Haque et al. · 0 citations
#machine learning Preprint Sep 2026

Cliff: Learning Process Rewards from the First Mistake

Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout, is proposed and established as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.

Pei-Xuan Han, Runnan Wang, Ketan Ramaneti et al. · 1 citation
#machine learning Preprint Sep 2026

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.

Hao Liang, Ming-Rui Chen, Hengyi Feng et al. · 1 citation
#artificial intelligence Preprint Oct 2026

Does Scaling Reinforcement Learning Really Require More Training?

Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL histor...

Bang Yang, Jia-Jun Fan, Hong-Bo Ma et al. · 0 citations
#machine learning Preprint Sep 2026

The Statistical Benefits of Multiple Responses for Learning from Demonstrations

It is found that multiple responses provide a qualitatively stronger benefit in this setting of generative systems, and is shown that a greedy multiplicative-weights learner achieving the upper bounds without any assumption on demonstrator quality is possible.

Chandramauli Chakraborty, Cong Ma · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.