Skip to content

Scaling Test-Time Compute via Generative Verification in Constrained Parameter Regimes

· 0 citations · 2 references

TL;DR

This project investigates scaling test-time compute through a Generative Verifier (GV) on the Countdown mathematical reasoning task using a computationally constrained 0.5B parameter regime, hypothesizing that the “verification gap” will widen at higher values of N due to the model’s limited semantic capacity.

View source

Similar papers

Conference Open access 2026

Crossing the Reward Bridge: Expanding Reinforcement Learning with Verifiable Rewards Across Diverse Domains

This work provides a scalable and effective framework for extending RLVR beyond the limitations of pattern-based verification to complex, noisy, real-world domains, and generalizes strongly to seven out-of-distribution benchmarks.

Yi Su, Dian Yu, Linfeng Song et al. · 1 citation

When Does Execution Feedback Transfer? Minimal-Sufficient

This project studies that failure mode in a controlled Countdown arithmetic setting and asks whether verifier-generated execution feedback can supply a minimally sufficient correction signal that later transfers into a no-feedback policy, confirming the proposal’s central concern.

Yu-Jie Yao · 0 citations
Conference Open access 2026

Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience

This paper introduces C LUE (Clustering and Experience-based Verification) , a training-free, non-parametric verifier that improves selection and reranking in Large Language Model outputs and finds that correct and incorrect solutions exhibit measurable geometric differences in their hidden-state trajectories.

Zhenwen Liang, Ruosen Li, Yujun Zhou et al. · 0 citations
Preprint Aug 2026

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

TrajVal, a lightweight probe-based estimator that approximates per-task learnability from a short probe run and two endpoint evaluations, is proposed and it is found that learnability is reproducible across independently sampled training contexts and predictive of downstream utility.

Ting Zhou, Zhenqing Ling, Daoyuan Chen et al. · 0 citations
Preprint Jul 2026

Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards

This work establishes the first non-vacuous generalization bounds for parameter-efficient RLVR fine-tuning at the billion-parameter scale, and proposes the Progressive RLVR framework, which integrates RLVR with on-policy distillation, TinyLoRA, and model quantization.

Yuxuan Zhu, Rohan Alur, Daniel Kang · 0 citations