Skip to content

Efficient Reasoning Exploration via State-Conditioned Latent Steering with Progress Guidance

Sep 2026 · 0 citations · 49 references
Computer Science

TL;DR

SPS is a training-free latent steering framework that constructs a state-conditioned Direction Bank containing multiple progress-guided steering vectors for different prefix-state regions and applies it at high-uncertainty transitions to guide the next reasoning step toward meaningful progress.

Abstract

Best-of-$N$ is a widely used inference strategy for complex reasoning, whose effectiveness depends on whether sampled candidates can cover diverse and high-quality reasoning paths. However, post-trained reasoning models often suffer from \emph{exploration collapse}, where independent rollouts repeatedly follow similar reasoning paths and limit the gains from increasing the rollout budget. Existing methods alleviate this issue by promoting broader exploration, but do not explicitly guide exploration toward continuations that make meaningful progress, resulting in limited exploration efficiency. To address this, we propose \emph{\underline{S}tate-conditioned \underline{P}rogress-guided \underline{S}teering} (SPS), a training-free latent steering framework. Specifically, SPS constructs a state-conditioned Direction Bank containing multiple progress-guided steering vectors for different prefix-state regions. During online inference, SPS retrieves a suitable steering vector based on the current prefix state and applies it at high-uncertainty transitions to guide the next reasoning step toward meaningful progress. Extensive experiments across multiple model scales and benchmarks demonstrate that SPS consistently outperforms strong baselines. Further analyses validate the effectiveness of its key designs and offer valuable insights for future research. The code is available at https://github.com/rattlesnakey/SPS.

View source

Similar papers

#machine learning Preprint Sep 2026

Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

A central question in LLM reasoning is whether reinforcement learning (RL) instills genuinely new capabilities or merely reshapes how existing knowledge is expressed during inference. Building on the distribution-sharpening hypothesis, which holds that RL reallocates probability mass toward high-reward trajectories alr...

Zhen-Dong Mi, Shao-Yi Huang · 0 citations
#artificial intelligence Preprint Sep 2026

Efficient Reasoning via Constrained Optimization in Latent Space

Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt to mitigate this is...

Zhi-Nan Hou, Xing-Chen Li, Ke-You You · 0 citations
#artificial intelligence Preprint Sep 2026

The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration

This work introduces Problem--Strategy Rollout Allocation (PSRA), which treats unguided and strategy-conditioned prompts as competing exploration arms and uses Bayesian sequential allocation to direct a fixed rollout budget toward arms most likely to yield informative, non-saturated groups.

Jin Cui, Xin-Yue Long, Bo-Ran Zhao et al. · 1 citation
#natural language process... Preprint Sep 2026

Explore Before Committing: Hypothesis-Guided Search for Deep Research Agents

HypoSearch is proposed, which generates lightweight hypotheses as soft search hints, explores them through bounded independent branches, and compares branch-level evidence before commitment, and consistently outperforms single-trajectory search and standard parallel baselines.

Ruo-Chen Zhou, Zheng-Zong Chen, Luan Zhang et al. · 1 citation
#natural language process... Preprint Sep 2026

Learning from Think-Mode Advantage via On-Policy Distillation

Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefixes. Privileged reaso...

Wan-Qi Ren, Jian-Xiang Wang, Dan-Xuan Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Dense Process Supervision for Search Agents via Fact Utility Estimation

A dense process supervision method based on fact utility estimation, which models the reasoning process as the accumulation of discrete evidence facts, and converts the estimated fact utilities into dense step-level rewards to guide RL training.

Rongzhi Zhu, Xiang-Yu Liu, Yi Liu et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.