Skip to content

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

Sep 2026 · 0 citations · 71 references
Computer Science

Abstract

Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, trading reduced computation on easy instances for accuracy loss on hard instances. We formulate efficient reasoning as an instance-adaptive computation allocation problem and propose When2Think, a post-training framework for hybrid reasoning that dynamically allocates computation based on problem difficulty. Our method introduces Instance-level Difficulty-Aware Control (IDAC), a reward-shaping mechanism that leverages pre-computed reference statistics (accuracy and token usage) to regulate reasoning depth. Combined with verifier-based rewards and batch-wise standardized advantages, IDAC enables stable critic-free optimization without learned reward models or online reference-model queries. When2Think encourages direct answering on easy instances while preserving extended reasoning on hard instances, thereby learning when to use System 1 (NoThink) versus System 2 (Think). Experiments on mathematical benchmarks demonstrate improved accuracy-efficiency trade-offs: on AIME24, Pass@3 increases by 10.0% while token usage is reduced by 27.9% relative to the base model, and on AIME25, When2Think achieves 40.0% Pass@3, outperforming compression and routing-only baselines.

View source

Similar papers

Preprint Aug 2026

Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

A reasoning model is built that adaptively chooses how much to reason for each problem, and the brief modes end up more accurate than \textsc{Long}, which shows that the router sorts problems by difficulty rather than at random.

Gijs Kassenaar, Zhao Yang, Vincent François-Lavet · 1 citation
#natural language process... Preprint Sep 2026

From Base Rollouts to RL Reasoning: A Budgeted Search Perspective

Results support a qualified internalized-search reading: under the recipe the authors test, much of the measured RL gain corresponds to a change in sampling efficiency toward operating points the base model can already reach under search.

Wen-He Sun, Cun-Xiang Wang, Zi-Jun Yao et al. · 0 citations
Preprint Aug 2026

Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning

ReCo (Reward-Coordinated Compression), a step-wise framework in which a lightweight process-reward estimator scores each completed step and drives three components: reward-adaptive KV-cache compression that shrinks the retained cache harder at high-reward steps and less at low-reward ones, and a confidence-based early...

Qi-Yuan Zhu, De-Zhi Li, Pengyu Cheng et al. · 1 citation
#reinforcement learning Book Open access Aug 2026

LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing

Recommendation systems thrive on personalization, where “correctness” is rarely a binary truth but a matter of subjective human preference. As Large Language Models (LLMs) are deployed as autonomous verifiers of safety and quality guidelines, they face a distinctive challenge: context-aware preference alignment. Recent...

Jun-Cheng Dong, Ding Tong, Ishan Gupta et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration

Reinforcement learning with verifiable rewards (RLVR) is often limited by insufficient exploration: difficult problems can yield uniformly incorrect rollout groups and therefore little learning signal. We show that such failures need not reflect missing capability. Instead, finite sampling often concentrates on a probl...

Jin Cui, Xin-Yue Long, Bo-Ran Zhao et al. · 0 citations
Preprint Aug 2026

Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning

Results show that test-time compute can be more effective when used to refine sampled trajectories rather than only to sample more candidates or rely on verifier-guided selection.

Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.