Skip to content
Preprint

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

Aug 2026 · 2 citations · 41 references
Computer Science

TL;DR

ReflectRL is proposed, a lightweight plug-and-play framework that learns from Golden Negative Trajectories during on-policy training, and first uses these trajectories to elicit Reflective Reasoning, then applies Reflective-to-Direct Policy Transition to transfer the acquired reasoning behavior back to Direct Reasoning.

Abstract

On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-guided methods lose their main source of supervision, and these failed trajectories are typically discarded as negative samples. We argue that such failures, which we call Golden Negative Trajectories, can still provide valuable reasoning signals when treated not as demonstrations to imitate, but as flawed trajectories to reflect upon. We identify a Reflection Advantage: for hard problems, reflecting on a flawed trajectory can be easier and more effective than solving the problem directly from scratch. Motivated by this, we propose ReflectRL, a lightweight plug-and-play framework that learns from Golden Negative Trajectories during on-policy training. ReflectRL first uses these trajectories to elicit Reflective Reasoning, then applies Reflective-to-Direct Policy Transition to transfer the acquired reasoning behavior back to Direct Reasoning. Experiments across 9 benchmarks, 4 LLM backbones, and 4 on-policy training methods show that ReflectRL consistently improves reasoning performance with minimal overhead.

View source

Similar papers

#natural language process... Preprint Sep 2026

Revisiting Complete Reasoning Traces for Post-Training

It is found that full trajectories provide only limited benefit, while partial trajectories are effective even under heavy truncation, and training LLMs using endpoints leads to consistent changes in reasoning behavior, and that it also benefits post-training methods based on reinforcement learning or on-policy distill...

Jaehui Hwang, Sangdoo Yun, Byeongho Heo et al. · 0 citations
Preprint Aug 2026

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

This work proposes Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD), which constructs two within-trajectory rankings of reasoning spans, one from teacher-derived rewards and the other from independently estimated progress reward.

Chen Yang, Hai-Yuan Wan, Rengrong Xiong et al. · 1 citation
Preprint Sep 2026

What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation

On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We fi...

Zi-Zhuo Lin, Quan-Ling Liu, Yi Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Extremely Sparse Supervision Incentivizes Reasoning Ability

Large language models demonstrate increasingly strong reasoning capabilities through effective post-training. Yet, prevailing post-training methods optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. We revisit this assumption in the on-policy distillation (OPD)...

Zhi-Shuai Liu, Xingzi Xu, Mehmet Saygin Seyfioglu et al. · 2 citations
Preprint Aug 2026

LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation

LoongReflect is proposed, a training framework that formulates reflection as a memory-control policy over a reversible trajectory tree using explicit reflect and backtrack actions and demonstrates consistent improvements over outcome-only reinforcement learning and self-distillation baselines.

Zhi-Xin Zhang, Xin-Ke Jiang, Zhi-Bang Yang et al. · 0 citations
#natural language process... Preprint Aug 2026

CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents

CAST, a critique-aware training framework that converts sparse task outcomes into action-level supervision for critique learning and policy optimization, is presented, demonstrating that critique-aware training improves the robustness of LLM agents in realistic dynamic environments.

Amir Saeidi, Zeng Zhang, Rishi Singh et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.