Skip to content
Preprint

Decoupling Policy Extraction for Offline Reinforcement Learning

Aug 2026 · 2 citations · 26 references
Computer Science

TL;DR

This work revisits the conventional offline RL paradigm and proposes decoupling policy improvement from actor training, and trains the actor solely to model the behavior distribution and perform policy improvement at inference time by reranking multiple actor-generated proposals with a separately learned critic.

Abstract

Offline RL methods commonly jointly train the actor and critic, where the critic is used to guide the actor toward higher-value actions. This coupled learning process is well motivated in online RL, where an improved actor collects new data that can further update the actor and the critic. However, training data remains fixed in offline RL, making actor-side policy improvement unable to generate new data to validate or correct the critic. Moreover, retaining this coupled paradigm leads to two related challenges. Firstly, actor updates can drift toward high-valued but potentially out-of-distribution (OOD) actions and amplify critic overestimation. Secondly, conservative value estimation or behavior-cloning regularization creates a difficult trade-off between suppressing OOD actions and selecting high-value actions within the data-supported region. Motivated by this observation, we revisit the conventional offline RL paradigm and propose decoupling policy improvement from actor training. Specifically, we train the actor solely to model the behavior distribution and perform policy improvement at inference time by reranking multiple actor-generated proposals with a separately learned critic. We refer to this paradigm as the decoupled policy extraction paradigm. Under such paradigm, the actor provides behavior-supported action candidates, while the critic performs value-based selection within this candidate set. Extensive experiments show that the decoupled policy extraction paradigm outperforms both behavior cloning and jointly learned offline RL methods, while remaining effective even with a naive Q-learning critic.

View source

Similar papers

Preprint Aug 2026

Simple Actors and Deep Critics for Scalable Reinforcement Learning

This work revisits where capacity should be invested in an offline actor--critic method and proposes LAC (Light Actor, deep Critic), a lightweight deterministic actor that matches the strongest diffusion- and flow-matching baselines while achieving up to 4x lower inference latency.

Gu-Heon Kang, Jaehwi Lee, Minhae Kwon · 0 citations
#machine learning Preprint Sep 2026

Role-Adaptive Policy Optimization for Offline Reinforcement Learning

Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic bootstrapping, yet methods such as TD3+BC couple these roles through a shared policy. We pro...

Seonvin Cho, Soohyun Choi, Songnam Hong · 0 citations
Preprint Aug 2026

Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning

Critic-Free Pretraining is introduced: an efficient paradigm that completely abandons the approach of offline critic training, allowing a freshly initialized critic to adapt without inheriting biased estimates.

Daoyi Li, Yi-Xian Zhang, Wen-Bo Ding et al. · 0 citations
#machine learning Preprint Sep 2026

From Static Policies to Adaptive Priors in Offline Reinforcement Learning

This position paper argues that, in such settings, the objective of offline RL should extend beyond immediate deployment and instead prioritize learning adaptive policy priors: policies that preserve the capacity to improve during subsequent interaction through memory, exploration, and self-correction.

Tian-Wei Ni, Vineet Jain, Akash Karthikeyan et al. · 1 citation
#machine learning Preprint Sep 2026

EasyPPO: Stabilizing the Critic Is Key

A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for la...

Xuan-Yi Zhou, Qiu-Yang Mang, Huan-Zhi Mao et al. · 0 citations
Aug 2026

Dual Advantage-Guided Offline Reinforcement Learning

Offline reinforcement learning aims to learn effective policies from fixed datasets without online interaction, necessitating conservative constraints to mitigate the out-ofdistribution issue. Although existing approaches alleviate this issue through conservative constraints or policy regularization, they still struggl...

Hui-Zhi Wang, Yan Kong · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.