Skip to content

Author

Rui-Qi Liu

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation

On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher-student agreement to weight, filter, or truncate the sampled trajectory. These signals do not directly determine whether t...

Ximo Zhu, Rui-Qi Liu, Rong Wang et al. · 3 citations
#machine learning Preprint Sep 2026

Nonparametric Variance-Penalized Actor-Critic: Statistical Inference for Risk-Sensitive Reinforcement Learning

Variance penalization is a principled approach to risk-sensitive reinforcement learning (RL) that explicitly trades expected return for policy stability. Existing methods require a dedicated second critic to estimate return variance online, adding architectural complexity and compounding estimation error during learnin...

Saunak Kumar Panda, Tong Li, Yi-Sha Xiang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Data-free On-policy Distillation

On-policy distillation (OPD) has become a standard component of frontier post-training pipelines, yet how much its training data actually contributes has gone largely unexamined. On the two teacher--student pairings most common in practice, we find OPD almost indifferent to its data: eight prompts already match a 17k-p...

Gengsheng Li, Mao Zheng, Ming-Yang Song et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.