Skip to content

Author

Laurence Aitchison

We have 1 of 6 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

This work proposes two complementary strategies to improve the performance of value function RL: Privileged Value Functions (PVF) which provide an elegant mechanism to inject additional task-relevant token-level signal without biasing the policy objective; and TETHER, a baseline that adaptively interpolates between group-relative and value baselines depending on the value function accuracy.

S. Venkatraman, Matthieu Dinot, Laurence Aitchison · 0 citations