Preprint
Sep 2026
Tail-Likelihood Reinforcement Learning
Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold, which gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-k gradients.
Shrinivas Ramasubramanian, Daman Arora, Fahim Tajwar et al.
· 1 citation