Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold, which gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-k gradients.
Abstract
Reinforcement learning typically optimizes average reward. For generative policies, the average can hide an important distinction: two policies can achieve the same mean reward while having very different chances of producing a rare but high-reward rollout. This matters as sampling increases during training and inference, since its benefit depends on retaining probability mass on high-reward outcomes. We propose to optimize this coverage directly. Rather than considering only expected reward, we consider all of its upper tails: for each reward threshold, how likely is the policy to exceed it? This turns a continuous reward into a family of binary success events. We introduce Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold. Its gradient gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-k gradients. TailRL requires only a simple modification to the advantage function, making it compatible with existing reinforcement learning pipelines. Across object localization, maze navigation, GUI grounding, and code optimization, TailRL leverages rare high-reward training samples to avoid suboptimal solutions and yields models that benefit more from additional samples at inference time.
This work proposes a general RL-based framework for Distribution Matching allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution and proposes reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justific...
Sourabh Kulkarni, Ksheeraj Sai Vepuri, B. Demir et al.· 0 citations
This work introduces eVTA, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations, and introduces RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA using newly collected roll...
Duo Wu, Hai-Feng Wang, Rongwei Lu et al.· 0 citations
Empirically, EPIG reduces gradient MSE in cloned-state control, winning in all nine dense continuous-control environments of a 13-environment sweep and recovering the reference gradient direction near-perfectly, and it improves frozen-LLM gradient calibration relative to entropy branching.
Nikita Khomich, L. Hermansson, Ido Hakimi· 0 citations
Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a sca...
Xin-Ke Jiang, Tao Feng, Zhi-Bang Yang et al.· 0 citations
Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature enginee...
This work proposes LaRes, a novel hybrid framework that achieves efficient policy learning through reward function search by leveraging large language models to generate the reward function population, guiding RL in policy learning.
Pengyi Li, Hongyao Tang, Jinbin Qiao et al.· Neural Information Processin...· 6 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.