Skip to content
Preprint

Tail-Likelihood Reinforcement Learning

Sep 2026 · 1 citation
Computer Science Mathematics

TL;DR

Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold, which gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-k gradients.

Abstract

Reinforcement learning typically optimizes average reward. For generative policies, the average can hide an important distinction: two policies can achieve the same mean reward while having very different chances of producing a rare but high-reward rollout. This matters as sampling increases during training and inference, since its benefit depends on retaining probability mass on high-reward outcomes. We propose to optimize this coverage directly. Rather than considering only expected reward, we consider all of its upper tails: for each reward threshold, how likely is the policy to exceed it? This turns a continuous reward into a family of binary success events. We introduce Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold. Its gradient gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-k gradients. TailRL requires only a simple modification to the advantage function, making it compatible with existing reinforcement learning pipelines. Across object localization, maze navigation, GUI grounding, and code optimization, TailRL leverages rare high-reward training samples to avoid suboptimal solutions and yields models that benefit more from additional samples at inference time.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning

This work proposes a general RL-based framework for Distribution Matching allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution and proposes reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justific...

Sourabh Kulkarni, Ksheeraj Sai Vepuri, B. Demir et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies

This work introduces eVTA, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations, and introduces RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA using newly collected roll...

Duo Wu, Hai-Feng Wang, Rongwei Lu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning

Empirically, EPIG reduces gradient MSE in cloned-state control, winning in all nine dense continuous-control environments of a 13-environment sweep and recovering the reference gradient direction near-perfectly, and it improves frozen-LLM gradient calibration relative to entropy branching.

Nikita Khomich, L. Hermansson, Ido Hakimi · 0 citations
#machine learning Preprint Sep 2026

When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation

Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a sca...

Xin-Ke Jiang, Tao Feng, Zhi-Bang Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents

Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature enginee...

Mu-Hang Tian, Sherry Yang · 0 citations
2025

LaRes: Evolutionary Reinforcement Learning with LLM-based Adaptive Reward Search

This work proposes LaRes, a novel hybrid framework that achieves efficient policy learning through reward function search by leveraging large language models to generate the reward function population, guiding RL in policy learning.

Pengyi Li, Hongyao Tang, Jinbin Qiao et al. · 6 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.