Uncertainty-Aware Reward Modeling: A Large-Scale Case Study in Video Recommendation
Abstract
Historically, recommendation systems have focused on maximizing precision by treating user preference as a static, predictable target. However, this approach ignores both the inherent randomness of human behavior and the model’s own varying levels of confidence. To compensate, many industrial systems rely on "reserved slots" for exploring user interests—a heuristic that typically utilizes uniform selection. This paper presents a large-scale study on the YouTube video recommendation platform, where we integrated principled exploration directly into the ranking scores of the Homepage Ranking model. To model uncertainty, we evaluate two distinct architectures: a Variational Bayesian Last Layer (VBLL) designed to capture model’s parameter uncertainty, and Quantile Regression (QR) utilized to model the variance within the target reward distribution. By strategically targeting the right exploratory candidates, both paradigms successfully break the traditional explore-exploit trade-off, by driving improvements in both user engagement and content discovery metrics. Large-scale online A/B testing reveals that both paradigms are highly viable for production, successfully balancing exploration and exploitation while driving significant improvements in overall user engagement and content discovery. By analyzing the distinct behaviors of the VBLL and QR deployments, we compare both paradigms and hypothesize how the distinct uncertainties being modeled affect the final reward distributions. Finally, we discuss our ongoing efforts to merge these two paradigms and share preliminary notes on their integration.