Skip to content
Book Open access

Uncertainty-Aware Reward Modeling: A Large-Scale Case Study in Video Recommendation

Sep 2026 · Proceedings of the 20th ACM Conference on Recommender Systems · 0 citations · 6 references

Abstract

Historically, recommendation systems have focused on maximizing precision by treating user preference as a static, predictable target. However, this approach ignores both the inherent randomness of human behavior and the model’s own varying levels of confidence. To compensate, many industrial systems rely on "reserved slots" for exploring user interests—a heuristic that typically utilizes uniform selection. This paper presents a large-scale study on the YouTube video recommendation platform, where we integrated principled exploration directly into the ranking scores of the Homepage Ranking model. To model uncertainty, we evaluate two distinct architectures: a Variational Bayesian Last Layer (VBLL) designed to capture model’s parameter uncertainty, and Quantile Regression (QR) utilized to model the variance within the target reward distribution. By strategically targeting the right exploratory candidates, both paradigms successfully break the traditional explore-exploit trade-off, by driving improvements in both user engagement and content discovery metrics. Large-scale online A/B testing reveals that both paradigms are highly viable for production, successfully balancing exploration and exploitation while driving significant improvements in overall user engagement and content discovery. By analyzing the distinct behaviors of the VBLL and QR deployments, we compare both paradigms and hypothesize how the distinct uncertainties being modeled affect the final reward distributions. Finally, we discuss our ongoing efforts to merge these two paradigms and share preliminary notes on their integration.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.