Off-Policy Evaluation of Contextual Bandit Algorithms for Personalized News Recommendation Using the MIND Dataset
: News recommendation systems rely on user feedback to improve ranking decisions. Running online tests adds cost. It can also interrupt the user experience. Off-policy evaluation (OPE) gives a safer option for estimating new policies. The paper models news recommendation as a contextual bandit task. The paper construct...