Exposure-Based Reinforcement Learning to Rank
This work considerably improves RL for LTR methodology by increasing its effectiveness, efficiency, and ease of application, and proposes an abstraction that places gradient estimation behind a document-exposure distribution that enables seamless plug-and-play integration with auto-differentiation.