Preprint
Aug 2026
Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons
The Boltzmann-rational model for pairwise comparisons is adopted, which extends the Bradley-Terry-Luce model by incorporating worker competencies, and an EM-based algorithm for learning is derived by introducing Polya-Gamma latent variables to transform the logistic likelihood into a conditionally Gaussian form, enabling tractable optimization and leading to a simplified $Q$ function in the E-step of the algorithm.
Kaustubh Shivshankar Shejole, Tanish Agarwal, Arpit Agarwal et al.
· 0 citations