Preprint
Jul 2026
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models
A consensus-based evaluation framework that measures relative preference among model-generated responses rather than absolute correctness rather than absolute correctness is introduced, offering an alternative perspective on response quality in scenarios where multiple valid answers exist.
Mohtashim Khan
· 0 citations