Skip to content

Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels

Sep 2026 · 1 citation · 32 references
Computer Science

Abstract

Who the judge is can affect an LLM-as-judge result, but measuring that effect without confusing it with candidate quality is difficult. We study four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) in a fully crossed pairwise design with 9,312 judgments. A common per-family statistic is strongly confounded with candidate quality and correlates with Bradley-Terry ability at r = 0.95. We derive a corrected estimator that holds the candidate family fixed and compares judges. All four families then show a positive same-family lift (3.4-8.4 percentage points), with global FPS 0.067 (95% CI [0.053, 0.084], permutation p = 0.0002). The effect remains under panel-based quality controls, an independent human-consensus anchor, and a float16 judging replication. Judge-side likelihood is closely related to the effect: adding likelihood advantage reduces the controlled coefficient by 61%, which we treat as descriptive attenuation rather than causal mediation. Position is a separate failure mode. Across the panel, 55.4% of AB/BA pairs reverse, and reversal above 50% is incompatible with a simple independent content-noise model. Relative to a family-balanced reference, panel composition changes 18.5% of pairwise outcomes. A complete reproducibility archive has been prepared for public release.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has studied biases in pairwise LLM-as-a-judge settings; in this paper, we focus on absolute scoring tasks, which mirror more realistic use case...

Gemma Zhang, Prachi Badarayani, Asmi Kumar et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes. We study how this dependency affects the reliability of consensu...

Elias Hossain, Niloofar Yousefi, Ser-Nam Lim · 1 citation
Preprint Aug 2026

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

The results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.

Jian-Lin Chen, Wen-Hui Chen, Zi-Yao Lin et al. · 3 citations
Open access Aug 2026

condfair: An R Package for Ability-Conditioned Fairness and Explanation Diagnostics in Automated Scoring.

Fairness in automated scoring is typically evaluated with a single global statistic contrasting a focal and reference group-an approach that can either mask a disparity that changes sign across the ability range, or overstate one by conflating it with genuine ability differences between groups (impact). We introduce co...

Tri Zahra Ningsih, Aman Aman, Ahmad Nasrulloh · 1 citation
#artificial intelligence Preprint Sep 2026

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and re...

Yubo Li, Yidi Miao, Ramayya Krishnan et al. · 12 citations · ⚡3
#artificial intelligence Preprint Sep 2026

Opening LLM Judges: Recovering Preference Signals Beyond the Final Verdict

LLM judges are widely used to evaluate model outputs, but their verdicts can be unreliable: a judge may favor the worse answer for its position, length, or other surface features. When a judge is wrong, is the information needed to judge correctly absent from the model, or present in its internal representations but no...

Sourabrata Mukherjee, Sunayana Sitaram · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.