Skip to content

When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings

Sep 2026 · 1 citation · 20 references
Computer Science

TL;DR

The results show that high self-consistency does not necessarily indicate high agreement with human judgments when using local LLMs as automatic judges, and highlight the need to evaluate both consistency and human alignment when using local LLMs as automatic judges.

Abstract

Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faster and cheaper than human evaluation. However, a judge may produce consistent scores without necessarily agreeing with human evaluators. In this work, we study this issue using two local open-weight LLM judges, LLaMA-3-8B and Qwen2.5-7B. We evaluate 300 responses generated by an instruction-tuned GPT-2 (124M) model for 100 questions covering five categories: factual knowledge, instruction following, mathematics, reasoning, and writing. Each response is scored by nine human annotators and is evaluated three times by each LLM judge using the same rubric. We compare the judge scores with the average human scores using Pearson correlation, Spearman correlation, mean absolute error (MAE), signed bias, and self-consistency. LLaMA-3-8B shows a Pearson correlation of 0.275 with human scores, while Qwen2.5-7B achieves 0.340. Their MAEs are 27.71 and 18.64, respectively. Despite this limited agreement, both judges show high self-consistency, with exact consistency rates of 97.3\% for LLaMA-3-8B and 92.3\% for Qwen2.5-7B. These results show that high self-consistency does not necessarily indicate high agreement with human judgments. Our findings highlight the need to evaluate both consistency and human alignment when using local LLMs as automatic judges.

View source

Similar papers

#natural language process... Preprint Sep 2026

LLJ Cards: Best practices for the Use of LLMs as Judges

In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-eff...

Khaoula Chehbouni, Melina Medjdoub, Florian Carichon et al. · 0 citations
#small language model Open access Sep 2026

Human versus machine

Although LLM-based benchmark ratings can approximate expert judgments and reduce the need for labor-intensive human triple coding, limitations remain regarding cost, genre- and task specificity, and sensitivity to text presentation and student grade level—factors that constrain immediate classroom use, particularly for...

Afra Sturm, Valentin Unger, Fabian Grünig · 0 citations
#artificial intelligence Preprint Sep 2026

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has studied biases in pairwise LLM-as-a-judge settings; in this paper, we focus on absolute scoring tasks, which mirror more realistic use case...

Gemma Zhang, Prachi Badarayani, Asmi Kumar et al. · 0 citations
Preprint Aug 2026

A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation

The results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.

Jian-Lin Chen, Wen-Hui Chen, Zi-Yao Lin et al. · 3 citations
Preprint Aug 2026

Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation

Results demonstrate that the proposed method outperforms the baselines, including an LLM without debiasing and previous calibration methods, and it is confirmed that scoring bias varies across LLMs, tasks, and score ranges, indicating the importance of measuring latent number bias as the case may be.

Yuma Asato, Kiyoaki Shirai, Natthawut Kertkeidkachorn · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.