Skip to content

Author

Julia Kreutzer

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Aug 2026

Dynamically Allocating Evaluation Effort for Model Ranking

While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation protocols waste effort by exhaustively evaluating all models on the entire benchmark, a safe but inefficient approach. In this work, we formalize multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms, where pulling an arm corresponds to human-evaluating a model. By sampling adaptively based on the intermediate model rankings obtained on the samples so far, we can focus the annotation budget on the most competitive models. We prove the optimality of the proposed algorithms and show that it improves discrimination between top-performing models. This makes evaluations faster, cheaper and more aligned with large-scale competition evaluation goals.

Vilém Zouhar, Julia Kreutzer, A. Lavie et al. · 0 citations
Preprint Jul 2026

JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation

JudgeArena is an open-source framework that unifies major LLM-judge benchmarks under a single interface with swappable judges and comprehensive metadata logging for increased transparency in reporting and reproducibility and enables systematic studies of judge choices.

Erlis Lushtaku, Bora Kargi, Ali Elganzory et al. · 0 citations