Skip to content

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

Sep 2026 · 0 citations · 13 references
Computer Science

TL;DR

This work systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026, using staged screening and automated full-text coding to examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms.

Abstract

Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The collection shows growing emphasis on action, interaction, and professional applications, while established and newer design elements frequently coexist. Model participation also develops unevenly: LLM-based scoring grows within both agent and non-agent groups, whereas model-generated materials show no comparable sustained increase in recent cohorts. These findings illuminate how public research translates capability expectations into concrete tests and criteria for success. As AI participates in constructing tests, performing tasks, and judging responses, they also raise a question: does expanding evaluation provide more independent evidence, or risk reproducing the preferences and blind spots of its participating models?

View source

Similar papers

#natural language process... Preprint Sep 2026

LLJ Cards: Best practices for the Use of LLMs as Judges

In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-eff...

Khaoula Chehbouni, Melina Medjdoub, Florian Carichon et al. · 0 citations
#machine learning Preprint Sep 2026

How Reusable Are Benchmarks with Richer Feedback?

We study whether benchmarks reliably guide model selection as developers adapt to evaluation feedback across multiple criteria. We find that the worst-case test-set size needed to estimate the best score among $k$ adaptively chosen models, under any convex combination of the criteria, grows exponentially with the numbe...

Youssef Allouah, John C. Duchi · 0 citations
Preprint Aug 2026

Who's Keeping Score? Interactive Steering of LLM-Powered Scoring with Attune

Attune is presented, a mixed-initiative system for steerable LLM-powered scoring that performs pairwise comparisons across records to develop a global understanding first, and then resolves these comparisons into consistent score assignments-deriving scoring criteria and rules bottom-up in the process.

Bhavya Chopra, Meng Chen, Rebecca Dang et al. · 0 citations
Conference Open access Sep 2026

Mapping the Efficiency Landscape of Small Language Models

This work evaluates 70+ SLMs from 2023–2025 on five task-specific benchmarks and compares them with two popular LLMs, revealing key trade-offs between energy, performance, and model selection and highlighting the need for informed, task-aware model selection rather than size-driven choices.

Fabian Reichwald, Lukas Schiesser, Christiane Plociennik et al. · 0 citations
Preprint Aug 2026

LangChoiceBench: Measuring and Explaining Programming-Language Choice in LLMs

LangChoiceBench is introduced, a project-level code-generation benchmark for measuring Python preference, recommendation-implementation consistency, and language diversity, and it is found that Python remains heavily over-selected, recommendation-implementation consistency is low, and smaller open-weight models general...

Lukas Twist, Twm Stone, Helen Yannakoudakis et al. · 0 citations

What does AI mean for Open Source?

The recent meteoric rise of LLMs (Large Language Models) and associated tools was largely unexpected and surprising to most. The rapid ascent of this technology has caught many software developers unawares, leaving them suddenly somewhat ignorant, and arguably under-skilled. LLMs, whilst still advancing, have recently...

Adam Retter · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.