Skip to content

Similar papers

Preprint Jul 2026

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

Under cross-benchmark transfer evaluation across twelve models completed by all four readouts, difference-in-means attains the highest model-level mean on ten models, while logistic regression wins the most capability-model cells.

Yanshi Li, Xue Bai, Shuman Liu et al. · 0 citations
Preprint Jul 2026

evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations

Evalci, a pure-Python library that turns a per-item results table into a publication-ready claim, and re-analyzes a public comparison of nine language models'MMLU accuracy to find that 3 of the 8 adjacent leaderboard-rank gaps are not statistically significant after correcting for the 36 pairwise comparisons the ranking implies.

Shreyas Chandrahas · 0 citations
Open access Jul 2026

PROBE: Benchmarking code generation in large language models

The findings show that, while LLMs achieve promising results, they struggle with harder problems and with programming languages that have fewer available resources for training, and they often fail due to fundamental and easily avoidable errors that underscore the unreliability of automatically generated code.

Rodrigo Pato Nogueira, Marco Vieira, João R. Campos · 1 citation
#software testing Preprint Aug 2026

XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models

XREPOTEST is introduced, a multilingual repository-level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby, and Invocation Rate is proposed to assess whether generated tests meaningfully exercise the intended functionality.

L. Dung, Dong Cao Van, Nam Le Hai et al. · 0 citations
Open access Jul 2026

Empirical benchmarking of large language models for data science coding: a multidimensional evaluation

The LLM4DS-Benchmark is introduced and a multidimensional empirical evaluation of seven large language models is conducted, highlighting the need for multidimensional, task-aware benchmarking and suggesting that model selection for data science coding should be guided by task characteristics and practical constraints rather than aggregate success rate alone.

Santhosh Anitha Boominathan, Sai Sanjna Chintakunta, Everton Guimarães et al. · 0 citations