Skip to content

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

Sep 2026 · 1 citation · 10 references
Computer Science

TL;DR

SemVerBench is introduced, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo), and six frontier models are evaluated: Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar).

Abstract

Large language model (LLM) coding agents constantly decide whether a version satisfies a constraint such as ^1.2.3 or>=2.0,<3, yet their grasp of version-constraint semantics has never been measured directly. We introduce SemVerBench, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo): 240 machine-checkable items with unique answers, built author-neutrally from four balanced sources (each ecosystem's official test suite plus three frontier LLM proposers) and labeled by a non-circular two-implementation oracle. Evaluating six frontier models, we find systematic, predictable per-mechanism blind spots: a partial-comparator carry rule (>1.2 means>=1.3.0) traps every model on Cargo (near 60%), and although standard PEP 440 prefix matching is universal, on zero-pad/post-release corner cases GPT-5.1 collapses (0/26) while Claude stays at 97-100% (verified on a 67-item oracle-validated set). Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar). The failures look more like an activation/application gap than a knowledge gap: injecting the rule or a light correct hint recovers most errors, whereas interval decomposition does not, and models are at ceiling on the basic forms of the same rules. An author-stratified analysis finds no statistically significant self-favoritism. Because the task is verifiable and a free, 100%-correct resolver exists, tool delegation reaches ~100%: coding agents should delegate version resolution to a resolver rather than reason about versions in-head.

View source

Similar papers

Preprint Aug 2026

SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction

SemanticAlign-Bench (SA-Bench), a diagnostic benchmark covering 30 papers from ICLR, ICML and NeurIPS 2025, is introduced and indicates that scaffolds optimized for executability provide limited leverage for scientific reproduction; narrowing the gap requires scaffolds that prioritize semantic specification verificatio...

Xueyan Hu, Ze-Wei Pan, Zeli Su et al. · 0 citations
#natural language process... Preprint Sep 2026

Constrained Decoding Eliminates Structural Failures in Small LLMs but Reveals a Scale-Dependent Semantic Gap

Small open-source large language models (LLMs) in the 0.6B-4B parameter range are increasingly deployed for structured output generation (JSON, function calling, data extraction), yet little is known about how constrained decoding (CD) interacts with model scale in this regime. We benchmark five models from three famil...

Akash Chavan · 0 citations
#artificial intelligence Preprint Aug 2026

PROOF: Profiling Reliability of Object-Level Facts in Large Language Models

Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for factual coverage in instruction-tuned language models. PROOF converts a frozen Wikidata snapsho...

A. Chetvergov, Mikhail Solovev, Timofei Sivoraksha et al. · 0 citations
#small language model Preprint Sep 2026

Code Transformation Rule Synthesis using LLMs: Potential and Limits

A novel empirical study targeting three domain- specific languages for transformation rules, namely Comby, GritQL, and Ast-Grep, which shows great potential for LLMs to generate sound, correct, generalizable, and reusable rules.

Axel Allain, Aymeric Blot, D. Khelladi et al. · 1 citation
Open access Sep 2026

EvoSort: An Audit Protocol for LLM-Driven Program Search—Correctness, Ablation, and the Cost of Too Few Seeds

EvoSort evolves compiled C++ sorting routines. A language model proposes mutations of a baseline sorter, a per-context upper confidence bound (UCB1) bandit picks which operator to try, and a four-gate harness decides whether a candidate may be timed at all. This paper reports what happened when that system was audited...

K. M. Khudhair, B. M. Khudhair · 0 citations
#machine learning Preprint Aug 2026

ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning

ClosureBench is introduced, a constructive benchmark for compositional graph-relational reasoning with programmatically verified ground truth with programmatically verified ground truth: each task's reference answer is computed by executing a program in the Ein tensor-logic language, ensuring machine-verified correctne...

S. Goria · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.