Skip to content
Preprint

QuoteBench: How Matched Scores Can Hide Command-Path Failures

Aug 2026 · 1 citation · ⚡ 1 influential · 40 references
Computer Science

TL;DR

Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.

Abstract

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser. Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must come from the model changing its generation. Across eight same-window configurations, replaying the same reply through the added parser lowers success by 55.4 to 73.2 percentage points; disclosure recovers 30.4 to 60.7 points for six configurations, and zero or slightly negative for the other two. Raw generation is nearly saturated at the frontier; boundary adaptation is what still separates models. GPT-5.6-sol's matched gap of -3.6 points hides -64.3 points of damage and +60.7 points of compensation. The deployment configuration reorders models: one reversal among 26 comparable pairs is unambiguous and four more sit on single-task margins. Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents

Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the...

XiaoYu Xu, Zi Liang, Min-Xin Du et al. · 0 citations
#artificial intelligence Preprint Sep 2026

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advanta...

Zi-Yang Xu, Haitian Zhong, Hao Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics

LLM-based code generation fails when correctness depends on execution-dependent coupling: the meaning of one routine is defined by the runtime behavior of another, a relationship that cannot be resolved from textual descriptions alone. This limitation, which we call static binding, is not confined to explicitly coupled...

Gnaneswar Villuri, Hashmath Shaik, Alex Doboli · 1 citation
#artificial intelligence Preprint Sep 2026

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

SemVerBench is introduced, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo), and six frontier models are evaluated: Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar).

Qi-Bai Chen, Ze-Ming Liu · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.