Skip to content
Preprint

What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend

Aug 2026 · 0 citations · 5 references
Computer Science

Abstract

Benchmark scores are reported as properties of a model, yet the inference framework used to produce them, such as HuggingFace, vLLM, or Ollama, are considered non-influential and their names and versions are almost never disclosed. In this work we investigate how much this choice can influence the model output. In a fully-crossed study (three instruction-tuned models x five inference frameworks x six benchmarks x four generation modes) we investigate how different tools (wrappers/backend) influence benchmark scores and how their score changes is influenced by generation hyper-parameters. We find backend to be a non-negligible factor where even under greedy, sampling-noise-free decoding, changing the backend can significantly alter models performance and this effect is structural and strongly model-dependent. Decomposing the variance according to generation mode reveal that considerable portion of the variability (roughly 39\%) a practitioner sees out-of-the-box can stem from the backend, while the remaining stems from sampling noise and each framework's default generation parameters, both of which are avoidable by disclosing and matching the generation configuration. These divergences are more pronounced on factual than on social-bias benchmarks. Overall, benchmark numbers are not backend-agnostic therefore, we recommend disclosing the backend, its version, and the full generation configuration, also using deterministic decoding for cross-backend comparison.

View source

Similar papers

Preprint Jul 2026

Position: It's Time to Optimize LLMs for Self-Consistency

This position paper observes that a wide variety of techniques designed to improve specific aspects of LM behavior-targeting properties as diverse as adversarial robustness and factual coherence-can be understood as special cases of a common "consistency optimization" procedure and addressed with a standard set of optimization tools.

Itamar Hagay Pres, Belinda Z. Li, L. Ruis et al. · 1 citation
Preprint Aug 2026

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

It is shown that rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer in both directions, so some failures become successes and some successes become failures, and BenchDrift generates meaning-preserving variations of benchmark problems along four axes, namely linguistic, referential, pragmatic, and structural.

Shailja Thakur, Sungeun An, Chad DeLuca et al. · 0 citations
Preprint Aug 2026

Rewriting or Reweighting? A Geometric Account in Language Models

Behavior manifold analysis is introduced, which isolates behavior-specific geometry by selecting sparse behavior-associated coordinates and lifting them into low-dimensional local charts, and provides a unified framework for understanding the mechanistic distinction between the two objectives.

Juntong Wang, Shengkun Yang, Xiyuan Wang et al. · 0 citations
#small language model Preprint Aug 2026

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost.

Yu-Fan Wu, Yinghui He, Zhengyi Hu et al. · 1 citation
Preprint Aug 2026

Why Large Language Models Fail at Tabular Prediction

The results show that the LLM's capability dissolves with dimension in a way no noise-corrupted classical learner mimics - which explains why LLMs, so capable elsewhere, keep losing to fifty-year-old baselines on tables, while leaving the mechanism of the prediction as an open question.

M. Garnelo, Wojciech M. Czarnecki · 1 citation
Preprint Jul 2026

Is Model Instability just Noise to be Tolerated or a Property that can be Managed?

In software analytics, rerunning the same analysis twice often yields different models and conclusions. This reduces trust in the model and limits its use. We find that model instability is a major problem. Across 127 multi-objective SE optimization problems (12,700 test cases), repeated runs of a state-of-the-art optimizer agree on only 13.7% of test cases, even under improved settings. We argue that this instability is not merely noise to tolerate, but a property that can be measured and managed. By adjusting how labels are spent, how complex the models become, and how splits are scored, we obtain models that agree 4.8 times as often as the default configuration. The standard deviation of optimization error falls by 22% on average (mean std 17.4 to 13.6), while recommendation quality improves rather than degrades. In terms of quality, the refined settings are statistically top-ranked on 119 of 127 datasets, compared to 74 for the defaults. We then test causal and data-locality interventions and find that they help only partially, suggesting a residual stability floor. Our evidence suggests there are fundamental limits to stability set by the data itself (noise, scarce labels, proxy objectives, and the many near-equivalent models a dataset admits). We conclude that instability should be treated as a standard evaluation axis in SE optimization, which should be routinely measured, reported alongside performance, and used to calibrate trust in any single run. The methods in this paper provide a baseline against which future efforts to reduce SBSE instability can be judged. To support open science, we offer the following reproduction package: https://tinyurl.com/Model-Instability

Amirali Rayegan, Lunxiao Li, Tim Menzies · 0 citations