Skip to content

Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Serving-Time Weight Compression

Sep 2026 · 0 citations · 14 references
Computer Science

TL;DR

This work serves three checkpoints at three weight precisions, holding the hardware, software, and sampling configuration constant, and collects approximately 71,000 completions paired by prompt and seed across two custom, leak-checked prompt batteries.

Abstract

Weight quantization largely determines the economics of serving open-weight LLMs. Its costs are usually assessed with capability benchmarks, on which 4-bit quantization of mid-sized models is often considered"nearly free."We examine a different question: when several answers are valid, does quantization change what a model chooses to say? We serve three checkpoints (Qwen3-8B/14B/32B) at three weight precisions (W4A16 AWQ, W8A16 FP8-Marlin, and bf16), holding the hardware, software, and sampling configuration constant, and collect approximately 71,000 completions paired by prompt and seed across two custom, leak-checked prompt batteries. We pre-specified the analyses in three waves in version control. At 8B, int4 reduces output diversity: the probability that two samples for the same scenario recommend the same brand increases by 5.1 percentage points (prompt-paired sign-flip test, Holm p = .023; reproduced at +4.4pp on a full regeneration of the arm), and lexical diversity falls substantially (TTR -0.011, standardized effect -0.51; robust to a length-controlled measure). At 14B and 32B, no content-concentration measure reaches significance; instead, stylistic drift emerges (em-dash rate +0.46/1k words at 14B and +0.61/1k at 32B, both Holm p<= .0024). Pre-specified tests of stereotype direction are null at every scale: outputs concentrate on the modal answer for each prompt rather than on stereotypical answers. Mechanistically, the token-level distribution becomes flatter (decision-token entropy +0.091 bits, p = .015) while the semantic distribution, measured directly from first-token log probabilities, becomes more concentrated (collision +2.6pp, p = .023): individual tokens become less predictable even as meanings become more repetitive. At 8B, the smallest size tested, AWQ-int4 serving measurably narrows the range of suggestions; audits should assess concentration as well as bias.

View source

Similar papers

#machine learning Preprint Sep 2026

Steering Under Compression: Dose-Response, Capability Cost, and Failure Asymmetry in Quantized LLMs

This work systematically study activation steering under weight-only quantization (INT8 and NF4) across four open-weight 7-9B models and two behavioral targets: judged sentiment and judge-free reasoning length, finding that sentiment steering survives quantization intact.

Saurav Bhandari, Benjamin Wade · 0 citations
#machine learning Preprint Sep 2026

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional precision budget. Using causal mixed-precision intervention as ground...

Jun-Hao Hu, S. Ramachandran · 1 citation
Conference Sep 2026

Quantization-Induced Accuracy Degradation in Small Language Models: A Parametric Scaling Analysis

Post-Training Quantization (PTQ) enables the deployment of large language models (LLMs) on resourceconstrained edge devices, yet the accuracy-efficiency trade-off for Small Language Models (SLMs, 3 B parameters) remains poorly characterized. In this study, we present a systematic empirical study of quantization effects...

Pushpendra Tripathi, Aditya Sharma, Himani Sharma et al. · 0 citations
Preprint Aug 2026

SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization

SCHUROPT is introduced, which analytically eliminates the suffix's optimal continuous response, yielding an exact groupwise quadratic with Schur-complement curvature, and achieves the highest mean zero-shot accuracy among the evaluated backpropagation free PTQ baselines.

Gunjun Lee, Sehwan Son, Younjoo Lee et al. · 0 citations
#natural language process... Preprint Aug 2026

Which Decisions Low-Bit Quantization Breaks, and How to Predict Them

This work tracks quantization across 16 models from 8 families under round-to-nearest, seven under AWQ, two under GPTQ and one under GGUF, at 8 down to 2 bits, and measures the margin, the picked option's score minus its best alternative's, which removes the protection a large margin affords.

Zekun Wu, Swati Dhiman, A. Koshiyama · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.