Quantization-Induced Accuracy Degradation in Small Language Models: A Parametric Scaling Analysis
Abstract
Post-Training Quantization (PTQ) enables the deployment of large language models (LLMs) on resourceconstrained edge devices, yet the accuracy-efficiency trade-off for Small Language Models (SLMs, 3 B parameters) remains poorly characterized. In this study, we present a systematic empirical study of quantization effects across the Qwen2.5-Instruct model family (0.5 B, 1.5 B, 3 B) at five GPT-Generated Unified Format (GGUF) precision levels (8-bit to 3-bit) on Apple M3 hardware with Metal GPU acceleration. On evaluating 15 model-quantization configurations on two benchmarks: a.) GSM8K (mathematical reasoning) and b.) ARC-Challenge (science reasoning), we obtain 30 controlled data points. We propose and validate a novel parametric scaling law, $A(P, Q, C)=\alpha \ln (P)+\beta\left(1 e^{-\gamma Q}\right)+\delta C+\epsilon$, which jointly models accuracy as a function of parameter count $P$, quantization bit-width $Q$, and task complexity $C$, achieving effective $\boldsymbol{R}^{\mathbf{2}}=\mathbf{0. 9 2 9}$ with a Leave- One-Out Cross-Validation Mean Absolute Error (LOOCV MAE) of 4.4 percentage points. Our analysis reveals that Q4_K_M preserves accuracy within 1-3 percentage points (pp) of the 8-bit baseline while halving memory; 3-bit quantization causes catastrophic degradation on multi-step reasoning (21 pp) yet remains resilient on classification (3.5 pp); and an ablation confirms degradation is largely task-independent $\left(\boldsymbol{\Delta} \boldsymbol{R}^{\mathbf{2}}=\mathbf{0. 0 0 3}\right)$. We identify 3B-Q4_KM (2 GB, 82.5% accuracy, 26 tok/s) as the Pareto-optimal configuration for fanless edge deployment.