quant_eval: A Behavioral Evaluation Harness for Full-Weight and Quantized Large Language Models
Abstract
This paper describes quant_eval, a behavioral evaluation harness for comparing full-weight and quantized large language models on agent-relevant structured tasks, and demonstrates the methodology on a single worked example. METHOD The harness scores a model against a versioned, content-addressed fixture set across eight test families: single-step JSON, multi-step JSON planning, stateful two-turn follow-up, mixed brief-plus-JSON output, multiple choice, two-stage tool calling, tool-call-only dispatch, and a property-based fuzz regression suite. Scoring is tiered. Tier 1 gates on planning feasibility and externally verified outcome correctness; Tier 2 tracks final-state reporting discipline and does not gate a Tier 1 pass. A deterministic oracle computes the correct final state independently of the model, so outcome correctness is verified rather than asserted. The paper draws a working distinction between evaluation artifacts and true degradation. A failure that a corrected normalization removes, and that a second inference backend does not reproduce, is a property of the scoring path rather than of the quantized model. Treating the two as the same thing overstates quantization damage. WORKED EXAMPLE Mistral-Nemo-Instruct-2407, evaluated across three inference backends: FP16 on Hugging Face Transformers, Q4_K_M GGUF on llama.cpp, and W4A16 AWQ on vLLM. Decoding used a fixed seed of 42 and sampled decoding at temperature 0.3, the setting recommended on the model publisher's card. Degradation under quantization is not uniform. It concentrates in particular families while others hold at parity, which is the practical argument for evaluating behavior per capability rather than reporting a single aggregate score. Stored artifact size falls from 24,504,279,808 bytes at F16 to 7,477,207,808 bytes at Q4_K_M, a 3.3x reduction. Wall-clock latency is reported for all three lanes, with the caveat that those lanes ran on three different backends, so the spread combines a backend change with a precision change. SCOPE These results document quant_eval version 7.21 under the fast_gate screening profile, which evaluates a small number of cases per family in order to triage candidates rather than to establish statistical confidence. Every family verdict is reported as provisional, and the paper states that limitation throughout. The findings characterize one base model; the framework's cross-model generality is argued from its architecture, not proved here. DATA AVAILABILITY The complete per-case output of both runs, and the provenance manifest for each, are deposited separately under concept DOI 10.5281/zenodo.22851375, CC BY 4.0. Every pass rate, bucket score, and latency figure reported in this paper can be recomputed from those files. quant_eval itself is proprietary software of PBH Applied Systems, LLC and is not published.