Calibration of Variational Quantum Classifiers Under Depolarizing Noise: Expected Calibration Error, Ansatz Expressibility, and Post-Hoc Temperature Scaling
Jul 2026· 2026 International Conference on Intelligent and Sustainable AI Systems (ICOSAAS)· pp. 297-303· 0 citations· 17 references
Abstract
Variational Quantum Classifiers (VQCs) have emerged as prime candidates for machine learning on NISQ systems. It stands to reason that the same depolarizing noise that drives quantum states toward the maximally mixed state would also alleviate the overconfidence of VQCs. This paper tests that hypothesis through an empirical study on three datasets at six noise levels, validated across ten random seeds and supported by a formal analysis of how the depolarizing channel contracts the measured logits. The primary finding is that depolarizing noise does not reduce overconfidence in VQCs that remain in the learnable regime: expected calibration error (ECE) never decreases with noise on any dataset, and on the lowest-variance dataset it increases significantly (Wilcoxon signed-rank p < 0.005 over ten seeds). We derive why the optimizer compensates for the channel and confirm the mechanism through a confidence-trajectory experiment. A secondary finding is that a less expressive ansatz can appear well-calibrated only because it collapses to a degenerate solution, demonstrating that ECE must always be reported alongside accuracy. Any effect of noise on accuracy is small and seed-dependent, and it is decoupled from calibration. Post-hoc temperature scaling reduces VQC ECE by 64 to 77 percent across all datasets and is the recommended calibration method for NISQ-era classifiers.
Noisy Intermediate-Scale Quantum (NISQ) devices impose structural constraints on the training of Variational Quantum Algorithms (VQAs). In the absence of full error correction, each quantum gate introduces a non-zero probability of error that accumulates with circuit depth, while gradient estimation through finite sampling adds additional statistical variability. As a result, convergence depends on a delicate balance between physical coherence and estimation variance. In this setting, the depth of the variational ansatz not only determines model expressivity, but also its operational feasibility under noise. Without an explicit characterization of the interaction between error accumulation, number of measurements (shots), and control strategies, increasing experimental resources may fail to improve performance and can even become counterproductive. In this work, we empirically analyze how circuit depth affects the practical signal of the gradient under depolarizing noise. Across multiple configurations ($q=4-7$ qubits), we observe a pattern consistent with an effective exponential attenuation of the coherent gradient contribution, characterized by a per-layer rate $\epsilon_{\text{eff}}$. The estimated magnitude of $\epsilon_{\text{eff}}$ depends on circuit complexity and remains largely independent of the measurement budget, indicating that it primarily reflects architectural and physical factors rather than sampling effects. The analysis of adaptive control strategies further suggests that accuracy gains are constrained by gradient attenuation, rather than scaling significantly with circuit depth. However, a crossmode comparison across three distinct training strategies reveals that, while accuracy improvements remain modest, adaptive perturbations substantially improve the statistical detectability of the attenuation signal, acting as diagnostic probes rather than mere optimization heuristics. Taken together, these results suggest that the usable depth of a variational ansatz can be interpreted as an emergent property of the effective signal-to-noise regime, governed by a measurable structural parameter whose ranking across circuit configurations is preserved independently of the training mode. This framework enables principled comparison of circuit architectures in terms of their effective trainability under noise, providing guidance for architectural and control design in Quantum Machine Learning under NISQ conditions.
C. Braga, Manuel A. Serrano, E. Fernández-Medina· 2026 IEEE International Conf...· 0 citations
It is shown that gradient-based PQCs can exhibit improved performance on unseen data as model size increases, displaying the phenomenon of double descent, which contrasts with the traditional view that larger models lead to degraded generalization.
Marie C. Kempkes, Elies Gil-Fuster, Carlos Bravo-Prieto et al.· 0 citations
In these small, idealised, classically simulated matched-family tasks, the support-basis DQFIM provides a useful data-dependent pre-training diagnostic of effective capacity on the retained data support and contributes predictive information beyond raw parameter count and structural metadata.
Shreyosha Ganguly, A. Masta, Shalini Devendrababu et al.· Academia Quantum· 0 citations
This work solves the inner layer of quantum-measurement reduction globally and certifiably as a second-order cone program (SOCP), and uses RANGE, a robust adaptive nature-inspired global optimizer, for the combinatorial and statistical outer layer.
This work operationalizes the conditions as SPECTRA, a two-tier certificate: a simulator-free structural screen followed by a decisive comparison between a matched quantum model and five tuned classical twins using paired-bootstrap confidence bounds, providing a practical recipe for identifying, engineering, and deploying quantum-advantage candidates in tabular data.
It is shown that finite quantum measurement statistics (shot noise) act as a built-in defense against gradient-based test-time attacks whose cost scales unfavorably for the attacker.
Bacui Li, Chandra Thapa, Tansu Alpcan et al.· 0 citations