Skip to content
Conference

Quantum Kernel Concentration Under Class Imbalance: Empirical Characterisation with Discrimination Ratio and Quantum Imbalance Vulnerability Score

Jul 2026 · 2026 International Conference on Intelligent and Sustainable AI Systems (ICOSAAS) · pp. 802-807 · 0 citations · 17 references

Abstract

Quantum kernel methods are a candidate approach for machine learning on near-term quantum hardware, but two practical problems limit their deployment: kernel values concentrate exponentially as the qubit count grows, and real-world datasets are often severely class-imbalanced. We present the first systematic empirical study of how these two effects interact. We define two diagnostic metrics, the Discrimination Ratio (DR) and the Quantum Imbalance Vulnerability Score (QIVS), which measure whether quantum kernels retain minority-class separability under concentration. Experiments span ten log-spaced imbalance ratios, five random seeds, five qubit counts (4 to 12), and five real-world datasets, and yield three results. First, DR stays above 1.0 at every qubit count tested (4 to 12), so the discriminative signal survives concentration. Second, at extreme imbalance (IR below 0.003), quantum kernel SVMs retain positive discriminative signal, crossing above DR=1.0 by IR≈0.0028, while the classical oversampling methods SMOTE and ADASYN produce zero minority-class recall throughout the same regime, a practical advantage for quantum kernels at the imbalance extreme. Third, QIVS follows a broadly monotonic decreasing trend, falling from 13.25 to 5.54 as the imbalance ratio increases, with a single minor fluctuation smaller than the cross-seed variability we measure elsewhere in the sweep. This trend makes QIVS a reliable diagnostic for practitioners choosing quantum kernels on imbalanced tasks.

View source

Similar papers

Conference Jul 2026

Benchmarking Classical and Quantum Machine Learning for Intrusion Detection Across Multiple Datasets

This paper presents a comparative benchmarking study of classical and quantum machine learning models for intrusion detection using three benchmark datasets: NSL-KDD, UNSW-NB15, and MQTTEEB-D2025. The study evaluates how preprocessing choices, feature selection strategies, and quantum encoding methods influence model performance across datasets with different levels of noise and complexity. A unified pipeline is adopted, incorporating normalization, imbalance handling, dimensionality reduction, and two feature selection approaches: Random Forest importance and a quantum-aware method based on Quantum Kernel Alignment with Mutual Information. Four models are assessed: Support Vector Machine, Random Forest, Quantum Support Vector Machine, and Pegasos Quantum SVM. Results show that classical models remain stable across datasets, while quantum models are more sensitive to feature representation and kernel alignment. Quantum performance improves significantly with quantum-aware feature selection, particularly on cleaner datasets, whereas heterogeneous datasets remain challenging. Pegasos Quantum SVM offers a favorable balance between accuracy and computational efficiency, highlighting the importance of preprocessing alignment for practical quantum intrusion detection.

Taha M. Mahmoud, N. Kaabouch · 0 citations
Preprint Aug 2026

Evaluating Quantum Kernel Methods for Track-Based Classification in High-Energy Physics

We present a systematic design for large-scale quantum kernel classification, demonstrated through a quantum support vector classifier (QSVC) for particle-track classification using centroid-based CLAS12 drift-chamber features. Each event is encoded into a six-qubit state via a fully entangled ZZFeatureMap, whose fidelities define a quantum kernel within a standard SVM framework. By decoupling state preparation from kernel construction and distributing evaluation across a multi-node MPI-based HPC allocation, the approach scales to 1.0x10^5 training and 4.0x10^5 test events with an exactly constructed kernel matrix, to our knowledge more than an order of magnitude larger than prior high-energy-physics quantum-kernel studies. Benchmarked against linear, polynomial, RBF, and sigmoid SVM kernels and extremely randomized trees (ERT), the ideal QSVC achieves the highest recall (99.99%) among all models. Under a calibrated hardware noise model (FakeMumbaiV2, 500 training / 2,000 test events), AUC falls from 0.9985 to 0.9671 and peak significance improvement falls from 17.5 to ~3.5, yet recall remains at 99.51% -- indicating this signal-retention advantage is attenuated but not eliminated by circuit-level decoherence. Geometric analysis of the quantum embedding shows near-orthogonal inter-class states with coherent intra-class neighborhoods under ideal simulation; under noise this structure compresses toward the maximally mixed state while preserving its relative ordering. These results demonstrate a scalable, reproducible workflow for quantum kernel experimentation at HEP-relevant scale, quantifying the practical cost of realistic hardware noise on quantum-enhanced classification.

Emmanuel Billias, Nikos Chrisochoides · 0 citations
Conference Open access 2026

Quantum Kernel Support Vector Machine for Quantum Dot State Recognition

: Semiconductor quantum dot platforms require rapid recognition of charge states during automated device tuning, especially when labeled data are scarce and device-to-device variation is strong. This study evaluated state recognition from 100 by 100 two-gate current maps using 30 by 30 labeled patches under a strict device-level split. Patch features were transformed with a signed logarithmic scale, standardized, compressed to four principal components, and scaled to the interval from zero to two pi. A fidelity quantum kernel support vector machine implemented in IBM Qiskit with a four-qubit ZZFeatureMap was compared against linear and radial basis function support vector machines across few-shot budgets of 10, 20, 40, and 80 samples per class. At 80 samples per class, the quantum kernel model achieved accuracy and macro F1 near 0.947 on held-out devices, outperforming the classical baselines in the evaluated setting. Noise-injection experiments showed stable macro F1 under perturbation. These findings support fidelity-based quantum kernels as practical components for automated quantum dot tuning pipelines requiring few-shot generalization.

M. Ashakin, Rubayat Khan, A. Mahata et al. · 0 citations
Preprint Jul 2026

Quantum-Enhanced Synthetic Data Generation Using Quantum Circuit Born Machines for Imbalanced Tabular Learning

Findings establish QCBM as a viable complementary tool for data augmentation, particularly for low-dimensional structured tabular data with class imbalance, particularly for low-dimensional structured tabular data with class imbalance.

Tanapol Nuatho, Narisorn Sangnakara, Prapong Prechaprapranwong et al. · 0 citations
Preprint Aug 2026

Quantum Kernel k-Means for Credit-Card Fraud Detection:A Controlled Benchmark on Real Transaction Data

It is shown that ordinary hyperparameter choices move performance by considerably more than the quantum kernel does, that additional qubits degrade rather than improve performance through kernel concentration, and that the clustering framing itself fails at realistic class imbalance though kernel-based anomaly scoring does not.

Muhammad Faryad · 0 citations
Preprint Aug 2026

Benchmarking Quantum Machine Learning for Power-System Attack Detection: Evaluation Choices Decide the Outcome Before the Models Do

Machine-learning detectors for power-system cyberattacks are themselves attack surfaces, and quantum machine learning has been proposed for them. We benchmark fidelity-kernel SVMs and variational classifiers against six tuned classical models on public power-system attack data (Mississippi State/ORNL), across white-box, transfer, decision-based black-box, and poisoning attacks. Our headline finding is methodological: the benchmark's answers are set by the evaluator's choices before the models. Eight choices -- six in the evaluation protocol, two in the tuning the benchmark itself runs -- each reversed or moved a conclusion at fixed models. The largest is the split: the row-level protocol scores 0.905 macro-F1 where holding whole source files out leaves 0.594, and in the capped matched-dimensionality regime the quantum arm sits within noise of chance with the classical arm 0.024 above it. A fidelity kernel looks most robust until attacked directly (retention 0.886 to 0.064); a mis-fitted surrogate manufactures a 10x asymmetry; an unseeded black-box attack moves 75% between restarts. A positive control explains the accuracy null: the labels, not the pipeline. We give the control that catches each choice and release the seeded benchmark.

Md Rezwanul Islam · 0 citations