Skip to content
Preprint

Auditing Machine-Learning Models and Their Training Data with Explainability and First-Principles Verification: Application to Spin Hall Conductivity

Jul 2026 · 0 citations · 31 references
Physics

TL;DR

A model-agnostic audit protocol is introduced, combining SHAP attribution, counterfactual partial dependence analysis, and Rashomon-style cross-model verification, with every finding adjudicated by targeted density functional theory (DFT).

Abstract

Machine-learning models for materials properties rest on two assumptions that standard validation never tests: that a model's features reflect the physics of the property rather than accidents of the training distribution, and that the training labels are themselves correct. We introduce a model-agnostic audit protocol for both, combining SHAP attribution, counterfactual partial dependence analysis, and Rashomon-style cross-model verification, with every finding adjudicated by targeted density functional theory (DFT). Demonstrated on intrinsic spin Hall conductivity using a composition-only Random Forest, the model needs no relaxed crystal structure, reaching accuracy competitive with structure-aware graph networks while remaining applicable to the far larger space of compositions for which no structure has been computed. The model audit reveals that the average p-valence descriptor becomes statistically entangled with Pt content - a property of the learned representation rather than the physics; DFT confirms the consequence, a Pt-free compound (HgOsPb$_2$) whose true SHC is nearly four times the prediction. The data audit exposes a thirtyfold error in the HfC training label, inherited undetectably by every black-box model trained on the same data. The protocol audits a model and its training data for the cost of a few DFT calculations, wherever one element dominates the high-property regime.

View source

Similar papers

Conference Open access 2026

Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience

This paper introduces C LUE (Clustering and Experience-based Verification) , a training-free, non-parametric verifier that improves selection and reranking in Large Language Model outputs and finds that correct and incorrect solutions exhibit measurable geometric differences in their hidden-state trajectories.

Zhenwen Liang, Ruosen Li, Yujun Zhou et al. · 0 citations
Preprint Aug 2026

Certifying Compressed Language Models: An Audit and a Statistical Toolkit

Paired equivalence testing at a declared margin is supply: paired equivalence testing at a declared margin, with certification tables giving the items an evaluation needs, computed from disagreement observed under compression, not from independent-binomial variance.

Amogh Singh · 0 citations
Preprint Aug 2026

Benchmarking Quantum Machine Learning for Power-System Attack Detection: Evaluation Choices Decide the Outcome Before the Models Do

Machine-learning detectors for power-system cyberattacks are themselves attack surfaces, and quantum machine learning has been proposed for them. We benchmark fidelity-kernel SVMs and variational classifiers against six tuned classical models on public power-system attack data (Mississippi State/ORNL), across white-box, transfer, decision-based black-box, and poisoning attacks. Our headline finding is methodological: the benchmark's answers are set by the evaluator's choices before the models. Eight choices -- six in the evaluation protocol, two in the tuning the benchmark itself runs -- each reversed or moved a conclusion at fixed models. The largest is the split: the row-level protocol scores 0.905 macro-F1 where holding whole source files out leaves 0.594, and in the capped matched-dimensionality regime the quantum arm sits within noise of chance with the classical arm 0.024 above it. A fidelity kernel looks most robust until attacked directly (retention 0.886 to 0.064); a mis-fitted surrogate manufactures a 10x asymmetry; an unseeded black-box attack moves 75% between restarts. A positive control explains the accuracy null: the labels, not the pipeline. We give the control that catches each choice and release the seeded benchmark.

Md Rezwanul Islam · 0 citations
#small language model Preprint Aug 2026

Phantom Gains: Auditing Self-Improvement Against a Measured Null

Auditing three rounds of rank-$32$ LoRA self-training on Qwen3-8B against a frozen control pushed through the identical pipeline, this work identifies seven measurement failures, each of which inverts a reported finding when its control is absent.

Cheng Xu, Nan Yan, Liming Chen et al. · 2 citations