A unified uncertainty-aware back-end comprising uncertainty-aware cosine scoring, uncertainty-aware AS-Norm (UAS-Norm), and uncertainty-aware Quality Measure Function calibration (UQMF).
Abstract
Speaker verification back-ends commonly combine similarity scoring, score normalization, and calibration. However, speaker embeddings extracted from real-world utterances have trial-dependent reliability because of factors such as duration, noise, and channel variation. Existing uncertainty-aware methods primarily improve the speaker encoder or the initial similarity score, while the estimated uncertainty is typically not propagated through subsequent normalization and calibration. We represent each utterance by a speaker embedding, interpreted as a posterior mean, together with its covariance as an uncertainty estimate. We present a unified uncertainty-aware back-end comprising uncertainty-aware cosine scoring, uncertainty-aware AS-Norm (UAS-Norm), and uncertainty-aware Quality Measure Function calibration (UQMF). Covariance information is incorporated throughout this pipeline to adjust score scaling, cohort statistics, normalized-score combination, and calibration features. Experiments with ECAPA-TDNN and ResNet show consistent EER reductions and improved target--non-target separation across both architectures.
Speaker distance estimators are trained almost exclusively on simulated room acoustics, because real recordings annotated with the true talker-to-microphone distance are scarce. We show that models trained this way transfer poorly. On three real corpora we evaluate, simply predicting the average distance of the corpus...
Michael Neri, A. Politis, Tuomas Virtanen· 0 citations
A trainable reliability-aware evidential fusion framework that estimates not only sentiment predictions but also modality-specific evidence, predictive uncertainty, observable input quality, cross-modal disagreement, and normalized sample-dependent fusion weights is presented.
Oracle analysis shows that substantially larger WER reductions are possible through utterance-level scale selection and identifies how beneficial adjustments vary with domain shifts.
While cascaded multi-talker ASR (MT-ASR) leverages state-of-the-art foundation models, its performance is often capped by speaker leakage during separation. Prior correction strategies primarily focus on lexical re-labeling for speaker attribution. We propose a complementary pruning-based paradigm that robustly identif...
H. Y. Nkouanga, Minwei Luo, Maggie B. Wigness et al.· 0 citations
DPQ is introduced, a lightweight pre-quantization recipe family that uses full-precision predictions to construct target-aligned calibration mixtures of high-doubt examples and generic anchors that better preserve broad multiple-choice QA behavior.
Zhen Yang, Sizai Hou, Kai-Wen Zheng et al.· 1 citation
Speaker verification systems encounter combinations of noise, channel distortion, and changes in speech. Evaluating each condition separately does not establish whether their effects add. InterBias-SV organises this question around a four-term comparison: joint error, two marginal errors, and a common reference. Its re...