Emotion-Robust Speaker Recognition Through Synthetic Emotional Spectrogram Augmentation
Speaker recognition models are typically enrolled using neutral speech, yet real users rarely speak under emotionally neutral conditions. Emotion alters prosody, spectral structure, articulation, and speaking rate, shifting utterances away from the acoustic distribution observed during enrolment. Conventional augmentat...