Comparing human and nonhuman judgments of perceived similarity of gender diverse talkers
Abstract
Listeners conduct fine-grained analyses of gender from speech, and their perceptual organization of voices is presumed to be grounded in structured acoustic-phonetic information. Here, we compare human perceived similarity to automated self-supervised speech representations from a deep-learning model (i.e., HuBERT), which provides pairwise distance scores that increase with divergence in acoustic representations. Comparing human similarity judgments with HuBERT-derived distances tests the extent to which listeners’ perceptual organization is driven by acoustic representations versus socially- or task-driven constraints. Stimuli consisted of one sentence produced by 20 talkers representing five gender identities (cisgender man, cisgender woman, transgender man, transgender woman, nonbinary). In a free classification task, listeners (29 cisgender and 29 gender diverse) first grouped talkers by general similarity and then by perceived gender identity. Across all listeners and tasks, higher perceived similarity was associated with smaller HuBERT distances. This relationship was stronger in the unconstrained than the constrained task. Reliable alignment between human and nonhuman similarity metrics suggests that holistic acoustic distance partially drives how listeners classify talkers. However, human listeners also likely draw on knowledge of social constructs that automated similarity algorithms fail to capture, especially when the social category is invoked by task instructions.