Comparing Human and Large Language Model Perceptions of Semantic Similarity: Insights from the SimilEx Dataset
Abstract
This paper investigates how human annotators and Large Language Models (LLMs) assign and justify semantic similarity judgments in a Semantic Textual Similarity (STS) task. To this end, we present a new version of SimilEx, the first Italian dataset containing human similarity judgments and natural language explanations for sentence pairs, extended with LLM-generated scores and explanations, enabling a direct comparison between human and LLM behaviour under parallel annotation conditions. Within this framework, we examine the extent to which humans and LLMs align in their perception of sentence similarity. We explore this question from multiple perspectives, including the relationship between sentence-level stylistic features and similarity scores, the consistency of judgments across annotator types, and the alignment of human and LLM explanations. Our findings show that LLMs tend to express more moderate judgments than humans, resulting in higher agreement. At the same time, stylistic features of the evaluated sentences are related to similarity judgments in both groups. As for explanations, humans typically produce shorter, often nominal constructions, reflecting more individually driven strategies for justifying similarity judgments, whereas LLMs generate more canonical sentence structures whose content is also more consistent across models, suggesting that justification is a more subjective process for humans.