In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.
Comparing human annotations with outputs from generative large language models and examining their alignment with theoretical models of textuality show how computational approaches can shed light on linguistic theories and, vice versa, how linguistic theories can guide the application of computational resources.
A. Bienati, Mariachiara Pascucci, Jennifer-Carmen Frey et al.· Italian Journal of Computati...· 0 citations
CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models, and argues that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.
Malvina Nissim, Danilo Croce, V. Patti et al.· Italian Journal of Computati...· 0 citations
This paper investigates how human annotators and Large Language Models (LLMs) assign and justify semantic similarity judgments in a Semantic Textual Similarity (STS) task. To this end, we present a new version of SimilEx, the first Italian dataset containing human similarity judgments and natural language explanations for sentence pairs, extended with LLM-generated scores and explanations, enabling a direct comparison between human and LLM behaviour under parallel annotation conditions. Within this framework, we examine the extent to which humans and LLMs align in their perception of sentence similarity. We explore this question from multiple perspectives, including the relationship between sentence-level stylistic features and similarity scores, the consistency of judgments across annotator types, and the alignment of human and LLM explanations. Our findings show that LLMs tend to express more moderate judgments than humans, resulting in higher agreement. At the same time, stylistic features of the evaluated sentences are related to similarity judgments in both groups. As for explanations, humans typically produce shorter, often nominal constructions, reflecting more individually driven strategies for justifying similarity judgments, whereas LLMs generate more canonical sentence structures whose content is also more consistent across models, suggesting that justification is a more subjective process for humans.
Chiara Alzetta, F. Dell’orletta, Chiara Fazzone et al.· Italian Journal of Computati...· 0 citations
This paper investigates the affect heuristic in online argumentative discourse through a corpus-based study combining manual and automatic annotation. Drawing on research in argumentation theory, psychology, and decision science, the study approaches affect not as an external addition to reasoning but as a recurring component of evaluative judgment. The analysis focuses on discussions of climate change and artificial intelligence collected from Reddit and X (formerly Twitter), domains characterised by uncertainty, risk perception, and public controversy. The study employs a bottom-up annotation methodology in which four human annotators identify instances of affect heuristic and related cognitive biases in a corpus of more than 30,000 posts and comments. Inter-annotator agreement is assessed using Fleiss’ κ, Cohen’s κ, Gwet’s AC1, and weighted F1 measures. In a second stage, the same annotation scheme is applied to GPT-4o, treated as a constrained fifth annotator operating within predefined categories and probabilistic classification rules. The results show that the affect heuristic is the most frequent heuristic pattern in the corpus, occurring more often than confirmation bias, availability heuristic, or representativeness heuristic. Although traditional κ coefficients remain low because of category imbalance, agreement measures robust to prevalence effects indicate substantial consistency among annotators. The automatic annotation stage reveals partial alignment between human and model judgments, while also exposing systematic discrepancies in the model’s distribution of categories. A lexical and discursive analysis further demonstrates that affect heuristics do not necessarily manifest through explicit emotion vocabulary. Instead, they frequently appear through evaluative framing, practical reasoning under uncertainty, and subtle stance-taking related to collective action and future-oriented judgment. The findings contribute to empirical research on emotional processes in argumentation and demonstrate how affective reasoning can be operationalised and studied through combined qualitative and computational methods.
Paulina Żelewska, Barbara Konat· Człowiek i społeczeństwo· 0 citations
Large Language Models (LLMs) are increasingly used in Information Retrieval, both within retrieval pipelines and for constructing evaluation resources. Existing studies on using LLMs for IR evaluation, however, focus almost exclusively on English, leaving their applicability to other languages, where evaluation resources are often limited and highly needed, unexplored. We examine the use of LLMs to generate relevance labels for an Arabic test collection (ArTest). Using about 10K relevance labels, generated by three LLMs, and used to order eight automated systems and nine simulated manual systems, we show that agreement on binary labels and system ordering is generally high, with no cases of significant opposite conclusions; however, there are cases of false or missed improvements and noticeable limitations when labelling manual, highly performing systems. These findings align with results reported for English and indicate that LLMs could, with some caveats, offer a viable approach to supporting IR evaluation in languages with limited evaluation resources.
Marwah Alaofi, Fatima Haouari· Annual International ACM SIG...· 0 citations