2025· Italian Journal of Computational Linguistics· Vol 11, pp. 35· 0 citations
TL;DR
Comparing human annotations with outputs from generative large language models and examining their alignment with theoretical models of textuality show how computational approaches can shed light on linguistic theories and, vice versa, how linguistic theories can guide the application of computational resources.
Abstract
This article presents a study on annotating explicit discourse relations in Italian student essays, comparing human annotations with outputs from generative large language models and examining their alignment with theoretical models of textuality. We review prior work on automatic discourse relation annotation in Italian, highlighting limitations in language coverage, especially in out-of-domain scenarios and how these have been addressed. Our experiments explore the use of generative models to mitigate the scarcity of domain-specific training data, while assessing their ability to reflect the intended theoretical framework. We evaluate two generative models in detecting connectives and classifying their senses, comparing results to human annotation. For our evaluation sample, we use a string-matching algorithm combined with a rule-based approach to pre-annotate essays with possible connective forms and their senses, based on their presence in the Lexicon of Italian Connectives (LICO). These annotations were manually corrected by two expert annotators, resulting in a publicly available evaluation sample. The study raises significant theoretical questions about the definition of connectives, its relationship to text segmentation and the challenges both human and machines face when annotating discourse relations. Our findings show how computational approaches can shed light on linguistic theories and, vice versa, how linguistic theories can guide the application of computational resources.
In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.
Mirela Imamović, Aenne Knierim, Khushi Pitroda et al.· 0 citations
This paper investigates the affect heuristic in online argumentative discourse through a corpus-based study combining manual and automatic annotation. Drawing on research in argumentation theory, psychology, and decision science, the study approaches affect not as an external addition to reasoning but as a recurring component of evaluative judgment. The analysis focuses on discussions of climate change and artificial intelligence collected from Reddit and X (formerly Twitter), domains characterised by uncertainty, risk perception, and public controversy. The study employs a bottom-up annotation methodology in which four human annotators identify instances of affect heuristic and related cognitive biases in a corpus of more than 30,000 posts and comments. Inter-annotator agreement is assessed using Fleiss’ κ, Cohen’s κ, Gwet’s AC1, and weighted F1 measures. In a second stage, the same annotation scheme is applied to GPT-4o, treated as a constrained fifth annotator operating within predefined categories and probabilistic classification rules. The results show that the affect heuristic is the most frequent heuristic pattern in the corpus, occurring more often than confirmation bias, availability heuristic, or representativeness heuristic. Although traditional κ coefficients remain low because of category imbalance, agreement measures robust to prevalence effects indicate substantial consistency among annotators. The automatic annotation stage reveals partial alignment between human and model judgments, while also exposing systematic discrepancies in the model’s distribution of categories. A lexical and discursive analysis further demonstrates that affect heuristics do not necessarily manifest through explicit emotion vocabulary. Instead, they frequently appear through evaluative framing, practical reasoning under uncertainty, and subtle stance-taking related to collective action and future-oriented judgment. The findings contribute to empirical research on emotional processes in argumentation and demonstrate how affective reasoning can be operationalised and studied through combined qualitative and computational methods.
Paulina Żelewska, Barbara Konat· Człowiek i społeczeństwo· 0 citations
In argument mining research, educational texts such as student essays have been a popular target genre from the beginning. The annotated corpora available so far, however, focus on essays written by older (or adult) students with relatively high proficiency levels. In our work, we expand the range of texts to German essays by school students in grade 9, which display very different qualities on all levels of analysis. We show that a common approach to representing argument structure as trees is not sufficient to capture the constellations of argument components in those essays, and we propose a suitable extension of that scheme, which we applied to an initial corpus of 50 essays. Furthermore, we conducted experiments with large language models on a fine-grained version of the argument component type classification task and show that medium-sized open-source models can achieve promising classification results using only two labeled essays in a few-shot prompting approach.
Xiaoyu Bai, Kemal Afzal, Dietmar Benndorf et al.· Argument & Computation· 1 citation
Abstract Generative AI tools have implications for corpus linguistics from practical and theoretical angles, both of which the current paper addresses. The first part of the paper establishes whether and how generative AI tools – here ChatGPT o3 – can in practice support a corpus-linguistic workflow comprising data extraction, cleaning, and annotation for the dative alternation. While certain steps in this workflow such as data extraction or object length annotation can be completed reliably in ChatGPT o3, more taxing processes like cleaning extracted data for relevant instances, identifying clause boundaries and clause elements, or annotating objects for animacy are marked by errors and inconsistencies, which even the addition of steps dedicated to facilitating AI support in the corpus-linguistic workflow cannot remedy. The second part of the paper relates to corpus-linguistic theory by discussing whether and how AI-generated texts should be represented in linguistic corpora. It argues that the notions of authenticity and representativeness as key features of linguistic corpora are compatible with the integration of AI-generated texts in linguistic corpora. As said integration marks a departure from corpus-linguistic tradition, various arguments in favour of and against featuring AI-generated texts are provided, yielding a pragmatic practical suggestion for future corpus projects.
This study investigates grammatical trends in texts generated by artificial intelligence and human learners. The study puts to the test a fundamental principle of usage-based grammar: language is learned through repeated exposure to patterns. A direct comparison is conducted between AI-generated writings and language learners' essays. Quantitative approaches count words, sentences, and grammatical errors. Qualitative analysis detects trends in sentence structure and specific qualities such as past tense. Finding out if AI models adhere to usage-based grammar rules is the aim. Comparing the two groups' mistake types is another objective. The results show that whereas human writing varies, AI output is very constant. Almost no grammatical errors were found in AI articles, according to the study. Expected errors in human texts include omissions and overgeneralizations. The findings also demonstrate that AI makes greater use of components like the past tense and plurals. These studies demonstrate that the outcomes of usage-based learning are operationally replicated by AI. The results of training the model on massive amounts of data are consistent and precise. The ongoing process of language acquisition is reflected in human output. The study comes to the conclusion that AI is a powerful instrument for confirming frequency-based linguistic theory.but does not model the human cognitive journey. Future research should investigate different AI models and learner proficiency levels
Assis. lect. Batool Abdul-Mohsin Miri· Journal of College of Educat...· 0 citations