It is shown that AI generation leaves a consistent ``stylometric footprint'': a small subset of features, primarily entropy and lexical diversity, consistently separates AI-generated text from human writing across 8 LLMs and 5 domains, while the remaining features depend heavily on the domain and generator.
Abstract
Text generated by large language models (LLMs) has been shown to be stylometrically distinct from human-written text \citep{andreDetectingAIAuthorship2023, shahDetectingUnmaskingAIGenerated2023, oparaStyloAIDistinguishingAIGenerated2024, soto2024fewshot, liLinguisticDifferencesAI2025, selviogluFeatureExtractionAnalysis2025}. But LLMs are increasingly used not only to generate text but also to edit human writing, and it is unclear whether the two leave the same trace. We show that AI generation leaves a consistent ``stylometric footprint'': a small subset of features, primarily entropy and lexical diversity, consistently separates AI-generated text from human writing across 8 LLMs and 5 domains, while the remaining features depend heavily on the domain and generator. AI editing, however, does not reproduce the same footprint. Relative to their human-written sources, AI-edited texts show only a small increase in lexical diversity and a decrease in entropy, rather than the joint increase that characterizes AI generation. Lexical density, which contributes little to generation, instead becomes the dominant editing-associated signal. Stylometric features therefore separate AI-edited text from AI-generated text but are substantially less effective at separating it from human-written text. Our results suggest that ``AI text''is not a single phenomenon: generation and editing leave qualitatively different stylometric traces and should be studied separately.
A common intuition holds that a region's music mirrors the temperament of its people, so that melancholic melodies mark melancholic populations. We test the measurable half of that intuition and reject the inferential half. Using the Essen Folksong Collection, a corpus of thousands of notated folk melodies, we extract real melodic and affect-related features from 2393 deduplicated melodies spanning 16 countries and 7 geographic regions, with the analysis performed on symbolic scores rather than audio. The mode of each melody is computed with a key-finding algorithm rather than read from the file, because the collection's own documentation warns its major and minor labels are unreliable. Cross-country differences in melodic structure are large and highly significant. All 8 tested features differ across countries at p<0.001, with the leap-related features reaching p<10^-90, and China carries a distinctive wide-leap, high-activity signature (arousal composite +1.24 standard deviations, mean absolute interval 2.77 semitones against Germany's 2.17). We then test the inferential half. We correlate the regional musical-affect measures with two published, validated national indices, the World Happiness Report ladder score and the Hofstede individualism index. None of the 6 correlations is significant (0 of 6). The geography of musical affect is real and measurable, but it does not predict how happy or how individualist a population is, and any claim that it does is an ecological fallacy. We release the full extraction and analysis pipeline, and a fail-closed checker re-derives every number in this paper from the data.
When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines. This gap between LLM and human writing has been noted across a variety of domains. LLMs tend to produce"average"writing, while human writing contains more diverse content that covers a broader distribution. Existing work has shown the existence of this distributional"gap", but no work has proposed a systematic way to measure it. Our paper proposes a human-grounded framework that uses the empirical distribution of human writing on a topic to measure the distributional breadth of LLM-generated content on that same topic. We propose two metrics, LLM Coverage (LLM-Cov) and In-Boundary Rate (IBR), that separate the plausibility of LLM content from its distributional breadth. Across ideation and narrative tasks, we find that current LLMs produce plausible but narrow content that concentrates near the center of the human response space. Our framework can enable researchers to better assess the distributional breadth of LLM-authored content, which we term its"cultural reach".
Abstract Journalists are mnemonically dispossessed in the age of large language models. To accept News Corp CEO Robert Thomson’s 2026 redefinition of journalism as an ‘input’ alongside semiconductors and datacentres is to deny the human provenance of what it means to make and to consume news. Licensing deals now being signed between major press groups and AI companies enact a transfer of mnemonic sovereignty from journalists as human ‘agents of memory’ (Zelizer, 2008) to the algorithm, to the machine. Generative AI produces an intention economy, restructuring how information flows through society, with each interaction with AI deepening its understanding not just of what you know, but of what you do not yet know to ask (Fang, 2025). This gives AI systems functional agency, anticipating the curiosities it wants us to formulate, generating a ‘past that never existed’. We are at a tipping point in the battle over journalism’s soul in AI’s seizure of human agency in the making of memory, and in how the production of news becomes infrastructure food. Journalists are caught up in the production of content used to feed AI models and systems, which will shape what people are trying to know, rather than their value being derived from the work of journalists as human agents of memory, of having created a past that they have a stake in. We set out how the battle for journalism’s soul is not yet lost in the emergence of a new memorial front of ‘archival journalism’, institutional glitches, and infrastructural exits.
Laura Guien, Andrew Hoskins· Memory, Mind & Media· 0 citations
This paper examines how generative Artificial Intelligence (AI) and algorithmic culture have reshaped theatre practice in Eldoret, Kenya. Theatre production in the town has for long benefited from a Peke Yangu model, where the producer is the writer, director, technician, designer, marketer and actor in his own production, notes the paper. This model has been in vogue, and is a necessity, in response to a dearth in theatre training and technical infrastructure within Eldoret.
Algorithmic culture helps practitioners read audience preferences by mirroring trends circulating in social media, while generative AI makes it possible to produce scripts, posters, and other promotional materials quicker and at lower cost, allowing small groups to imitate popular forms and compete with larger companies, presumably from bigger cities like Nairobi. Consequently, this paper demonstrates that these gains come with significant costs, biggest among them, reduced production value, and cultural relevance. Using "My Robot Brother", a children’s play recently performed at the Kenya National Theatre, by the 64 Theatre Production house from Eldoret, about a boy and a sentient robot, the paper demonstrates that the use of algorithmic culture and generative AI is not totally misplaced. Younger audiences already familiar with AI respond with both fascination and disappointment, to the performance of the character of the robot. This reveals a meta-modern oscillation between acceptance and critique.
Octavious Onyango· Writing the Arts & Human...· 0 citations
While large language model outputs are frequently analysed as a collective super variety termed"AI language,"this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human idiolects. We analyse two datasets of LLM-generated texts on societal topics: a 2024 corpus of six models (Improta et al. 2024) and a newly generated 2026 corpus using the same prompts featuring six contemporary models. Our findings, utilising computational descriptors and stylometric principal component analysis reveal a generational shift between the style of the 2024 and 2026 cohorts, while demonstrating that each individual model maintains a unique linguistic profile. This multi-layered interplay is illustrated by contraction frequencies, which vary from over 1,200 to over 30,000 per million words within the same cohort of models (2026). Ultimately, we conclude that treating LLM output as idiolectal in nature provides a valuable framework with potential implications for research on variation and change, LLM-generated text detection, forensic linguistics and usage-based approaches to language.
Karolina Rudnicka, Thomas Stephan Juzek· 0 citations
Large Language Models (LLMs) recognise patterns but do not natively track the path of exclusions that a coherent discourse demands. When an input rests on a fabricated authority, a misapplied mechanism, or a surreptitious analogy, an unsteered LLM tends to engage with it as if it were grounded, and to reify the misstep into any structured output it generates. Logic-Augmented Generation (LAG) with POLANYI++, an LLM-steering method that uses heuristics, ontologies and problem solving methods for tacit knowledge extraction, produces an Extended Knowledge Graph (XKG) in OWL2, but inherits the same vulnerability: a sophisticated nonsensical input is reified into the graph alongside the legitimate triples, and is hardly detectable by automated reasoners since the XKG is generated jointly with the wrong assumptions. We introduce DARKSIDE, a coherence auditing method on top of POLANYI++. It formalises the trail as an explicit data structure of accumulated exclusions over discourse time, complemented by a warrant axis that classifies each named referent as Warranted, Unattested, Misattributed or Fabricated, with an escalation rule that pushes the DelegationRiskAssessment to UNSAFE when the fabricated rate is positive or the unsupported rate exceeds a threshold. We evaluate DARKSIDE as a steering layer over a Gemini 3 on BSBench, a 100-item adversarial corpus of sophisticated-sounding nonsense across software engineering, finance, healthcare, physics and law, with Claude Sonnet 4.6 as an independent judge. The empirical evidence supports an architectural claim: when an LLM forward pass is wrapped in an ontology-mediated negative-trail apparatus, the structural pattern-vs-path gap can be partially scaffolded. The XKG functions as the missing memory, and the warrant axis as an epistemic firewall.
A. Gangemi, E. Bottazzi· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.