Skip to content

Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

This work analyzes three public corpora: CoAuthor (1,447 keystroke-level co-writing sessions), RealHumanEval (editor telemetry from 243 programmer records), and a pre-LLM CS1 corpus as a human-only baseline, comparing minimal-AI work, collaborative AI use, and simulated wholesale delegation.

Abstract

The rapid adoption of generative AI has made final artifacts unreliable evidence of student learning, and AI detectors that examine only the finished product are inaccurate and ethically contentious. Process data offers an alternative, but prior work covers only English essay writing. We ask whether AI assistance carries a temporal signature, whether it generalizes from writing to programming, and whether it distinguishes ordinary collaboration from wholesale delegation. We analyze three public corpora: CoAuthor (1,447 keystroke-level co-writing sessions), RealHumanEval (editor telemetry from 243 programmer records), and a pre-LLM CS1 corpus (5.1 million keystrokes) as a human-only baseline, comparing minimal-AI work, collaborative AI use, and simulated wholesale delegation. Three findings emerge. First, the signature generalizes: AI contributions arrive in bursts far outside the author's own baseline in both mediums (paired d_z = 1.13 and 3.54). Second, engagement diverges by medium: 93% of AI-inserted characters survived to writers'final documents, while only 14% of accepted code suggestions survived intact. Third, classifiers using only observable temporal features separate simulated delegation from authentic work nearly perfectly (F1 $\geq$ 0.997; at most 0.5% of real work misclassified), while ordinary collaboration remains hard to distinguish from unassisted work. Temporal evidence flags wholesale delegation rather than assistance, positioning process visibility as a candidate evidentiary basis for academic integrity, pending validation in authentic coursework.

View source

Similar papers

Preprint Sep 2026

Who Wrote This? Turing, Total Variation, and the Mathematics of AI-Text Detection

What can a finished text reveal about the process that produced it? Drawing on Turing's imitation game and statistical decision theory, this article examines the limits of AI-text detection as an inference from a completed object to an unobserved history. For two known source distributions with equal prior probabilitie...

S. Schnell · 0 citations
Preprint Sep 2026

AI-Research Agents in the Wild. From GitHub and arXiv to Regularities and Gaps

AI-research agents, or autoresearch systems, combine language models with tools, search, evaluation, and iterative modification of research artifacts. Their public software ecology is hard to compare because repositories, papers, benchmarks, libraries, and companion artifacts are often counted as one population. We con...

Aleksey Komissarov, A. Ustyuzhanin · 0 citations
#generative ai Review Open access Sep 2026

The pitfalls of AI detection in academic writing: bias, false positives, and the need for inclusive assessment

AI-text detectors are increasingly used to police authorship in student assessment and scholarly publishing. This paper argues that they are the wrong instrument for that task. Its design combines a critical synthesis of empirical, information-theoretic, and policy evidence with an original documented exploratory multi...

Victor Angelier · 0 citations
#machine learning Preprint Sep 2026

When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora

Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option is therefore genuine human prose; every distractor is synthetic. Correctness and provenance become entangled, and a policy could in principl...

Esther Xin · 0 citations
#artificial intelligence Preprint Sep 2026

Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase

We present: (i) a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI, with no human-authored code or tests, (ii) two code-provenance tracing tools, (iii) three taxonomies for instruction intent, commit provenance, and response reliability, (iv) application of...

D. Leith · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.