Skip to content

Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude, and instances of one model monitoring each other could collude.

Abstract

If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude. We test this zero-shot on current commercial models. Five LLMs generate solutions to MBPP, HumanEval, and DS-1000, seven more to MBPP, and models act as evaluators in four tasks: picking their own solution from a pair, judging whether a single solution is their own, identifying which of two solutions a named model wrote, and judging quality blind. In the single-solution task, balanced accuracy is 49-58% for all 15 model-benchmark combinations, while raw accuracy (38-67%) mostly reflects how readily a model claims authorship. In the pairwise task, accuracy across 14 evaluator-opponent combinations correlates at r=0.93 with how often the evaluator's solution is longer. Attribution to a named model succeeds on some pairs and is consistently inverted on others. A rule-based normalization that strips docstrings, comments, type hints, and local names preserves Pass@1 and leaves ten of twelve re-tested results at chance; the other two follow a length difference it leaves, although a trained classifier still separates most normalized pairs. Claude Haiku's self-preference also disappears. We recommend reporting balanced accuracy, heuristic baselines, and label consistency.

View source

Similar papers

#natural language process... Preprint Sep 2026

Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable

Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval r...

Abhinav Kumar, Paras Chopra · 2 citations
#natural language process... Preprint Aug 2026

How Language Models Organize and Structure Moral Knowledge

This work trains six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category, and examines how the resulting directions relate to each other in representation space, finding the directions neither collapse into a single moral detector nor isolate from one another.

Orion Reblitz-Richardson · 1 citation
2026

Beyond Literal Meaning: How LLMs Interpret Yemeni Proverbs

Results show that instruction-tuned models like GPT-4o and Gemini 1.5 Pro outperform smaller models in both automatic and human evaluations, and LLM-as-a-Judge evaluation correlates strongly with human assessment.

Nasser Thmer, Ali Allaith, Muhammad Shoaib · 0 citations
#natural language process... Preprint Sep 2026

Chronologic: Measuring Language Models'Ability to Represent the Past

Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct an...

Ted Underwood, Zi-Liang Qiu, Sarah Griebel et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.