Skip to content
Preprint

Human-Like Anaphor Resolution in Large Language Models

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects.

Abstract

Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation-model properties, and semantic factors. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models (LLMs) with open weights: GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, and Mistral-24B. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors. The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects. These findings delimit the conditions under which LLMs approximate human anaphor resolution.

View source

Similar papers

Preprint Aug 2026

Reversing Arrows in Large Language Models

This work presents the first systematic study of inverse relation directionality in LLMs, using a benchmark consisting of 5,457 instances spanning 27 distinct inverse relation labels and reveals systematic asymmetries in inverse relation classification across LLMs.

Sefika Efeoglu, A. Paschke · 0 citations

There is No Spoon: Existential Presupposition in Large Language Models

It is found that while all models show sensitivity to existential presupposition across syntactic embeddings, determiner types and contextual cues, their behaviour differs markedly in strength and systematicity, with NLI-fine-tuned autoregressive models exhibiting the most coherent and stable projection patterns.

Marie-Léontine Wörgötter, Shikai Lai, Sebastian Schuster · 0 citations
Book Open access Jul 2026

Attend to Fragments: How Key Information Affects Large Language Models for Factual Inconsistency Detection

A new benchmark, KIFI, is designed, which comprises 1032 carefully selected instances from the TRUE and ScreenEval datasets, with key information annotated, and it is shown that LLMs frequently fail to use the appropriate information to make correct decisions.

Xindi Guo, Zhen Xie, Patrick H. Chen · 0 citations
Open access Jul 2026

Logical Misconceptions, Pragmatic Insufficiencies in LLMs and How to Fix Them

Despite great performance on many tasks, language models (LMs) still struggle with reasoning, sometimes providing responses that cannot possibly be true because they stem from logical incoherence. Extending on the arguments of Asher and Bhar (2024), we show that logical incoherencies follow from an LLM’s computation of its internal representations, in particular from an LLM’s failure to take account of the different roles that different expressions may play in determining content. Linguistics and logicians have shown the importance of the fact that logical operators provide a structure on which to compute content recursively. We extend this view of logical tokens to structure at the discursive level with an eye to improving pragmatic reasoning as well as deductive reasoning. The key in reasoning is that these structures introduce operations over an LM’s latent representations that constrain how they may evolve . We show how LLMs can leverage those structures.

Nicholas Asher, Swarnadeep Bhar · 0 citations
Jul 2026

A Semantic Specification of Expressive Small Clauses in Japanese

This paper develops a semantic analysis of Japanese expressive small clauses (ESCs)—phrases comparable to English You fool!—within Portner et al.’s (2019) participant-structure framework. We analyze ESCs and the nominal particle -me as conventional-implicature operators that update the hierarchical ranking of individuals. An ESC assigns its referent the lowest rank, whereas -me requires only that at least one individual outrank its referent. This entailment relation explains why -me contributes no distinct status-related content within an ESC but remains meaningful in clausal environments. The analysis also unifies the derogatory third-person and humble first-person uses of -me. We further analyze the honorific suffix -sama as an upward operator that assigns its referent the highest rank. This captures both deferential uses such as Tanaka-sama and self-aggrandizing uses such as ore-sama, while allowing multiple individuals to share the highest rank.

Yu Izumi · 0 citations