It is found that while all models show sensitivity to existential presupposition across syntactic embeddings, determiner types and contextual cues, their behaviour differs markedly in strength and systematicity, with NLI-fine-tuned autoregressive models exhibiting the most coherent and stable projection patterns.
Despite great performance on many tasks, language models (LMs) still struggle with reasoning, sometimes providing responses that cannot possibly be true because they stem from logical incoherence. Extending on the arguments of Asher and Bhar (2024), we show that logical incoherencies follow from an LLM’s computation of its internal representations, in particular from an LLM’s failure to take account of the different roles that different expressions may play in determining content. Linguistics and logicians have shown the importance of the fact that logical operators provide a structure on which to compute content recursively. We extend this view of logical tokens to structure at the discursive level with an eye to improving pragmatic reasoning as well as deductive reasoning. The key in reasoning is that these structures introduce
operations over an LM’s latent representations that constrain how they may evolve
. We show how LLMs can leverage those structures.
Nicholas Asher, Swarnadeep Bhar· Topoi· 0 citations
The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. It further argues that prompting interventions cause a Calibration Crisis. We reexamine the benchmark and conclusions and show that it is substantially affected by conceptual and evaluation mis-specifications. We identify three conceptual mis-specifications. In particular, Aspectual Reduction affects the benchmark construction, analysis, experiments, and conclusions. Under a strict NLI standard, 76% of Group A instances do not explicitly rule out culmination. In our native-speaker annotation, 38% of Group A examples and 29% of the Group C examples were judged to permit an alternative interpretation. To control these issues and lexical variation, we construct Lexically Matched Minimal Pairs. At the evaluation level, we formulate event-semantic NLI as a Multi-step Reasoning Problem and assess both intermediate semantic decisions and final predictions. Our results show that models often do not affirm culmination but nevertheless accept the corresponding simple-past hypothesis, a pattern we characterize as Sufficiency Bias. We further show that prompting interventions produce a Decision Shift among labels without reliably improving the underlying semantic understanding and reasoning. Intermediate and oracle-guided analyses identify two additional failure modes: errors in compositional aspectual classification and Surface-form Attraction toward surface-associated answers. Our experiments on Qwen-7B with suitable prompts, GPT-5.4, and Qwen-72B provide initial evidence for the context sensitivity of aspectual classification and suggest that these models can achieve performance comparable to that of human annotators.
This paper examines the syntax of quantificational expressions in Chuj, a Mayan language spoken primarily in Guatemala and Mexico. Building on Royer, Buenrostro & Jenks (2026), we argue that Chuj quantifiers fall into two distinct syntactic categories. Some quantifiers function exclusively as
determiners
, while others belong to the well-established category of
nonverbal predicates
(Grinevald & Peake 2012, Coon 2016, Armstrong 2017, Mateo Toledo 2023). This syntactic split is supported by a range of diagnostics, which we systematically apply to another type of quantifier:
wh
-items. The results indicate that Chuj
wh
-items should be treated exclusively as nonverbal predicates, leading us to reanalyze
wh
-questions as constructions that necessarily involve pseudoclefts. Our novel approach challenges the prevailing view in the Mayanist literature, which has treated
wh
-items as components of the extended nominal domain, and supports a less common one instead (Zavala 1992, Tonhauser 2003, 2007). We demonstrate that our proposal also derives three (seemingly idiosyncratic) properties of Mayan
wh
-questions in a unified and theoretically appealing way: (i) the apparent ban on
wh
-
in situ
(Caponigro, Torrence & Zavala Maldonado 2021, Coon, Baier & Levin 2021), (ii) the ban on multiple
wh
-questions (Caponigro, Torrence & Zavala Maldonado 2021, Coon, Baier & Levin 2021), and (iii) the phenomenon known as pied-piping with inversion (Smith-Stark 1988, Aissen 1996).
Justin Royer, C. Buenrostro, Rodrigo Ranero· Journal of Linguistics· 0 citations
Prior work has shown that LLMs encode the truth of factual propositions along linear directions in activation space. It's unclear how these representations extend to contextual truth: propositions whose truth is determined by in-context evidence rather than world knowledge. We show that LLMs maintain a linear representation of contextual truth that persists across structurally different output policies, even when the output doesn't require the model to determine a proposition's truth, and show causal evidence via steering experiments. Using the transcripts from a collaborative vision-language task that requires two LLMs to maintain a shared common ground, we show that truth representations of a proposition are significantly swayed by partner assertions about that proposition, even when the LLM has enough evidence to determine its truth. We find evidence that propositions near the decision boundary are more susceptible to having their truth shifted through partner assertions. Separating representation from output distinguish two forms of sycophancy that output behavior alone cannot: the model may accommodate a false proposition while continuing to represent it as false, or shift its representation across the boundary. The latter is 2.59x more common when the model agrees by restating the false claim explicitly than when it agrees implicitly.
The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects.
Keane Zhang, Varshini Chinta, Raj Sanjay Shah et al.· 0 citations