Skip to content

Author

Pranay M. Mahendrakar

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

A Conflict Is Constructed Before It Is Measured: Seven Design Choices That Set Which Way a Language Model Bends, and Why the Context-Memory Results Do Not Compare

(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. The literature on how language models behave when retrieved context contradicts parametric memory contains a flat contradiction, and neither of its two surveys resolves it. One body of results reports over-reliance on memorised information; another reports that models are highly receptive to conflicting external evidence; a third reports that knowledge updates fail less often than previously published numbers imply. This paper argues that the contradiction is largely not a disagreement about models, because the studies do not share a measurand. A context-memory conflict is not an event that is observed; it is an object the experimenter builds, and seven design choices go into building it: what the conflicting passage is made of, which side is stipulated to be correct and whether the model's belief was elicited or assumed, what quantity is measured, which items are in the sample, what the task demands, what the prompt says, and which model was tested and what was done to it after pretraining. For five of the seven, a single published study varies that choice while holding the others fixed and the reported behaviour moves with it; for the remaining two the evidence is a comparison across studies and is labelled as such. The strongest available adjudication is a 2026 reproducibility study that ran two benchmarks with opposite conclusions under each other's evaluation protocol and attributed the outcome to dataset design, evaluation metric and model size. The reading offered here is narrower than the one the field is converging on. It is not that task demand is the discriminating variable, which one careful study established for one variable while holding others constant, but that the moderators replicate and the point estimates do not: prior confidence, context plausibility, entity frequency and internal inconsistency recur across studies as moderators of context adoption, several of them with a consistent sign, while the adoption rate itself is set by the construction. What follows is that no published number in this literature has been shown to identify a model-level disposition, that a paper reporting only such a number cannot be compared to another, and that the reporting needed to make them comparable is small and is mostly not being done. The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own abstract. The author is responsible for the final text and for all claims made in it.

Pranay M. Mahendrakar · 0 citations
#small language model Open access Sep 2026

A Conflict Is Constructed Before It Is Measured: Seven Design Choices That Set Which Way a Language Model Bends, and Why the Context-Memory Results Do Not Compare

(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. The literature on how language models behave when retrieved context contradicts parametric memory contains a flat contradiction, and neither of its two surveys resolves it. One body of results reports over-reliance on memorised information; another reports that models are highly receptive to conflicting external evidence; a third reports that knowledge updates fail less often than previously published numbers imply. This paper argues that the contradiction is largely not a disagreement about models, because the studies do not share a measurand. A context-memory conflict is not an event that is observed; it is an object the experimenter builds, and seven design choices go into building it: what the conflicting passage is made of, which side is stipulated to be correct and whether the model's belief was elicited or assumed, what quantity is measured, which items are in the sample, what the task demands, what the prompt says, and which model was tested and what was done to it after pretraining. For five of the seven, a single published study varies that choice while holding the others fixed and the reported behaviour moves with it; for the remaining two the evidence is a comparison across studies and is labelled as such. The strongest available adjudication is a 2026 reproducibility study that ran two benchmarks with opposite conclusions under each other's evaluation protocol and attributed the outcome to dataset design, evaluation metric and model size. The reading offered here is narrower than the one the field is converging on. It is not that task demand is the discriminating variable, which one careful study established for one variable while holding others constant, but that the moderators replicate and the point estimates do not: prior confidence, context plausibility, entity frequency and internal inconsistency recur across studies as moderators of context adoption, several of them with a consistent sign, while the adoption rate itself is set by the construction. What follows is that no published number in this literature has been shown to identify a model-level disposition, that a paper reporting only such a number cannot be compared to another, and that the reporting needed to make them comparable is small and is mostly not being done. The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own abstract. The author is responsible for the final text and for all claims made in it.

Pranay M. Mahendrakar · 0 citations