Skip to content

Euston: Training Away Mathematical Sycophancy Without Losing the Mathematics

Sep 2026 · 0 citations · 13 references
Computer Science

TL;DR

Euston is presented, an 8B mathematical claim-verification model trained to resist the all-false composition of the official evaluation sets and the low precision implied at realistic error prevalence, which is principally the all-false composition of the official evaluation sets.

Abstract

Reasoning language models are trained to produce solutions, not to refuse them, and this bias persists when the problem they are handed is false. Asked to prove a corrupted theorem, a strong model will typically comply and produce a confident derivation of something untrue. We present Euston, an 8B mathematical claim-verification model trained to resist exactly this. Training data were generated with GraphSynth, a probabilistic factor-graph generator that couples attribute-level diversity to decode-time structural masking and span-synchronized verification, yielding 3{,}026 matched true/corrupted statement pairs (6,052 statements) drawn from arXiv papers spanning 2010--2025. We fine-tuned DeepSeek-R1-8B with GRPO under a rule-based, zero-API reward for 189 steps on four H100 GPUs. On a balanced 200-true/200-false held-out split, balanced accuracy rises from 29.50% to 63.75% and the discrimination gap---the difference between the rate of calling false statements false and the rate of calling true statements false moves from -0.5% (z=-0.1) to +27.5% (z=+6.0). Critically, the gain is not purchased with general mathematical ability: AIME 2026 accuracy under official semantics is 65.00% against a 69.17% base, a difference of -4.17% that is not statistically significant, whereas an earlier run of the same recipe on a smaller GraphSynth corpus collapsed to 40.00%. Median response length also falls from 19,217 to 18,296 tokens and the truncation rate from 25.8% to 8.3%, so the improvement does not come from thinking longer. We report the result together with the confounds that bound its interpretation, principally the all-false composition of the official evaluation sets and the low precision implied at realistic error prevalence.

View source

Similar papers

#machine learning Preprint Sep 2026

Beyond Solver Verdicts: Generative Reward Models for Autoformalization

Generative Verification (GenV) is introduced, which distills an offline Z3-equivalence oracle into a reference-free, continuous reference-equivalence score by repurposing the language model's native vocabulary space and theoretically proves that structural, verdict-only verification heuristics are mathematically bounde...

Vikash Singh, Debargha Ganguly, Aman Goel et al. · 0 citations
Preprint Aug 2026

TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint

When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce TRAPSBench, a procedurally generated video benchmark of 1,404 matched physics pairs in which a...

Fnu Pramono, J. Cai, Sourabh Kulkarni · 3 citations · ⚡1
#natural language process... Preprint Sep 2026

Towards Robust Numerical Claim Verification

Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically perturbed examples. Usin...

Peter Røysland Aarnes, Vinay Setty · 0 citations
#natural language process... Preprint Sep 2026

LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models

Language models can now prove theorems, but people still decide which problems to pursue. We ask whether a model's internal representations can help identify promising mathematical connections. We develop LANTERN, a fast, cost-efficient pipeline that uses a classifier over pretrained-model activations to rank candidate...

Pavel Tikhonov, Elena Tutubalina, I. Oseledets et al. · 0 citations
#artificial intelligence Preprint Aug 2026

HSRM: Hidden-State Reward Models for Test-Time Verification

HSRM is introduced, a lightweight hidden-state reward model that verifies candidate solutions by directly reading the generator's internal representations rather than re-processing its text, providing an efficient alternative to text-only verification by reusing representations already computed during generation.

Xianzhi Li, Xiao-Dan Zhu · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.