Evaluating three frontier VLMs in both homogeneous and cross-model adversarial settings, it is found that even the strongest agent hallucinates 15.1% of its verifiable spatial claims and 11.5% of accusations are strictly unsupported.
Ye Yuan, Ruiqi Song, Wei-En Li et al.· arXiv.org· 2 citations
NEI-CAP is introduced, a construction-aware diagnostic protocol for insufficient-evidence evaluation that audits shortcut cues, validates hard cases through human adjudication, and tests whether competence transfers across constructions.
It is argued that a key difference from Python interpreters, web search, and JSON APIs is interface feedback: their failures often leak natural-language signal the model saw in pretraining, and that same-family reward redesigns do not fix.
This work proposes a context-aware evaluation framework in which human-likeness is assessed using a two-sample problem between the linguistic feature distribution of a human reference corpus for a given register and a corresponding LLM-generated corpus.
Björn Nieth, Marianna Gracheva, Michaela Mahlberg et al.· arXiv.org· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
A released verifier system: a three-stage contamination audit that certifies the train/test boundary behind an audited multimodal training corpus and a held-out olympiad benchmark; a binary answer verifier that supplies the reinforcement-learning training reward; and an answer-judging harness that brackets every open-ended score between a deterministic strict layer and a large-language-model liberal layer.
ExACT, a supervision-allocation objective that assigns extra weight to long effective-context targets by inverse frequency within the long tail, supports a supervision-centric thesis: long-context adaptation depends on how strongly training supervises long-context predictions.
Jinchang Zhu, Jindong Li, Chengyu Zou et al.· arXiv.org· 4 citations
The Patent Heterogeneous Attention-Guided Graph Encoder (PHAGE), which constructs a typed claim graph that distinguishes legal citations from technical relations, and fine-tunes the encoder using a dual-granularity contrastive objective that combines inter-patent taxonomy with intra-patent topology.
DRIP-R is introduced, a benchmark that systematically exploits real-world retail policy ambiguities to construct scenarios in which no single correct resolution exists, and shows that frontier models fundamentally disagree on identical policy-ambiguous scenarios, confirming that ambiguity poses a genuine and systematic challenge to LLM decision-making.
Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang et al.· arXiv.org· 0 citations
Perspective API closes at the end of 2026, removing the de facto standard for toxicity measurement and exposing researchers'dependence on a tool they did not control. Drawing on this case, we argue that a research field must build and govern its own measurement infrastructure rather than borrow it. Surveying 241 papers that use or study Perspective, we show what depending on it cost the research community: claims reaching past what the tool could support, results that shifted when its model was silently retrained, and disparities researchers could measure but not explain. These failures were amplified throughout the LLM lifecycle, where Perspective supplied the labels, filtered the corpora, and graded the systems trained on each, rewarding errors rather than catching them. To keep the instrument open to study after shutdown, we release Perspective scores for 5.9 million text snippets from 77 datasets. We further specify ten requirements for measurement infrastructure a field owns, and argue that what blocks such infrastructure is not technical capability but the value the field places on infrastructure work.
David Hartmann, Manuel Tonneau, Angelie Kraft et al.· 0 citations
This work provides a theoretical formalization of the relationship between model calibration and predictive accuracy, and derives the sufficient conditions for greedy decoding optimality, and proposes Greedy Decoding for Reasoning Models, which outperforms both stochastic sampling and standard greedy decoding in multimodal reasoning scenarios.
Bo-Qi Chen, Xudong Liu, Yunke Ao et al.· arXiv.org· 0 citations
Graph2Counsel is introduced, a framework for generating synthetic counseling sessions grounded in Client Psychological Graphs that encode relationships among clients'thoughts, emotions, and behaviors, and Therapy-Eval is introduced, a multi-turn evaluation framework, and the effectiveness of the fine-tuned model in realistic therapeutic conversations is demonstrated.
Aishik Mandal, Hiba Arnaout, Clarissa W. Ong et al.· arXiv.org· 3 citations
This paper introduces KoALa-Bench, a comprehensive benchmark for evaluating Korean speech understanding and speech faithfulness of LALMs, and incorporates listening questions from the Korean college scholastic ability test as well as content covering Korean cultural domains.
Jinyoung Kim, Hyeongsoo Lim, Eunseo Seo et al.· arXiv.org· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.