Skip to content

Category

natural language processing

2,394 papers

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

Evaluating three frontier VLMs in both homogeneous and cross-model adversarial settings, it is found that even the strongest agent hallucinates 15.1% of its verifiable spatial claims and 11.5% of accusations are strictly unsupported.

Ye Yuan, Ruiqi Song, Wei-En Li et al. · 2 citations

Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification

NEI-CAP is introduced, a construction-aware diagnostic protocol for insufficient-evidence evaluation that audits shortcut cues, validates hard cases through human adjudication, and tests whether competence transfers across constructions.

Jing Qiu, Ze-Yu Han, Chen Huang · 0 citations

How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

This work proposes a context-aware evaluation framework in which human-likeness is assessed using a two-sample problem between the linguistic feature distribution of a human reference corpus for a given register and a corresponding LLM-generated corpus.

Björn Nieth, Marianna Gracheva, Michaela Mahlberg et al. · 0 citations
#natural language process... Preprint May 2026

Physics-R1: An Audited Olympiad Corpus and Released Verifiers for Visual Physics Reasoning

A released verifier system: a three-stage contamination audit that certifies the train/test boundary behind an audited multimodal training corpus and a held-out olympiad benchmark; a binary answer verifier that supplies the reinforcement-learning training reward; and an answer-judging harness that brackets every open-ended score between a deterministic strict layer and a large-language-model liberal layer.

Shan Yang · 0 citations

Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing

ExACT, a supervision-allocation objective that assigns extra weight to long effective-context targets by inverse frequency within the long tail, supports a supervision-centric thesis: long-context adaptation depends on how strongly training supervises long-context predictions.

Jinchang Zhu, Jindong Li, Chengyu Zou et al. · 4 citations

Heterogeneous Dependency Graph-Guided Attentionfor Patent Representation Learning

The Patent Heterogeneous Attention-Guided Graph Encoder (PHAGE), which constructs a typed claim graph that distinguishes legal citations from technical relations, and fine-tunes the encoder using a dual-granularity contrastive objective that combines inter-patent taxonomy with intra-patent topology.

Yongmin Yoo, Qiongkai Xu, Zhangkai Wu et al. · 0 citations

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

DRIP-R is introduced, a benchmark that systematically exploits real-world retail policy ambiguities to construct scenarios in which no single correct resolution exists, and shows that frontier models fundamentally disagree on identical policy-ambiguous scenarios, confirming that ambiguity poses a genuine and systematic challenge to LLM decision-making.

Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang et al. · 0 citations

Bye Bye Perspective API: Lessons for Building and Governing Measurement Infrastructure

Perspective API closes at the end of 2026, removing the de facto standard for toxicity measurement and exposing researchers'dependence on a tool they did not control. Drawing on this case, we argue that a research field must build and govern its own measurement infrastructure rather than borrow it. Surveying 241 papers that use or study Perspective, we show what depending on it cost the research community: claims reaching past what the tool could support, results that shifted when its model was silently retrained, and disparities researchers could measure but not explain. These failures were amplified throughout the LLM lifecycle, where Perspective supplied the labels, filtered the corpora, and graded the systems trained on each, rewarding errors rather than catching them. To keep the instrument open to study after shutdown, we release Perspective scores for 5.9 million text snippets from 77 datasets. We further specify ten requirements for measurement infrastructure a field owns, and argue that what blocks such infrastructure is not technical capability but the value the field places on infrastructure work.

David Hartmann, Manuel Tonneau, Angelie Kraft et al. · 0 citations

Revisiting Greedy Decoding for Visual Question Answering: A Calibration Perspective

This work provides a theoretical formalization of the relationship between model calibration and predictive accuracy, and derives the sufficient conditions for greedy decoding optimality, and proposes Greedy Decoding for Reasoning Models, which outperforms both stochastic sampling and standard greedy decoding in multimodal reasoning scenarios.

Bo-Qi Chen, Xudong Liu, Yunke Ao et al. · 0 citations

Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs

Graph2Counsel is introduced, a framework for generating synthetic counseling sessions grounded in Client Psychological Graphs that encode relationships among clients'thoughts, emotions, and behaviors, and Therapy-Eval is introduced, a multi-turn evaluation framework, and the effectiveness of the fine-tuned model in realistic therapeutic conversations is demonstrated.

Aishik Mandal, Hiba Arnaout, Clarissa W. Ong et al. · 3 citations

KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness

This paper introduces KoALa-Bench, a comprehensive benchmark for evaluating Korean speech understanding and speech faithfulness of LALMs, and incorporates listening questions from the Korean college scholastic ability test as well as content covering Korean cultural domains.

Jinyoung Kim, Hyeongsoo Lim, Eunseo Seo et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.