Anchored Decoding is proposed, a plug-and-play inference-time method for suppressing verbatim copying that enables decoding from any risky LM trained on mixed-license data by keeping generation in bounded proximity to a permissively trained safe LM.
Jacqueline He, J. Hayase, Wen-tau Yih et al.· arXiv.org· 0 citations
ClinMPO improved performance across two complementary schemes covering ICD-11 diagnostic categories and psychiatric practice competencies and Blinded assessment by three clinicians showed improved rationale quality across CPTS criteria.
Xinxin Lin, Guangxin Dai, Y. Zhong et al.· 1 citation
The first large-scale study of 50 quantized models evaluated on PostTrainingBiasBench, a unified benchmark of 13 closed- and open-ended bias datasets, shows that compression fundamentally alters bias patterns, requiring crucial post-quantization evaluation and interventions to ensure reliability in practice.
Stanley Bryan Z. Hua, Sanae Lotfi, Irene Y. Chen· 4 citations
It is taken that TLMs encode a non-trivial amount of syntactic knowledge, which shows strong performance on formal syntactic phenomena, but weaker and more variable performance on phenomena at the syntax-semantics interface.
It is suggested that demographic conditioning in LLMs is not a cue-invariant category-level parameter but depends fundamentally on how identity is cued, reflecting responses to linguistic signals rather than stable demographic categories.
Manuel Tonneau, Neil K. R. Seghal, Niyati Malhotra et al.· 4 citations
This work introduces MentorQA, the first multilingual dataset and evaluation framework for mentorship-focused question answering from long-form videos, and defines mentorship-focused evaluation dimensions that go beyond factual accuracy, capturing clarity, alignment, and learning value.
Parth Bhalerao, D. D’souza, Rui Guan et al.· arXiv.org· 0 citations
An LLM-based pipeline to automatically annotate longitudinal information in radiology reports is proposed, which outperforms existing annotation solutions, achieving 11.3\% and 5.3\% higher F1-scores for longitudinal information detection and disease tracking, respectively.
Xin-Yi Wang, G. Figueredo, Ruizhe Li et al.· Expert systems with applicat...· 0 citations
This work introduces EMemBench, a programmatic benchmark generator for evaluating long-term episodic memory of agents through interactive games, and evaluates memory agents with strong LMs/VLMs as backbones, using in-context prompting as baselines.
Xinze Li, Zi-Yue Zhu, Siyuan Liu et al.· arXiv.org· 7 citations· ⚡1
LocalNewsQA is introduced, an 18,700-item English-news benchmark that pairs the same question across two locales and scores whether a model actually switches its answer when the locale changes, and pretraining with metadata in MAPLE produces measurable switching and improves accuracy on questions whose correct answer depends on locale.
A. Mukherjee, Ziwei Zhu, Antonios Anastasopoulos· 0 citations
This work proposes CORE-T, a scalable, training-free framework that enriches tables with LLM-generated purpose metadata and pre-computes a lightweight table-compatibility cache, and uses 1.20x fewer total selection tokens than LLM-intensive baselines.
Hassan Soliman, Vivek Gupta, Dan Roth et al.· arXiv.org· 2 citations· ⚡1
CoReflect is introduced, which unifies dialogue simulation and evaluation into an adaptive, iterative process that allows evaluation protocols to adapt alongside the rapidly advancing capabilities of dialogue models.
Mechanistic analyses of attention routing, MLP contributions, and layer-wise probability trajectories reveal an asymmetry: inducing copying is an easy ``reactivation''process that can be triggered at different locations in the input, while restoring recall is a ``suppression''process that is more fragile and strongly tied to object-token interventions.
M. Farahani, Franziska Penzkofer, Richard Johansson· arXiv.org· 1 citation
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.