Skip to content

Category

natural language processing

2,491 papers

#natural language process... Preprint Feb 2026

An evidence-guided reinforcement learning method to improve psychiatric reasoning in small language models

ClinMPO improved performance across two complementary schemes covering ICD-11 diagnostic categories and psychiatric practice competencies and Blinded assessment by three clinicians showed improved rationale quality across CPTS criteria.

Xinxin Lin, Guangxin Dai, Y. Zhong et al. · 1 citation
#natural language process... Preprint Feb 2026

Investigating Social Bias Changes in Quantized Language Models

The first large-scale study of 50 quantized models evaluated on PostTrainingBiasBench, a unified benchmark of 13 closed- and open-ended bias datasets, shows that compression fundamentally alters bias patterns, requiring crucial post-quantization evaluation and interventions to ensure reliability in practice.

Stanley Bryan Z. Hua, Sanae Lotfi, Irene Y. Chen · 4 citations
#artificial intelligence Review Jan 2026

The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models

It is taken that TLMs encode a non-trivial amount of syntactic knowledge, which shows strong performance on formal syntactic phenomena, but weaker and more variable performance on phenomena at the syntax-semantics interface.

Nora Graichen, Iria de-Dios-Flores, Gemma Boleda · 2 citations
#natural language process... Preprint Jan 2026

Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias

It is suggested that demographic conditioning in LLMs is not a cue-invariant category-level parameter but depends fundamentally on how identity is cued, reflecting responses to linguistic signals rather than stable demographic categories.

Manuel Tonneau, Neil K. R. Seghal, Niyati Malhotra et al. · 4 citations

Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content

This work introduces MentorQA, the first multilingual dataset and evaluation framework for mentorship-focused question answering from long-form videos, and defines mentorship-focused evaluation dimensions that go beyond factual accuracy, capturing clarity, alignment, and learning value.

Parth Bhalerao, D. D’souza, Rui Guan et al. · 0 citations
#artificial intelligence Open access Jan 2026

Standardizing Longitudinal Chest X-ray Report Evaluation via Large Language Model Annotation

An LLM-based pipeline to automatically annotate longitudinal information in radiology reports is proposed, which outperforms existing annotation solutions, achieving 11.3\% and 5.3\% higher F1-scores for longitudinal information detection and disease tracking, respectively.

Xin-Yi Wang, G. Figueredo, Ruizhe Li et al. · 0 citations
#computer vision Jan 2026

EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents

This work introduces EMemBench, a programmatic benchmark generator for evaluating long-term episodic memory of agents through interactive games, and evaluates memory agents with strong LMs/VLMs as backbones, using in-context prompting as baselines.

Xinze Li, Zi-Yue Zhu, Siyuan Liu et al. · 7 citations · ⚡1
#natural language process... Preprint Jan 2026

MAPLE: Metadata Conditioned LLM Pretraining for Locale-Aware Question Answering

LocalNewsQA is introduced, an 18,700-item English-news benchmark that pairs the same question across two locales and scores whether a model actually switches its answer when the locale changes, and pretraining with metadata in MAPLE produces measurable switching and improves accuracy on questions whose correct answer depends on locale.

A. Mukherjee, Ziwei Zhu, Antonios Anastasopoulos · 0 citations

CORE-T: COherent REtrieval of Tables for Text-to-SQL

This work proposes CORE-T, a scalable, training-free framework that enriches tables with LLM-generated purpose metadata and pre-computes a lightweight table-compatibility cache, and uses 1.20x fewer total selection tokens than LLM-intensive baselines.

Hassan Soliman, Vivek Gupta, Dan Roth et al. · 2 citations · ⚡1

To Copy or Not to Copy: Copying Is Easier to Induce Than Recall

Mechanistic analyses of attention routing, MLP contributions, and layer-wise probability trajectories reveal an asymmetry: inducing copying is an easy ``reactivation''process that can be triggered at different locations in the input, while restoring recall is a ``suppression''process that is more fragile and strongly tied to object-token interventions.

M. Farahani, Franziska Penzkofer, Richard Johansson · 1 citation
#artificial intelligence Preprint Jan 2026

To Retrieve or To Think? Cross-Boundary Context Evolution for Multi-hop Complex Reasoning

EvoCtx dynamically decides whether the next reasoning transition should cross the current evidence boundary through retrieval or refine the reasoning state within the existing context, and strategically alternates between boundary expansion and intra-boundary trajectory refinement.

Rubing Chen, Jian Wang, Wenjie Li et al. · 2 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.