Skip to content

Category

natural language processing

2,394 papers

#artificial intelligence Preprint Aug 2025

From Isolation to Alignment: Unified LoRA for Efficient Multi-Task Learning

This work reveals a paradox where a simplified, router-free multi-head model with high inter-head redundancy outperforms complex, diversity-driven baselines and proposes Align-LoRA, a unified and efficient framework that shifts the focus from architectural isolation to representation alignment.

Jinda Liu, Bo Cheng, Yi Chang et al. · 1 citation

Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems

The first fully automated framework that synthesizes high-quality, proof-centric benchmarks from natural language mathematical corpora and a new type of hybrid-formatted questions, named ``$m$-out-of-$n$ multiple judge questions'', specifically designed to enable robust, automatic evaluation while being resilient to guessing and superficial pattern matching inherent in traditional formats are proposed.

Ye-Bo Peng, Zixiang Liu, Yao-Ming Li et al. · 1 citation

Learning Composable Chains-of-Thought

It is found that simply training models on CoT data of atomic tasks leads to limited generalization, but minimally modifying CoT formats of constituent atomic tasks to be composable can lead to improvements.

Fangcong Yin, Zeyu Liu, Liu Leqi et al. · 1 citation
#natural language process... Preprint May 2025

Popular but Wrong: Understanding and Mitigating LLM Overconfidence through Knowledge Popularity

It is shown that popularity-related signals can mitigate overconfidence and improve overall confidence estimation, and incorporate knowledge popularity reduces average confidence on incorrect answers and improves overall confidence estimation.

Shiyu Ni, Keping Bi, Jiafeng Guo et al. · 6 citations

Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning

This paper aims to present a comprehensive overview of this emerging paradigm and establish a systematic taxonomy of latent CoT methods, categorizing them from token-wise horizontal approaches to layer-wise vertical strategies.

Xinghao Chen, Anhao Zhao, He-Ming Xia et al. · 57 citations · ⚡3

Vectors from Larger Language Models Predict Human Reading Time and fMRI Data More Poorly when Dimensionality Expansion is Controlled

The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conjecturing a positive'quality-power'relationship, in which language models'(LMs') fit to psychometric data continues to improve as their ability to predict words in context increases. This is important because it might suggest that elements of LLM architecture reflect the architecture of the human sentence processing faculty, and that any inadequacies in predicting human reading time and brain imaging data may be attributed to insufficient model complexity, which recedes as larger models become available. But recent studies have shown this scaling inverts after a point, as LMs become excessively large and accurate, when information-theoretic surprisal is used as a predictor. Other studies propose the use of entire vectors from differently sized LLMs, still showing positive scaling, casting doubt on the value of surprisal as a predictor, but do not control for dimensionality expansion using untrained LLMs with more than 1.6B parameters. This study evaluates scaling of LLM vector predictors controlled using untrained LLMs with up to 66B parameters. Results show that inverse scaling obtains, and moreover the contribution of trained LMs over corresponding untrained LMs drops to zero at around a few billion parameters on most datasets.

Yi-Chien Lin, Hongao Zhu, William Schuler · 4 citations
#natural language process... Preprint Apr 2025

FLAME: A New Dataset on FLemish Accounts of Momentary Experiences

We introduce FLAME (FLemish Accounts of Momentary Experiences), a corpus of nearly 25,000 personal narratives in Belgian-Dutch (Flemish), collected through experience sampling to support Natural Language Processing (NLP) research on an underrepresented variety. Such everyday narratives are rich in culturally grounded themes, but their informal register and low-resource setting make thematic extraction hard. Comparing K-Means, LDA, and BERTopic, we find that human evaluation favors BERTopic, which produces the most coherent, culturally resonant topics. FLAME, thereby, offers a new resource for studying everyday language use in a low-resource variety.

Ratna Kandala, Niels Vanhasbroeck, Katie Hoemann · 3 citations

HeTGB: A Comprehensive Benchmark for Heterophilic Text-Attributed Graphs

The HeTGB is introduced, a novel benchmark comprising five real-world heterophilic graph datasets from diverse domains, with nodes enriched by extensive textual descriptions that enables systematic evaluation of GNNs, pre-trained language models (PLMs) and co-training methods on the node classification task.

Shujie Li, Yuxia Wu, Chuan Shi et al. · 5 citations · ⚡2

Man Made Language Models? Evaluating LLMs' Perpetuation of Masculine Generics Bias

Instruct-based large language models (LLMs) have been shown to propagate and even amplify gender bias when prompted with contextually constrained instructions (e.g., writing a text from a description or selecting a gendered pronoun). However, little attention has been paid to biases in responses to contextually unconstrained (generic) instructions conveyed by gendered language, particularly masculine generics (MG). MG, found in many gender-marked languages, denote the use of the masculine gender as a supposedly neutral reference to mixed-gender groups or individuals whose gender is unknown or non-binary. Yet, psycholinguistic studies demonstrate that MG are not neutral and systematically induce gender bias. This study investigates how both local and proprietary LLMs are MG-biased when responding to generic prompts in French, examining LLMs'MG bias rates and use of gender-fair language (GFL). We create a 16k+ human noun database from existing lexical resources and evaluate six LLMs on four instruction-response datasets under two conditions: prompts with and without MG. Overall, we find that $\approx$27.57% of LLMs'responses to MG-filtered generic instructions are MG-biased ($\approx$78.55% with MG-containing prompts). Moreover, we find that LLMs rarely use GFL spontaneously. These findings highlight the persistence of MG bias in LLM outputs and models'limited tendency towards GFL strategies.

Enzo Doyen, Amalia Todirascu-Courtier · 5 citations · ⚡1
#natural language process... Preprint Feb 2025

HintEval: An Open-Source Python Toolkit for Hint Generation and Hint Evaluation

HintEval, an open-source Python library for unified hint generation and evaluation, facilitates systematic research on hint-based question answering (QA) in NLP and IR through human studies in which participants assess generated hints and use them to answer questions.

Jamshid Mozafari, Bhawna Piryani, Abdelrahman Abdallah et al. · 4 citations
#natural language process... Preprint Dec 2024

Learning from Many Voices: Literary MT Using Multi-Reference Human and Synthetic Data

This work finds fine-tuning on human expert translations outperforms fine-tuning on synthetically augmented data in automatic metrics and human evaluations, demonstrating the indispensable value of human expert translations for fine-tuning literary machine translation models.

Sijing Wu, J. Wieting, David A. Smith · 7 citations
#artificial intelligence Preprint Dec 2024

Learning Personalized Prompts for Healthcare Guidance

The results show that the proposed personalized prompt learning (PPL) approach produces more personalized healthcare guidance and wins 97 out of 100 comparisons in expert evaluation, demonstrating its potential for broader healthcare applications.

Ruize Shi, Hong Huang, Wei Zhou et al. · 6 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.