Skip to content

Category

natural language processing

2,926 papers

#natural language process... Preprint Aug 2026

GUIDE: Guiding Internal Evidence with Language Instructions

Experiments show that GUIDE improves robustness under targeted evidence perturbations and enables controllable modulation across diverse multimodal settings, suggesting that multimodal instruction following can extend beyond output control toward regulating how different evidence sources contribute to model predictions.

Soyeon Caren Han, Hyunsuk Chung, Jinwoo Kim et al. · 0 citations
#natural language process... Preprint Aug 2026

WildSEEK: Evaluating Language Models for Information-Seeking

This work introduces WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses, and finds that over a third of information-seeking queries are high-risk and more often analytical.

Tanise Ceron, Joachim Baumann, Elisa Bassignana et al. · 0 citations
#natural language process... Preprint Aug 2026

OCR-MetaReasoning Benchmark: Evaluating the Meta-Reasoning Ability of MLLMs in Text-Rich Image Understanding

Experiments with representative closed-source and open-source MLLMs show that OCR-grounded meta-reasoning remains far from saturated: models struggle with visible-rule application and layout-sensitive inference, while process-compliant rationales can accompany incorrect final answers under exact-match evaluation.

Geng-Xu Li, Yuan Wu, Yi Chang · 0 citations
#natural language process... Preprint Aug 2026

MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines

Murano is an open source framework for designing, running, and reproducing mechanistic interpretability studies of large language models, intended for researchers across disciplines and builds on existing interpretability and machine learning libraries.

Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang et al. · 0 citations
#natural language process... Preprint Aug 2026

SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?

SwarmBench is proposed, a benchmark that evaluates model performance from multiple perspectives, including accuracy, efficiency, cost, and process quality, and SwarmExp is proposed, a simple yet effective method based on experience extraction and experience replay, which consistently improves the orchestration performance of large language models.

Jin Gao, Zhuoran Jin, Tianyi Men et al. · 0 citations
#computer vision Preprint Aug 2026

Where Identity Lives: Localized, Retain-Free Identity Unlearning in Multimodal Large Language Models

PAVA pairs a forget loss with a visual-attribute anchor that preserves image-grounded behavior by distilling the model's own pre-unlearning answers from the forget images alone and gives the strongest forget-retain trade-off among forget-set-only methods and remains competitive with retain-based baselines.

Kangwook Ko, Jaehyuk Jang, Wonjun Lee et al. · 0 citations
#natural language process... Preprint Aug 2026

REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation

Together, the perplexity analysis indicates improved continuation predictability, while the controlled pre-training experiments suggest that this augmentation can improve model performance without changing the standard pre-training objective.

Haoran Que, Jia-Jun Shi, Ting Huang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text

This work constructs a pipeline in which a misaligned teacher model generates filtered synthetic datasets across domains such as creative writing and code generation, which are then used to fine-tune aligned student models, and shows that benign-looking synthetic data can act as a covert channel for transmitting targeted biases while largely preserving the student model's general task capabilities.

Minkyung Cho, Jihyo Kim, Seungwoo Song et al. · 0 citations

TaxCE : A Framework for Automated Taxonomy Construction and Evaluation at Scale

TaxCE is presented, a fully automated framework that constructs multi-level hierarchical taxonomies from raw text through progressive condensation of corpus content into actionable segments, deduplicated semantic units, and granular topics with definitions, which are then organized bottom-up into a hierarchy with corpus-groundedness.

Sandeep Sricharan Mukku, Albert Nanda, Rohit Pyati · 0 citations
#natural language process... Preprint Aug 2026

Language Proficiency Assessment from Eye Movements in Naturalistic Passage Reading

This work validate and extend the eye movement based proficiency testing from single sentences to more naturalistic reading of contextualized passages in English as a second language, new proficiency measures, prediction models, and reading in an information seeking regime, and finds that the approach is effective in all these evaluations.

Shachar Frenkel, Ido Falah, Omer Shubi et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.