Skip to content

Category

natural language processing

2,926 papers

#artificial intelligence Review Jul 2026

Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models'training cutoffs.

Emad Alharbi · 0 citations
#artificial intelligence Preprint Jul 2026

Asymmetric Within-Document Predictive Learning for Scientific Document Representation

This work proposes SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations.

You Zuo, Éric de la Clergerie, Benoît Sagot · 0 citations
#artificial intelligence Preprint Jul 2026

MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson's Disease Assessments

Results demonstrate that domain-specialized multi-agent reasoning enables reliable query-driven summarization of structured longitudinal clinical assessment data, and that domain-specialized multi-agent reasoning enables reliable query-driven summarization of structured longitudinal clinical assessment data.

Sana Alamgeera, Denise Goberta, M. Irshad et al. · 0 citations
#artificial intelligence Preprint Jul 2026

Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

The results show that sycophancy can corrupt the reasoning chain independently of the final answer, so answer-level evaluation alone is insufficient, and a failure taxonomy separating reasoning-chain from answer-level sycophancy is introduced, and a complementary sentence-level taxonomy locating where in the chain drift first emerges.

Mahir Numayeer Islam, G. Okuyama, Nikolaus Siauw et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

STAGEET: Stage-wise Typed Edit Tagging for Grammatical Error Correction with Arabic as a Case Study

Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rather than generating a full corrected sentence. Their interpretability, however, is primarily operational: a label specifies how the string should change, but a single edit vocabulary does not always reveal the type of correction being made. We propose STAGEET, a stage-wise typed edit-tagging framework that reorganizes Seq2Edit supervision into typed executable stages and extends edit operations to correction categories. STAGEET decomposes correction into an ordered sequence of medium-grained typed stages; each stage predicts from its own label space, rewrites the current hypothesis once, and passes the resulting intermediate sentence to the next stage. We instantiate the framework as both an end-to-end shared-encoder multi-head model with stage-specific adapters and a fully specialized variant with one independent tagger per stage. Experiments on QALB-2014 and ZAEBUC show that category-aware staged correction retains competitive edit-based GEC performance while exposing a more inspectable correction trajectory, and attains state-of-the-art results on QALB-2014.

Wenjie Lou, Alaa Mamdouh Akef · 0 citations
#artificial intelligence Preprint Jul 2026

Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System

This work curates a syllabus-aligned QA dataset based on NCERT textbooks for classes 9-12, capturing the content, context, and teaching style of Indian curricula, and introduces GurukulAI, an open-access platform that enables Indian students to chat with the model, get doubts cleared, practice exam-style questions, receive contextual answers, and interact in both English and Hindi.

I. Narang, Sneha S. Gosai, Mayank Singh · 0 citations
#artificial intelligence Preprint Jul 2026

Parametric Multimodal User Memory: Storing What Captions Cannot Carry

This work ground perceptual memory in the model, decomposing recall into two subproblems: a vision-language model grounds the referent in context (what and where), and a dedicated encoder extracts an identity key (who), stored as one inline token read by attention at generation with no external round-trip.

Bojie Li, Noah Shi · 0 citations
#artificial intelligence Preprint Jul 2026

NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts

The research here utilizes Natural Language Processing methods like Named Entity Recognition (NER), BERTopic modeling, and Knowledge Graph development in Neo4j to extract, categorize, and visualize important concepts based on translated versions to make ancient Indian medical wisdom more accessible and understandable.

M. Rajeevan, B. Devi, V. Anoop et al. · 2 citations
#machine learning Preprint Aug 2026

GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented Reasoning

This work presents \textsc{GTA-RAG}, a graph-trajectory-augmented RL framework for multi-turn retrieval-augmented reasoning that consistently outperforms RL-based RAG baselines with both Qwen2.5-3B and Qwen2.5-7B backbones, while substantially improving evidence-chain coverage.

Jun Chen, Yongchao Liu, Pengyu Qiu et al. · 0 citations
#machine learning Preprint Aug 2026

Dynamically Allocating Evaluation Effort for Model Ranking

This work formalizes multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms, where pulling an arm corresponds to human-evaluating a model, and proves the optimality of the proposed algorithms and shows that it improves discrimination between top-performing models.

Vilém Zouhar, Julia Kreutzer, A. Lavie et al. · 0 citations

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

WorldCupArena is presented, a dynamic benchmark for language models and deep-research agents that can be reused for future leagues and cups, and shows only small gains in result and exact-score accuracy, but a clearer gain in Scoreline.

Zhaokai Wang, T. Gui, Jiayuan Rao et al. · 1 citation · ⚡1

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.