Skip to content

Category

natural language processing

2,394 papers

Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship

SPIRE (Scholarly-Primitives-Inspired Research Engine), a multi-agent framework that realizes Scholarly Primitives, a typology of basic humanities scholarship practices, as cooperating roles over a multi-scale close-reading substrate of passages, intra-context graph communities, and cross-context semantic clusters, is introduced.

Yating Pan, Jiajun Zhang, Jun Wang et al. · 0 citations

Who Am I? History-Aware Profiles for Student Simulation in Tutoring Dialogues

This work proposes a two-component framework in which a profile generator summarizes a student's history and a simulator predicts student turns conditioned on the resulting profile, which trains both components with reinforcement learning (RL), yielding profiles optimized for faithful student simulation.

Zhangqi Duan, Shuyan Huang, Alexander Scarlatos et al. · 0 citations

Skill-Conditioned Gated Self-Distillation for LLM Reasoning

This work proposes Skill-Conditioned Gated Gated Self-Distillation (SGSD), which formulates skill-based SD as teacher hypothesis validation rather than unconditional imitation, and shows that SGSD consistently improves over GRPO and remains competitive with answer-conditioned OPSD under a weaker PI assumption.

Jiazhe Huang, Xiao Chen, Xiao Luo et al. · 5 citations

Reverse Probing: Supervised Token-level Uncertainty Quantification for Large Language Models in Clinical Text

Reverse Probing is proposed, the first UQ framework specialized for clinical summarization, which estimates token-level uncertainty directly from pre-existing labeled summaries, and reveals that delta energy and neighborhood context are the most consistent predictors across all models.

Bushi Xiao, Sarvesh Soni, Daisy Zhe Wang · 0 citations

GraphLit: Learning Text-Enriched Dynamic Character Network Representations for Literary Study

Dynamic Heterogeneous Character Networks are introduced, which organize long novels into temporally localized heterogeneous graphs that align characters with their textual contexts, and GraphLit is proposed, a self-supervised learning framework that learns rich literary representations through a masked graph autoencoder objective.

Gaspard Michel, Elena V. Epure, Romain Hennequin et al. · 1 citation

Why We Need Speech to Evaluate Speech Translation

It is argued that progress requires dedicated speech-specific training data and models that genuinely condition on speech, and both text- and speech-based quality estimation metrics on two contrastive datasets targeting gender agreement and prosody fall short.

Maike Zufle, Danni Liu, Vilém Zouhar et al. · 1 citation
#artificial intelligence Review May 2026

BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law

BenGER (Benchmark for German Law), a benchmark and dataset for evaluating LLM systems on subsumption-based legal reasoning in German law, is introduced and 12 contemporary LLM systems are evaluated with a rubric-aligned LLM-as-a-Judge cross-validated against a multi-rater human-grading layer.

Sebastian Nagl, A. Mayrhofer, Martin Heidebach et al. · 0 citations

Integrated and Cross-Architecture Interpretation of LLM Reasoning

This work proposes to use bandwidth-calibrated MIP coupled with Tukey IQR peak-detection to isolate reasoning-crucial tokens at the output layer, and applies a Jaccard stability metric over multi-domain problems to verify if the MIP-identified tokens are reasoning quality-guaranteed.

Leonardo Matthew Yauw, Wei-Bin Kou, Yujiu Yang · 0 citations

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

It is found that English-Korean performance gaps vary substantially across models and task families, and that SpokenQA and audio understanding rankings diverge, revealing complementary weaknesses invisible to English-only evaluation.

Haechan Kim, Seung-Jun Chung, Inkyu Park et al. · 0 citations

Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses

Semantic Flow Regularization (SFR), a lightweight auxiliary objective that supervises the backbone with continuous sentence-encoder embeddings of future segments via conditional flow matching, improves output diversity, style fidelity, and response quality over SFT on a large-scale industrial dialogue dataset.

Ke Peng, Feifei Li, Xing Fan et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.