SPIRE (Scholarly-Primitives-Inspired Research Engine), a multi-agent framework that realizes Scholarly Primitives, a typology of basic humanities scholarship practices, as cooperating roles over a multi-scale close-reading substrate of passages, intra-context graph communities, and cross-context semantic clusters, is introduced.
Yating Pan, Jiajun Zhang, Jun Wang et al.· arXiv.org· 0 citations
This work proposes a two-component framework in which a profile generator summarizes a student's history and a simulator predicts student turns conditioned on the resulting profile, which trains both components with reinforcement learning (RL), yielding profiles optimized for faithful student simulation.
Zhangqi Duan, Shuyan Huang, Alexander Scarlatos et al.· arXiv.org· 0 citations
This work introduces AgentREVEAL, a diagnostic framework for analyzing retrieval-induced safety degradation in LLM agents, and uncovers the Safe Source Paradox, a safety-utility trade-off for retrieval-enabled agents.
Aditya Nawal, Manit Baser, M. Gurusamy· arXiv.org· 0 citations
This work proposes Skill-Conditioned Gated Gated Self-Distillation (SGSD), which formulates skill-based SD as teacher hypothesis validation rather than unconditional imitation, and shows that SGSD consistently improves over GRPO and remains competitive with answer-conditioned OPSD under a weaker PI assumption.
Jiazhe Huang, Xiao Chen, Xiao Luo et al.· arXiv.org· 5 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Reverse Probing is proposed, the first UQ framework specialized for clinical summarization, which estimates token-level uncertainty directly from pre-existing labeled summaries, and reveals that delta energy and neighborhood context are the most consistent predictors across all models.
Dynamic Heterogeneous Character Networks are introduced, which organize long novels into temporally localized heterogeneous graphs that align characters with their textual contexts, and GraphLit is proposed, a self-supervised learning framework that learns rich literary representations through a masked graph autoencoder objective.
Gaspard Michel, Elena V. Epure, Romain Hennequin et al.· arXiv.org· 1 citation
It is argued that progress requires dedicated speech-specific training data and models that genuinely condition on speech, and both text- and speech-based quality estimation metrics on two contrastive datasets targeting gender agreement and prosody fall short.
Maike Zufle, Danni Liu, Vilém Zouhar et al.· arXiv.org· 1 citation
BenGER (Benchmark for German Law), a benchmark and dataset for evaluating LLM systems on subsumption-based legal reasoning in German law, is introduced and 12 contemporary LLM systems are evaluated with a rubric-aligned LLM-as-a-Judge cross-validated against a multi-rater human-grading layer.
Sebastian Nagl, A. Mayrhofer, Martin Heidebach et al.· arXiv.org· 0 citations
This work proposes to use bandwidth-calibrated MIP coupled with Tukey IQR peak-detection to isolate reasoning-crucial tokens at the output layer, and applies a Jaccard stability metric over multi-domain problems to verify if the MIP-identified tokens are reasoning quality-guaranteed.
Leonardo Matthew Yauw, Wei-Bin Kou, Yujiu Yang· arXiv.org· 0 citations
It is found that English-Korean performance gaps vary substantially across models and task families, and that SpokenQA and audio understanding rankings diverge, revealing complementary weaknesses invisible to English-only evaluation.
Haechan Kim, Seung-Jun Chung, Inkyu Park et al.· arXiv.org· 0 citations
Semantic Flow Regularization (SFR), a lightweight auxiliary objective that supervises the backbone with continuous sentence-encoder embeddings of future segments via conditional flow matching, improves output diversity, style fidelity, and response quality over SFT on a large-scale industrial dialogue dataset.
Ke Peng, Feifei Li, Xing Fan et al.· arXiv.org· 0 citations
Results show that self-contamination is a trainable component of the LiC gap, and propose MAIGO, an on-policy self-distillation method that reduces this contamination using history-cleaned references from the model's own policy.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.