This work curates a syllabus-aligned QA dataset based on NCERT textbooks for classes 9-12, capturing the content, context, and teaching style of Indian curricula, and introduces GurukulAI, an open-access platform that enables Indian students to chat with the model, get doubts cleared, practice exam-style questions, receive contextual answers, and interact in both English and Hindi.
I. Narang, Sneha S. Gosai, Mayank Singh· 0 citations
This work ground perceptual memory in the model, decomposing recall into two subproblems: a vision-language model grounds the referent in context (what and where), and a dedicated encoder extracts an identity key (who), stored as one inline token read by attention at generation with no external round-trip.
The research here utilizes Natural Language Processing methods like Named Entity Recognition (NER), BERTopic modeling, and Knowledge Graph development in Neo4j to extract, categorize, and visualize important concepts based on translated versions to make ancient Indian medical wisdom more accessible and understandable.
M. Rajeevan, B. Devi, V. Anoop et al.· 2 citations
This work presents \textsc{GTA-RAG}, a graph-trajectory-augmented RL framework for multi-turn retrieval-augmented reasoning that consistently outperforms RL-based RAG baselines with both Qwen2.5-3B and Qwen2.5-7B backbones, while substantially improving evidence-chain coverage.
Jun Chen, Yongchao Liu, Pengyu Qiu et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work formalizes multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms, where pulling an arm corresponds to human-evaluating a model, and proves the optimality of the proposed algorithms and shows that it improves discrimination between top-performing models.
Vilém Zouhar, Julia Kreutzer, A. Lavie et al.· 0 citations
WorldCupArena is presented, a dynamic benchmark for language models and deep-research agents that can be reused for future leagues and cups, and shows only small gains in result and exact-score accuracy, but a clearer gain in Scoreline.
Zhaokai Wang, T. Gui, Jiayuan Rao et al.· arXiv.org· 1 citation· ⚡1
This paper proposes MultiHashFormer, a new framework that allows hash-based autoregression that consistently outperforms standard Transformer LMs across multiple benchmarks and shows that the model handles multilingual vocabulary expansion with a constant parameter footprint without any modifications.
It is argued that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning, and the utility discrepancy between a model's deployed reasoning strategy and its rational counterpart whose responses maximise utility in the steepest direction is formalised.
Comprehensive empirical evaluations demonstrate PEAR significantly improves average accuracy over the strongest debate baselines, and theoretically characterize PEAR as an equivariant sparse router: it preserves accuracy under agent relabeling while reducing routing complexity and improving generalization.
Yang Feng, Ziwei Xu, Xia Hu et al.· arXiv.org· 0 citations
Retrospective Harness Optimization is introduced, a self-supervised method that optimizes the agent harness using only past trajectories and alters the agent's behavior patterns and sustains higher accuracy during long-horizon sessions.
Wenbo Pan, Shujie Liu, Chin-Yew Lin et al.· arXiv.org· 8 citations· ⚡1
Linear attention reduces the quadratic cost of softmax attention by maintaining a recurrent fast-weight state, but it consistently lags on in-context retrieval and long-context tasks. Existing remedies act on the write side of memory through gating, delta updates, or kernel feature maps, but the read step is left unchanged: every past key contributes additively to the output, so useful targets are diluted by the bulk of stored vectors. We borrow one specific piece of softmax's geometry to construct a cheap read-time contraction of the query. A second-order Taylor expansion of the softmax log-partition at the isotropic-attention point gives a local quadratic model whose curvature coincides with the running key covariance, a quantity that can be maintained with the same recurrent/chunkwise mechanism as the linear-attention state. The associated linear operator contracts the query along the high-variance directions of memory before it reads the state. We call this mechanism Curvature-Conditioned Query (CCQ). CCQ modifies only the read step and is composable with any linear-attention backbone. Attached to GLA and Gated DeltaNet, it improves perplexity, zero-shot downstream accuracy, S-NIAH retrieval at and beyond the training context, length-extrapolation perplexity from 4K to 20K, and LongBench accuracy.
Dong Le, Thong Nguyen, Cong-Duy Nguyen et al.· 0 citations
This study formalizes Autonomous Agentic Data Engineering, a novel task designed to evaluate LLMs as autonomous data engineers that drive model specialization through end-to-end data curation, and charts a path toward agent-driven model specialization.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.