LLM annotation at scale outperforms human-supervised classifiers at roughly one-tenth the cost, for both a closed-source and an open-weight LLM, and the advantage is robust under soft-label evaluation.
Ahmad Dawar Hakimi, Lea Hirlimann, Isabelle Augenstein et al.· 0 citations
ProbeDrift is introduced, a systematic evaluation framework for supervised uncertainty probes covering a wide range of OOD settings across models, tasks, and distributional shifts, and it is argued that robust uncertainty estimation requires robust evaluation.
Joe Stacey, Hadas Orgad, Kentaro Inui et al.· 0 citations
This work systematically characterize what makes an effective multilingual teacher, and combines intrinsic measures of data quality with extrinsic student model performance in a metric the authors call Polyglot Score, which reveals that model scale alone does not significantly predict teacher effectiveness.
Lester James Validad Miranda, Ivan Vulic, Anna Korhonen· arXiv.org· 1 citation
MedConceal, a benchmark with an interactive patient simulator for evaluating hidden-concern reasoning in medical dialogue, comprising 300 curated cases and 600 clinician-LLM interactions is presented, identifying hidden-concern reasoning under partial observability as a key unresolved challenge for medical dialogue systems.
DLR is proposed, a reinforced latent reasoning framework that dynamically decomposes queries into textual premises, extracts premise-conditioned continuous visual latents, and deduces answers through grounded rationales to enable effective exploration in the latent space.
A systematic analysis of expert routing patterns in MoE models reveals Language Routing Isolation, in which high- and low-resource languages tend to activate largely disjoint expert sets, and proposes RISE, a framework that exploits routing isolation to identify and adapt language-specific expert subnetworks.
Kening Zheng, Wei-Chieh Huang, Jiahao Huo et al.· arXiv.org· 4 citations· ⚡2
GRADE (GRAdient Dynamics for knowlEdge gap detection), which quantifies the knowledge gap via the cross-layer rank ratio of the gradient to that of the corresponding hidden state subspace, motivated by the property of gradients as estimators of the required knowledge updates for a given target.
Yujing Wang, Yuanbang Liang, Yu-Kun Lai et al.· arXiv.org· 1 citation
NeuRIT is proposed, a Neuron-guided Robust Instruction-Tuning framework built on a localization-first perspective that mines context-aware neurons associated with relevant and irrelevant context processing, and uses them as anchors to selectively adapt both the identified neuron groups and the layers in which they concentrate.
Jae Lee, Jaemin Kim, Sumyeong Ahn et al.· 0 citations
CARE is proposed: a multi-stage privacy-compliant agentic reasoning framework in which a proprietary LLM provides guidance by generating structured categories and transitions without accessing sensitive patient data, while a local LLM uses these categories and transitions to support evidence acquisition and final decision-making.
Hao Liu, Wei-En Li, Rui Song et al.· arXiv.org· 0 citations
This work explores a self-supervised framework that encourages models to predict concepts, approximated as sets of semantically equivalent tokens, suggesting that concepts enhance semantic alignment while preserving language modeling quality.
Christine Zhang, Daniel Jurafsky, Sha-Ni Chen· 1 citation· ⚡1
SwiAttn is a novel hybrid transformer that enables dynamic and fine-grained routing between full attention and sliding window attention, and dynamically routes the computation to either a full-attention branch for global information aggregation or a sliding-window branch for efficient local pattern matching.
WSF-ARG+, the first dataset which combines hate speech with check-worthiness information is released, and a novel LLM-in-the-loop framework to facilitate the annotation of check-worthy claims is introduced.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.