Experiments show that knowledge-aligned SFT can reduce factual hallucinations on WildHalu and Biography while largely preserving general capabilities and confirm that SFT targets beyond the base model's knowledge drive hallucination behavior.
AR Becker, Jakob Kemmler, David Thulke et al.· 0 citations
It is observed that the correlation between speakers'L1 distance and ASR error rates yields a systematic effect on English Speech, with its strength varying across datasets and models.
Tingyu Cheng, L. Clemmensen, Sneha Das· 0 citations
TopoCompress is introduced, a training-free and model-agnostic framework that compresses long contexts by selecting coherent semantic spans by selecting coherent semantic spans and achieves performance comparable to the strongest baseline while using a 4x smaller compression budget.
SingProbe is introduced, a lightweight intrinsic runtime guard that directly reuses hidden states produced during LLM inference and operates alongside autoregressive decoding and extends this paradigm to medical generation through SingProbe-Med, which selectively activates risk-directed decoding interventions only when clinically relevant risks emerge.
Singg Team· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Results show that route-supervised frontier selection can improve budgeted search without altering biochemical generation, although performance remains dependent on frontier construction and reaction ranking.
Philippe Meyer, Guillaume Gricourt, T. Duigou et al.· 0 citations
Both warm-rebuilding trained compressions of key-value caches and serving specifically-phrased updates beside a memory, as pasted text or injected cache state, show particular promise for keeping precomputed memories current, the latter as an interim measure between rebuilds.
In these experiments, BiG-SURE improves average abstention AUROC over prior black-box uncertainty estimators, while remaining simple, unsupervised, and applicable to black-box model settings.
It is found that training on the top 20% tokens ranked by GMTS consistently outperforms entropy-based token selection across three reasoning domains and various model sizes, suggesting that GMTS provides a more fine-grained estimate of token contribution for RLVR training.
This work investigates continued pre-training for adapting large language models to Swedish journalism, using a high-quality dataset that is curate from millions of news articles and demonstrates the importance of targeted evaluation in the adaptation process.
Lukas Borggren, Jenny Kunz, Marco Kuhlmann· 0 citations
This work analyzes 6,531 speeches over 200 years of UK parliamentary debate by using large language models to classify a speaker's perspective towards women's suffrage and political representation, as well as analyse sexist speech in parliament from the lens of the Ambivalent Sexism Inventory.
Mohammad Omar Khursheed, Mandira Sawkar, Ashiqur R. KhudaBukhsh· 0 citations
MineAmongUs is introduced, a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action, and ARIA is proposed, a configurable VLM-agent harness that exposes five cognitive-component ablation axes and opens a new path for embodied VLM-agent alignment research.
Jaewoo Ahn, Junseo Kim, Hyunseo Kim et al.· 0 citations
DBloom speeds up generation with an efficient draft model (drafter) that proposes tokens for a target model to verify in one pass, preserving the target's output distribution and recommending the histogram as a preflight check before spending training compute.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.