Evaluation of six frontier models reveals that even the strongest model reaches only 63.1\% pass rate, with many tasks unsolved by any model, highlighting that multilingual coding competence is a distinct and underexplored capability axis.
Yunsu Kim, Kaden Uhlig, Ashwin Purohit et al.· 0 citations
The Prompt Phrases Prediction Network (PPN) is introduced, an encoder-decoder architecture designed to effectively extract keyword prompts embeddings and infuse the prompt embedding into the Prompt-guided KWS encoder by utilizing a Prompt-acoustic Multi-head Cross-attention (MHCA).
G. Xu, Cheng-Fei Li, Xian-Liang Wang et al.· 0 citations
The results indicate that current MLLMs can process visually clear English and structured numeric content, but reliable native Khmer document understanding remains an open challenge.
PAUSE (Pause-And-Update Strategy Editing) is an intervention that exposes an editable adaptation strategy as a human control surface for cultural decisions in long-form story adaptation, a structured artifact that can be inspected, edited, and then projected through downstream character, entity, and chapter-localization stages.
Taaha Kazi, Vasu Sharma, M. Saifullah et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work proposes LPS-TC, a Lightweight Proactive Speech Turn Controller for plug-and-play integration, and introduces a two-tier evaluation scheme that assesses both chunk-level timing precision and turn-level interaction quality under realistic streaming constraints.
Tianrui Pan, Qinglin Zhang, Chong Deng et al.· 0 citations
This study establishes an end-to-end prototype from raw BIM data input, through defect identification, to repair suggestion generation, and establishes an end-to-end prototype to identify and repair various defects in BIM via domain-specific LLMs.
Jia-Rui Lin, Yunzhen Cai, Xiang Ni et al.· 0 citations
This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models'training cutoffs.
This work proposes SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations.
You Zuo, Éric de la Clergerie, Benoît Sagot· 0 citations
The results show that sycophancy can corrupt the reasoning chain independently of the final answer, so answer-level evaluation alone is insufficient, and a failure taxonomy separating reasoning-chain from answer-level sycophancy is introduced, and a complementary sentence-level taxonomy locating where in the chain drift first emerges.
Mahir Numayeer Islam, G. Okuyama, Nikolaus Siauw et al.· 0 citations
Layer-by-layer analysis of GenAI VP dialogue logs can reveal process patterns associated with high rated history taking and support process-focused feedback in medical education.
Xinyu Li, Zijian Li, Mengyu Xia et al.· 0 citations
Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rather than generating a full corrected sentence. Their interpretability, however, is primarily operational: a label specifies how the string should change, but a single edit vocabulary does not always reveal the type of correction being made. We propose STAGEET, a stage-wise typed edit-tagging framework that reorganizes Seq2Edit supervision into typed executable stages and extends edit operations to correction categories. STAGEET decomposes correction into an ordered sequence of medium-grained typed stages; each stage predicts from its own label space, rewrites the current hypothesis once, and passes the resulting intermediate sentence to the next stage. We instantiate the framework as both an end-to-end shared-encoder multi-head model with stage-specific adapters and a fully specialized variant with one independent tagger per stage. Experiments on QALB-2014 and ZAEBUC show that category-aware staged correction retains competitive edit-based GEC performance while exposing a more inspectable correction trajectory, and attains state-of-the-art results on QALB-2014.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.