Experiments show that GUIDE improves robustness under targeted evidence perturbations and enables controllable modulation across diverse multimodal settings, suggesting that multimodal instruction following can extend beyond output control toward regulating how different evidence sources contribute to model predictions.
Soyeon Caren Han, Hyunsuk Chung, Jinwoo Kim et al.· 0 citations
This work introduces WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses, and finds that over a third of information-seeking queries are high-risk and more often analytical.
Tanise Ceron, Joachim Baumann, Elisa Bassignana et al.· 0 citations
This work proposes Long Chain-of-Thought Graph Verifier (LCoT-GV), a graph-based framework that represents LCoTs as reasoning graphs, each node in the graph represents a reasoning step and the edges encode semantic and logical relations.
Experiments with representative closed-source and open-source MLLMs show that OCR-grounded meta-reasoning remains far from saturated: models struggle with visible-rule application and layout-sensitive inference, while process-compliant rationales can accompany incorrect final answers under exact-match evaluation.
Geng-Xu Li, Yuan Wu, Yi Chang· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Murano is an open source framework for designing, running, and reproducing mechanistic interpretability studies of large language models, intended for researchers across disciplines and builds on existing interpretability and machine learning libraries.
Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang et al.· 0 citations
SwarmBench is proposed, a benchmark that evaluates model performance from multiple perspectives, including accuracy, efficiency, cost, and process quality, and SwarmExp is proposed, a simple yet effective method based on experience extraction and experience replay, which consistently improves the orchestration performance of large language models.
Jin Gao, Zhuoran Jin, Tianyi Men et al.· 0 citations
PAVA pairs a forget loss with a visual-attribute anchor that preserves image-grounded behavior by distilling the model's own pre-unlearning answers from the forget images alone and gives the strongest forget-retain trade-off among forget-set-only methods and remains competitive with retain-based baselines.
Kangwook Ko, Jaehyuk Jang, Wonjun Lee et al.· 0 citations
Together, the perplexity analysis indicates improved continuation predictability, while the controlled pre-training experiments suggest that this augmentation can improve model performance without changing the standard pre-training objective.
Haoran Que, Jia-Jun Shi, Ting Huang et al.· 0 citations
This work constructs a pipeline in which a misaligned teacher model generates filtered synthetic datasets across domains such as creative writing and code generation, which are then used to fine-tune aligned student models, and shows that benign-looking synthetic data can act as a covert channel for transmitting targeted biases while largely preserving the student model's general task capabilities.
Minkyung Cho, Jihyo Kim, Seungwoo Song et al.· 0 citations
TaxCE is presented, a fully automated framework that constructs multi-level hierarchical taxonomies from raw text through progressive condensation of corpus content into actionable segments, deduplicated semantic units, and granular topics with definitions, which are then organized bottom-up into a hierarchy with corpus-groundedness.
Sandeep Sricharan Mukku, Albert Nanda, Rohit Pyati· 0 citations
This work validate and extend the eye movement based proficiency testing from single sentences to more naturalistic reading of contextualized passages in English as a second language, new proficiency measures, prediction models, and reading in an information seeking regime, and finds that the approach is effective in all these evaluations.
Shachar Frenkel, Ido Falah, Omer Shubi et al.· 0 citations
These metrics and case study offer empirical observations that could help inform data selection and script adaptation choices when working with pixel-based models in similar low-resource settings.
Ran Zhang, Miryam de Lhoneux, Wessel Poelman· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.