WebWorld is presented, the interface that lets a VLM prior interact with this browser-as-world-model autonomously and decides which interactions become supervision and reaches the level of strong frontier systems such as Kimi-K2.6 and GPT-5.4 on interactive HTML generation.
Jiajun Wu, Jian Yang, Ya-Xin Du et al.· 0 citations
UtilMem is introduced, a diagnostic benchmark comprising 1,717 instances across five domains, designed to evaluate four underexplored aspects of memory utilization: reasoning over dense histories, identifying implicitly relevant memories, synthesizing distributed evidence into summaries, analyses, or plans, and resisting interference from semantically similar distractors.
Pera describes a persistent agent organized around perception and control components that continually perceive service-relevant signals from episodic task executions, internal context, and changes in the surrounding environment, and use these signals to construct lifecycle tasks.
Shihan Dou, Haoxiang Jia, Shichun Liu et al.· 0 citations
An overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation is presented, covering spoken visual question answering and image-grounded hallucination detection in English and Modern Standard Arabic, and CRAI-Bench, evaluating the cultural accuracy of text-to-image generation.
Hi-Q is introduced, an evidence-conditioned framework for hierarchical query refinement that grows a query tree whose topology is determined by corpus support signals rather than by a fixed decomposition template or a pre-built graph.
Jueun Kim, Sungho Park, Wook-Shin Han· 0 citations
The first multilingual benchmark of paired solvable and unsolvable mathematical problems, extending ReliableMath to French and Greek, finds that Solvability Belief is encoded as a largely universal, language-agnostic feature, and that higher-resource languages, despite achieving stronger mathematical reasoning performance, exhibit lower solvability-detection faithfulness.
Maria-Eleni Zoumpoulidi, N. Xiros, Georgios Paraskevopoulos· 0 citations
A mechanistic intervention framework for identifying and transferring task-relevant sparse latent features across languages and reframes some cross-lingual reasoning gaps as failures of mechanism elicitation rather than capability absence, and offers a causally testable route to feature-mediated transfer without translation, fine-tuning, or changing the user-facing language.
Minju Song, Hyeon Hwang, Junhyun Lee et al.· 0 citations
The key observation is that although expert trajectories are scarce, high-quality final artifacts such as literature reviews, analyst reports and legal judgments, are abundant in pre-training data and can be viewed as compressed traces of the evidence-seeking processes that produced them.
Junjie Huang, Jiarui Qin, Di Yin et al.· 0 citations
This work presents S$^2$GE as an instance showing that diagnosis-driven interface design can improve native decoder usability and introduces an intervention triangle with three matched conditions: readable graph evidence, shuffled graph evidence, and no-graph input that separates evidence inclusion, structural readability, and decoder-usable topology.
Xiao-Yu Guo, Peng-Cheng Chen, Jiong Yu et al.· 0 citations
An unsupervised fine-tuning pipeline that harvests reasoning trajectories via in-context learning inference via in-context learning inference is proposed, enabling Large Language Models (LLMs) to access external knowledge and produce factual responses.
A novel SeqLab framework is proposed that enhances cross-lingual ABSA using a sequence-to-sequence model with an auxiliary sequence-labelling task performed by the encoder, enhancing aspect term recognition and sentiment predictions.
SemPOI-RL is proposed, a framework that aligns LLM semantic reasoning with structured sequence generation for interpretable OOT recommendation and consistently outperforms both traditional recommenders and direct LLM baselines, while providing interpretable style attribution across different phases of a trip.
Yunqi Liu, Yang Zhang, Ruixing Zhang et al.· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.