Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually translate an entire example into one language. This hides failures at language boundaries inside the reasoning chain. We introduce XHotpotQA, a controlled benchmark for cross-lingual knowledge composition over mixed-language evidence. Each instance is modeled as an evidence-dependency graph whose question, bridge evidence, answer-bearing evidence, and distractors have explicit language assignments. The audited resource contains 15,661 training and 7,405 validation instances, with sentence-level support supervision and supplied distractors. In validation, 99.81% of items cross the question-to-gold-evidence language interface and 95.60% use gold paragraphs in different languages. Across three reader artifacts, full question-evidence mismatch is associated with 10.25 to 15.79 lower Unicode-aware answer F1 than partial alignment, and different-script evidence with deficits of 11.98 to 23.70 points; the corresponding adapted-selector contrasts are 1.71 and 1.78 points. Under this supplied-candidate design, the evaluated readers therefore show substantially larger condition-associated deficits than the selector. XHotpotQA provides role-aware diagnostics, modular evaluation, and an audited test bed for knowledge-based systems that must integrate evidence across languages.
Iman Barati, A. Ghafouri, B. Minaei-Bidgoli· 0 citations
A systematic comparison of retrieval strategies for candidate generation under a shared LLM-based selection stage, combining sparse retrieval (BM25), Web KB search, and a state-of-the-art trained dense retriever with several open- and closed-source LLMs is presented.
Fina Polat, Daniel Daza, Pengyu Zhang et al.· 0 citations
The UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records, is described and an answer-first pipeline in which the model generates candidate answers citing specific note sentences is proposed, exploiting the asymmetry between judging relevance in the abstract versus relative to a generated answer.
Mohammad Arvan, Hossein Haeri, Natalie Parde et al.· Proceedings of the Language...· 0 citations
Experiments show that PACE outperforms scalable non-manual baselines while approaching the quality of manually engineered publisher-specific parsers, demonstrating that agentic configuration learning can automate publisher-specific extraction for LLM-ready page representations beyond article text.
Zhanli Liu, Munirathnam Srikanth· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Through a controlled design that separates emotion from conversational context, it is shown that emotional context increases LLM sycophancy even in top-tier flagship models.
Fine-grained Adaptive Implicit Hate speech Detection (FAID) is proposed, a novel framework that first performs fine-grained classification and then adapts to specific categories, and significantly outperforms SOTA baselines.
Han Wang, Yu-Hu Cheng, Xue-Song Wang et al.· 0 citations
The proposed DMRA, a deficit-based diagnostic framework that quantifies the contribution of these components to identify the primary cause of unsuccessful cases, reveals that relational reasoning is the primary source of error across all models, followed by memory limitations.
N. Yilmaz, Naga Sai Abhiram Kusumba, Stella Wenxing Liu et al.· 0 citations
It is shown that ASR errors lead to significant safety risks for embodied AI, and automatic correction of ASR errors can reduce the risk, but this is not always effective.
This work proposes a lookahead-guided decoding framework for context-free grammars based on pushdown automata based on bounded pushdown summaries with reachability labels and upper-bound distances to acceptance.
Vincenzo Collura, Karim Tit, Eleonora Giunchiglia et al.· 0 citations
This perspective clarifies where current alignment methods genuinely benefit from game-theoretic analysis, where the framework is looser, and what challenges remain in building robust, adaptive, and verifiable AI systems.
Yaxin Cai, Zhongyue Zhao, Zhigang Lu et al.· 0 citations
CEDAR is presented, a counterexample-guided framework that grounds instructions as regular languages over environment event traces and represents both skills and specifications as deterministic finite automata, suggesting that regular languages offer a practical verification layer between natural-language instructions and embodied-agent policies.
Le Chen, Alvaro Velasquez, Ashutosh Trivedi· 0 citations
This work separates this failure into two precisely defined quantities: occurrence, how often the model makes an unsupported claim on its own, measured from the visible evidence and final claim without using the hidden correct answer; and conditional repair, how often those same naturally occurring unsupported claims are repaired when the missing evidence is supplied.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.