It is asked whether judges detect omissions in clinical notes, and two methods reach it independently and trade off: a per-fact pipeline, and a GEPA-evolved prompt doing the same in one call.
Sebastian Fox, L. Markham, Ryan Lail et al.· 0 citations
The Evidence Package Benchmark is introduced, integrating 1,870 packages across six heterogeneous sources with explicit modality masks and evidence permissions, and EviBound, a protocol-aware evidence control framework is proposed, a protocol-aware evidence control framework for safer clinical NLP research.
Cheng-Yuan Gao, Jiang Wu, Tao Lu et al.· 0 citations
This work evaluates Qwen2.5-7B-Instruct under INT8 and INT4 quantization on RGB and HotpotQA, measuring both accuracy and faithfulness with a hallucination detector, NLI entailment, and an LLM judge.
Atta Ul Asad, Ahsan Bilal, Muhammad Ali et al.· 0 citations
To improve self-modeling skill, a scalable synthetic-data pipeline is developed that produces self-modeling training data, and reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks.
Si-Qi Zeng, André Assis, Rowan Wang· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass, lowering the unit cost of AI-native education at scale.
Shangqing Tu, Daniel Zhang-Li, Yucheng Wang et al.· 0 citations
It is found that students'detection accuracy improves over time, driven by a shift from relying on linguistic cues to leveraging shared social and contextual signals.
Dan Schumacher, Pragathi Durga Rajarajan, Haven Kotara et al.· 0 citations
It is reported that the ranking advantage of ASCR-H does not extend to citation accuracy, where DTF matches the oracle ceiling, and three negative results on lemmatisation, pseudo-relevance feedback and query rewriting.
This work proposes TRIPPULSE1, a multi- agent framework for review-grounded travel planning, and introduces Review-Grounded Per- sona Alignment (RGPA), an LLM-as-a-Judge metric for evaluating alignment with human- centric travel experiences.
This work introduces MMDS-Bench, a diagnostic benchmark for multimodal dynamic stance classification in social media parent-reply interactions, and evaluates 12 closed-source and open-source multimodal large language models and proposed reference-grounded LLM-judge protocol for assessing reasoning quality.
This paper investigates how language models encode preference information in their intermediate representations, finding that activations from chosen and rejected responses form distinct clusters across layers, even in pretrained models.
ECGQuest provides a reproducible benchmark for contextual ECG knowledge and shows that parameter-efficient fine-tuning can make smaller language models competitive with substantially larger commercial models.
M. Hassannia, Matthew A. Reyna, R. Sameni· 0 citations
A multilingual German-English benchmark dataset that combines community-sourced stereotypes from German-speaking queer individuals with a German translation of WinoQueer is introduced, showing that language models reproduce anti-queer stereotypes, with variation across identities and models.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.