ADAS is proposed, a training-free reranking rule that leaves the base sampler's stopping rule unchanged and greedily discounts each token-wise confidence score according to its attention to already selected positions, weighted by their prediction uncertainty.
Y. Şahin, Ahmed R. Saikia, V. Cevher et al.· arXiv.org· 1 citation
Prefilling-dLLM is proposed, a training-free prefill-decode disaggregation framework for dLLMs that partitions the prefix into N chunks, caches their KV representations once, and selects the top-K most relevant chunks with intra-chunk token sparsity for decoding, showing that sparse prefilling can outperform dense attention while reducing per-step complexity.
Jingfei Xiong, Qilong Han, Shansan Gong et al.· arXiv.org· 0 citations
This work introduces a clinically grounded framework that evaluates leakage along a graded axis of adversarial access, ranging from publicly inferable demographics to leaked note fragments, and provides a practical, reusable framework for contextual privacy evaluation of medical LMs.
Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle et al.· arXiv.org· 0 citations
This work introduces LexRubric, a rubric-based benchmark for evaluating open-ended Chinese legal tasks and evaluates 18 recent general and legal-domain LLMs on LexRubric, showing that different models exhibit distinct capability profiles, and that open-ended legal tasks remain challenging for current LLMs.
Yifan Chen, Haitao Li, Yiran Hu et al.· arXiv.org· 1 citation
Reach audiences
Advertise in front of researchers, engineers, and readers.
A genre-diverse sentence segmentation corpus spanning eight genres and a wide range of punctuation and document structure conditions is introduced, using AraSEG to evaluate LLMs, lightweight encoder models, and dependency parser-based models under increasingly challenging segmentation settings.
M. Elkholy, Khalid N. Elmadani, Nizar Habash et al.· arXiv.org· 1 citation
A multi-track evaluation covering diverse datasets and state-of-the-art LLMs reveals a more nuanced landscape in which human references continue to demonstrate advantages in informativeness and faithfulness, whereas LLM outputs are preferred mainly for surface-level coherence and fluency.
ABLE (Attribution-Based Large-model Embedding), a framework that leverages the interpretability space to construct model representations by aggregating gradient-based feature attributions via a tokenizer-agnostic word-level alignment, captures model-specific input-sensitivity patterns rather than only surface-level outputs.
This work studies counterfactual context revision as a framework for auditing LLM-based stance simulation and reveals effective and robust stance transitions in both text-only and multimodal strategies across different polarization-preference mechanisms.
Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are grounded in different cultural backgrounds. Existing cultural evaluation focuses on value alignment: how closely a single agent matches a target culture. Yet alignment is a per-agent property and cannot reveal whether a system, taken as a whole, preserves the cultural plurality it is meant to represent. We propose value diversity as a system-level evaluation axis for multicultural agent systems, defined through the dissimilarity between culturally conditioned agents'responses on a shared value survey. Using the World Values Survey, we evaluate 19 cultures and 18 backbone models across a wide range of system configurations. We find that diversity is largely uncorrelated with alignment, indicating that the two capture complementary system properties, and that current multicultural agent systems fall substantially below human societies in value diversity. Mixed-backbone systems narrow this gap but do not close it, and the gap persists across culture compositions and agent scales. Social interaction further erodes diversity by driving agents toward consensus, and a participatory budgeting case study shows that this homogenization narrows the breadth of collective decision-making. Together, our results establish value diversity as a distinct evaluation axis for multicultural multi-agent systems and reveal a persistent homogenization tendency in current LLM-based societies. Our code and data are publicly available at https://github.com/iNLP-Lab/MultiAgent-Diversity.
Shaoyang Xu, Jingsheng Zhang, Ho-Ang P. Long et al.· arXiv.org· 0 citations
Existing jailbreaks against aligned LLMs are discrete artifacts whose surface forms are easy to fingerprint and patch. We argue that the broader failure mode may lie not in any specific prompt, but in natural writing registers that safety tuning under-covers. Building on this insight, we introduce VAR, the first jailbreak family that uses real fanfiction subgenres as universal attack carriers: a creative-writing meta prompt is conditioned on passages from one of twelve Archive of Our Own (AO3) subgenres, and the harmful behavior is embedded as the climax of the resulting scene. The construction requires neither an adversarial attacker LLM nor optimization. On eight aligned LLMs over the union of HarmBench and JailbreakBench, this attack lifts mean ASR from 0.278 to 0.731 under a four-judge ensemble; a factorial decomposition shows the gain is carried by register rather than length or structure. Two active defenses widen rather than narrow the vernacular-to-baseline ratio, indicating that template-targeting defenses merely steer attackers toward register-based attacks like ours. We also propose VAR-A4, a static four-turn extension that attains a mean ASR of 0.924, substantially exceeding three existing multi-turn methods. Our code and data are safely open-sourced at https://github.com/T-Lab-CUHKSZ/VAR.
Zhongze Luo, Rui-He Shi, Zhenshuai Yin et al.· 0 citations
The results suggest that length generalization is a meaningful stress test for creative-writing models and a useful lens for distinguishing otherwise close models.
QUBRIC, a framework that co-designs queries and rubrics can make rubric-based RL a practical complement to RLVR beyond strictly verifiable tasks, provides evidence that co-designing queries and rubrics can make rubrics a practical complement to RLVR beyond strictly verifiable tasks.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.