This work borrows definitions from narratology to analyze eight intricate dimensions of character, such as stylization and wholeness, which consider more than just basic characteristics of characters within LLM and human-written stories.
A. Brei, Abhisheik Sharma, Nicholas Sanaie et al.· Annual Meeting of the Associ...· 0 citations
TokenPilot is presented, a dual-granularity context management framework that reduces costs by 61% and 56% in isolated mode, and 61% and 87% in continuous mode, while maintaining competitive performance compared to prior systems.
Buqiang Xu, Z. Xue, Dian Chen et al.· arXiv.org· 1 citation
Sycophancy co-occurs with degraded judged truthfulness (rho=0.40), a coupling that strengthens across generations, and a single direct instruction outperforms an elaborate reasoning protocol in seven of eight variants.
DiffuSent is presented, a non-auto-regressive diffusion framework that systematically formulates all ABSA subtasks as boundary denoising diffusion processes, progressively refining boundaries over noisy states, and introduces a contrastive denoising training strategy which effectively address duplicate predictions with subtle variations introduced by diffusion process.
S. Long, Yanglei Gan, Xuchuan Zhou· arXiv.org· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
LongDS is introduced, a benchmark for long-horizon, multi-turn data analysis where agents must maintain, update, restore, and compose evolving analytical states, suggesting that the key bottleneck is maintaining a correct analytical state rather than increasing interaction budget.
This work presents SkillSafetyBench, a runnable benchmark for evaluating skill-facing safety failures, and suggests that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.
G-Loss is presented, a graph-guided loss function that incorporates semi-supervised label propagation to use structural relationships within the embedding manifold to build a document-similarity graph that captures global semantic relationships.
LLMs have limitations when it comes to cultural coverage and competence, and in some cases, show specific cultural biases. Although prior studies have examined the cultural capabilities of LLMs, none have specifically investigated their regional preferences in generic culture-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ), with questions available in 24 languages. We evaluate LLMs by prompting them to answer questions from CROQ and provide a sample location. The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan in their answers. Moreover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs. Low-resource languages, on the other hand, show more inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training. Dataset available at https://huggingface.co/datasets/HiTZ/CROQ
Joseba Fernandez de Landa, Carla Pérez-Almendros, J. Camacho-Collados· arXiv.org· 1 citation
This work formalizes Active Testing in NLP and conducts an extensive benchmarking of existing approaches across 18 datasets and 4 embedding strategies spanning 4 different NLP tasks, revealing variations in method effectiveness across different data characteristics and task types.
Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu et al.· arXiv.org· 1 citation
There is potential to improve cross-lingual parametric knowledge transfer during post-training by providing the LLMs with the key entities of the questions in their source language and finding that this disproportionately improves cross-script questions.
Lucas Bandarkar, Alan Ansell, Trevor Cohn· arXiv.org· 3 citations
The results show that improving IF in LRMs can significantly enhance privacy, suggesting a promising direction for future privacy-aware LRMs, and introduces an SFT dataset that teaches models to follow general instructions throughout their reasoning process.
Haritz Puerto, Haonan Li, Xudong Han et al.· 0 citations
FENCE, a bilingual (Korean-English) multimodal dataset for training and evaluating jailbreak detectors in financial applications, provides a focused resource for advancing multimodal jailbreak detection in finance and for supporting safer, more reliable AI systems in sensitive domains.
Mirae Kim, Seonghun Jeong, Youngjun Kwak· arXiv.org· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.