Skip to content

Category

natural language processing

3,089 papers

#artificial intelligence Conference Open access Jun 2026

CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories

This work borrows definitions from narratology to analyze eight intricate dimensions of character, such as stylization and wholeness, which consider more than just basic characteristics of characters within LLM and human-written stories.

A. Brei, Abhisheik Sharma, Nicholas Sanaie et al. · 0 citations

TokenPilot: Cache-Efficient Context Management for LLM Agents

TokenPilot is presented, a dual-granularity context management framework that reduces costs by 61% and 56% in isolated mode, and 61% and 87% in continuous mode, while maintaining competitive performance compared to prior systems.

Buqiang Xu, Z. Xue, Dian Chen et al. · 1 citation

DiffuSent: Towards a Unified Diffusion Framework for Aspect-Based Sentiment Analysis

DiffuSent is presented, a non-auto-regressive diffusion framework that systematically formulates all ABSA subtasks as boundary denoising diffusion processes, progressively refining boundaries over noisy states, and introduces a contrastive denoising training strategy which effectively address duplicate predictions with subtle variations introduced by diffusion process.

S. Long, Yanglei Gan, Xuchuan Zhou · 0 citations

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

LongDS is introduced, a benchmark for long-horizon, multi-turn data analysis where agents must maintain, update, restore, and compose evolving analytical states, suggesting that the key bottleneck is maintaining a correct analytical state rather than increasing interaction budget.

Kewei Xu, Xiaobe Lu, Shuofei Qiao et al. · 1 citation

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

This work presents SkillSafetyBench, a runnable benchmark for evaluating skill-facing safety failures, and suggests that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.

Chang Jin, Anr'an W'ang, Zeming Wei et al. · 11 citations · ⚡1

G-Loss: Graph-Guided Fine-Tuning of Language Models

G-Loss is presented, a graph-guided loss function that incorporates semi-supervised label propagation to use structural relationships within the embedding manifold to build a document-similarity graph that captures global semantic relationships.

Aditya Sharma, Vinti Agarwal, Rajesh Kumar · 0 citations

Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs

LLMs have limitations when it comes to cultural coverage and competence, and in some cases, show specific cultural biases. Although prior studies have examined the cultural capabilities of LLMs, none have specifically investigated their regional preferences in generic culture-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ), with questions available in 24 languages. We evaluate LLMs by prompting them to answer questions from CROQ and provide a sample location. The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan in their answers. Moreover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs. Low-resource languages, on the other hand, show more inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training. Dataset available at https://huggingface.co/datasets/HiTZ/CROQ

Joseba Fernandez de Landa, Carla Pérez-Almendros, J. Camacho-Collados · 1 citation

Select, Label, Evaluate: Active Testing in NLP

This work formalizes Active Testing in NLP and conducts an extensive benchmarking of existing approaches across 18 datasets and 4 embedding strategies spanning 4 different NLP tasks, revealing variations in method effectiveness across different data characteristics and task types.

Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu et al. · 1 citation

Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts

There is potential to improve cross-lingual parametric knowledge transfer during post-training by providing the LLMs with the key entities of the questions in their source language and finding that this disproportionately improves cross-script questions.

Lucas Bandarkar, Alan Ansell, Trevor Cohn · 3 citations
#artificial intelligence Preprint Feb 2026

From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves

The results show that improving IF in LRMs can significantly enhance privacy, suggesting a promising direction for future privacy-aware LRMs, and introduces an SFT dataset that teaches models to follow general instructions throughout their reasoning process.

Haritz Puerto, Haonan Li, Xudong Han et al. · 0 citations

FENCE: A Financial and Multimodal Jailbreak Detection Dataset

FENCE, a bilingual (Korean-English) multimodal dataset for training and evaluating jailbreak detectors in financial applications, provides a focused resource for advancing multimodal jailbreak detection in finance and for supporting safer, more reliable AI systems in sensitive domains.

Mirae Kim, Seonghun Jeong, Youngjun Kwak · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.