Skip to content

Category

natural language processing

2,394 papers

LLP: LLM-based Product Pricing in E-commerce

Inspired by recent breakthroughs in Large Language Models (LLMs), this work introduces LLP, the first LLM-based generative framework for second-hand product pricing that substantially surpasses existing methods while generalizing well to unseen categories.

Hairu Wang, Sheng You, Qi-Heng Zhang et al. · 3 citations · ⚡1

FocusAgent: Simple Yet Effective Ways of Trimming the Large Context of Web Agents

A variant of FocusAgent significantly reduces the success rate of prompt-injection attacks, including banner and pop-up attacks, while maintaining task success performance in attack-free settings, highlighting that targeted LLM-based retrieval is a practical and robust strategy for building web agents that are efficient, effective, and secure.

Imene Kerboua, S. Shayegan, Megh Thakkar et al. · 15 citations

Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models

This work conducts the first large-scale, systematic studies of multilingual calibration across six model families and over 100 languages, revealing that non-English languages suffer from systematically worse calibration.

Ej Zhou, Caiqi Zhang, Tiancheng Hu et al. · 10 citations · ⚡1

Transformer-Encoder Trees for Efficient Multilingual Machine Translation and Speech Translation

Transformer Encoder Tree is introduced, a hierarchical, non-autoregressive encoder-only architecture trained with Connectionist Temporal Classification for multilingual translation that eliminates the sequential bottleneck of autoregressive models and supports fully parallel decoding of all tokens across all target languages.

Yiwen Guan, Jacob Whitehill · 0 citations
#natural language process... Preprint Sep 2025

Privacy-Preserving Generation of Clinical Narratives from Medical Terminologies

Compared to existing DP text generation baselines, Term2Note substantially improves both fidelity and utility, without relying on label distribution assumptions, highlighting its effectiveness as a practical privacy-preserving alternative to real clinical notes.

Yuping Wu, Viktor Schlegel, Warren Del-Pinto et al. · 2 citations

MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions

The growing demand for scalable psychological counseling highlights the need for high-quality, privacy-compliant data, yet such data remains scarce. Here we introduce MAGneT, a novel multi-agent framework for synthetic psychological counseling session generation that decomposes counselor response generation into coordinated sub-tasks handled by specialized LLM agents, each modeling a key psychological technique. Unlike prior single-agent approaches, MAGneT better captures the structure and nuance of real counseling. We further propose a unified evaluation framework that consolidates diverse automatic metrics and expands expert assessment from four to nine counseling aspects, thus addressing inconsistencies in prior evaluation protocols. Empirically, MAGneT substantially outperforms existing methods: experts prefer MAGneT-generated sessions in 77.2% of cases on average across the nine aspects over the strongest baseline, and sessions generated by MAGneT using Llama3-8B-Instruct backbone yield 3.2% higher general counseling skills and 4.3% higher CBT-specific skills on cognitive therapy rating scale (CTRS). An open source Llama3-8B-Instruct model fine-tuned on MAGneT-generated data also outperforms models fine-tuned using baseline synthetic datasets by 6.9% on average on CTRS. We make our code, data and fine-tuned model public.

Aishik Mandal, Tanmoy Chakraborty, Iryna Gurevych · 7 citations
#natural language process... Open access Aug 2025

Political Ideology Shifts in Large Language Models

Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideological biases. In this work, we examine how synthetic persona conditioning shapes ideological expression across seven open-weight instruction-tuned models (7B-72B parameters) using the Political Compass Test (62 statements) as a standardized behavioral probe. Across three studies involving 200,000 synthetic personas and more than 260 million model responses, we analyze implicit and explicit malleability, as well as theme-associated variations. We find that: (i) larger models exhibit broader implicit ideological coverage, increasing from 14-35% for 7-8B models to up to 49% for 70B+ models; (ii) explicit ideological priming induces large and statistically significant shifts, with right-authoritarian cues moving all models in the intended direction and producing larger effects in most model-axis comparisons; (iii) left-libertarian priming produces more heterogeneous responses, including counter-directional economic shifts in three of four 7-8B models, while all 70B+ models move in the intended direction; and (iv) theme-associated semantic content in persona descriptions is linked to systematic and interpretable directional shifts in ideological space. While our results identify an upstream mechanism through which persona conditioning can alter model responses under a standardized ideological probe, we do not test whether such shifts affect users beliefs, decisions, or political behavior. Our findings are best understood as evidence of ideological malleability at the generation layer, highlighting the need to account for interactional factors when evaluating political neutrality, fairness, and safety in English-prompted, persona-conditioned language models.

Pietro Bernardelle, Stefano Civelli, Leon Fröhling et al. · 5 citations

Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever

Behavior Aligned Retrieval (BAR) is proposed, a backbone-agnostic training recipe that teaches a dense retriever a behavior-aware similarity, keeping semantically related candidates close only when their tool-use behavior is compatible.

Yixin Chen, Ying Xiong, Shangyu Wu et al. · 0 citations
#natural language process... Preprint Aug 2025

Towards AI-Assisted Research Writing: Benchmarking LLMs for AI/ML Introduction Generation

This work introduces Scientific Introduction Generation (SciIG), a task that evaluates LLMs'ability to produce coherent introductions from titles, abstracts, and related works, and combines automated metrics with LLM-as-a-judge evaluations.

Krishna Garg, Firoz Shaik, Sambaran Bandyopadhyay et al. · 5 citations

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

The introduction of MM-BrowseComp, a novel benchmark comprising 400 challenging, hand-crafted questions designed to evaluate multimodal retrieval and reasoning capabilities, is introduced, establishing MM-BrowseComp as a rigorous new standard for the field.

Shilong Li, Xingyuan Bu, Wenjie Wang et al. · 37 citations · ⚡7
#natural language process... Conference Open access Aug 2025

SinLlama - A Large Language Model for Sinhala

This research extends an existing multilingual LLM (Llama-3-8B) to get a better coverage for Sinhala and enhances the LLM tokenizer with Sinhala specific vocabulary and performs continual pre-training on a 10 million sentence Sinhala corpus, resulting in the SinLlama model.

H.W.K. Aravinda, Rashad Sirajudeen, Samith Karunathilake et al. · 10 citations · ⚡1

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.