Skip to content

Category

natural language processing

2,926 papers

#natural language process... Preprint Aug 2026

Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols

The Evidence Package Benchmark is introduced, integrating 1,870 packages across six heterogeneous sources with explicit modality masks and evidence permissions, and EviBound, a protocol-aware evidence control framework is proposed, a protocol-aware evidence control framework for safer clinical NLP research.

Cheng-Yuan Gao, Jiang Wu, Tao Lu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Evaluating and Improving LLM Self-Modeling

To improve self-modeling skill, a scalable synthetic-data pipeline is developed that produces self-modeling training data, and reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks.

Si-Qi Zeng, André Assis, Rowan Wang · 0 citations
#artificial intelligence Preprint Aug 2026

CogEvol: Towards Efficient and Reliable Learning Environment Generation

CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass, lowering the unit cost of AI-native education at scale.

Shangqing Tu, Daniel Zhang-Li, Yucheng Wang et al. · 0 citations
#natural language process... Preprint Aug 2026

Annotated Surrogate Retrieval for Polish Statutory Law

It is reported that the ranking advantage of ASCR-H does not extend to citation accuracy, where DTF matches the oracle ceiling, and three negative results on lemmatisation, pseudo-relevance feedback and query rewriting.

Orkun Yiğit Cengiz · 0 citations

TRIPPULSE: Multi-Agent Travel Planning with Review-Grounded Reasoning

This work proposes TRIPPULSE1, a multi- agent framework for review-grounded travel planning, and introduces Review-Grounded Per- sona Alignment (RGPA), an LLM-as-a-Judge metric for evaluating alignment with human- centric travel experiences.

Priyanshu Karmakar, Borru Vijay Sai, Shubhojit Mallick et al. · 0 citations
#natural language process... Preprint Aug 2026

MMDS-Bench: Benchmarking Multimodal Large Language Models on Dynamic Stance in Social Media Interactions

This work introduces MMDS-Bench, a diagnostic benchmark for multimodal dynamic stance classification in social media parent-reply interactions, and evaluates 12 closed-source and open-source multimodal large language models and proposed reference-grounded LLM-judge protocol for assessing reasoning quality.

Yuzhe Ding, Kang He, Li Zheng et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models

A multilingual German-English benchmark dataset that combines community-sourced stereotypes from German-speaking queer individuals with a German translation of WinoQueer is introduced, showing that language models reproduce anti-queer stereotypes, with variation across identities and models.

M. Morch, Daniel Braun · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.