Skip to content

Category

natural language processing

2,926 papers

#natural language process... Preprint Aug 2026

Configurable Semantic Chunking for Biomedical Information Extraction in Retrieval-Augmented Generation

Cross-dataset analysis shows that semantic chunking improves extraction datasets with explicit relation cues, such as GM-CIHT and DDI, while fixed chunking remains competitive or stronger for dense biochemical extraction and binary classification settings such as ChemProt and ADE.

Riya Ahuja, Tim Kacprowski, Roya Shiasi Sardoabi Institute of Data Science in Biomedicine et al. · 0 citations
#natural language process... Preprint Aug 2026

DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening

DIASENTINEL demonstrates a practical framework for reliable, auditable, and privacy-preserving LLM-based clinical decision support for type 2 diabetes mellitus risk screening and guideline-grounded report generation from electronic health records (EHRs).

Yung Wei Shueh, Zhi-Jie Chen, Chiang-Hsuan Hsu et al. · 0 citations
#natural language process... Preprint Aug 2026

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

This work introduces PaperGym, a unified framework that turns each research paper into a complete training environment, and releases the pipeline, the 20,000-instance corpus PaperGym-20k, and the benchmarks PaperGym-Innov and PaperGym-Design.

Yu-Han Wang, Zhengxi Lu, Yuchen Yan et al. · 0 citations
#natural language process... Preprint Aug 2026

Aspire: Can Models Self-Evolve from Vague Goals?

This work introduces ASPIRE, a benchmark for vague-goal-driven self-evolution and shows that vague goals redirect search effort toward goal interpretation, and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals.

Yu-Hao Wu, Jingyuan Zhang, Jia-Jun Shi et al. · 0 citations
#natural language process... Preprint Aug 2026

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

These findings show that recognizing successful actions is insufficient; agents must also transform feedback into executable and transferable policies, and provide a unified framework for diagnosing this process and identifying the bottlenecks that prevent agents from translating interaction experience into reliable self-improvement.

Jia-Jun Shi, Siyang Tao, Yu-Hao Wu et al. · 0 citations
#natural language process... Preprint Aug 2026

Every Token Leaves a Ripple in the Stream of Thought: Eliciting Model-Internal Token Saliency for Chain-of-Thought Compression

Across four reasoning benchmarks and four models, \textsc{MIST} consistently outperforms baseline methods, suggesting that model-internal saliency provides an effective proxy for reasoning-token importance.

Tianyi Zhao, Yinhan He, Wendy Zheng et al. · 0 citations
#natural language process... Preprint Aug 2026

Improving Information Extraction with Learned Queries

This paper shows that another part of the pipeline matters at least as much: the queries used to elicit information extraction, and introduces List of Questions (LoQ), which generates document-specific question sets, and FeedQ, a feedback-driven optimization method that iteratively refines questions against extraction outcomes.

Omar Sharif, S. Vosoughi, Nikhil Singh · 0 citations
#natural language process... Open access Aug 2026

Type-Balanced Contextual Learning for Incremental Named Entity Recognition

This analysis shows that, in new sentences, the contextual associations of tokens representing old entity types exhibit a significantly stronger bias towards new entity types compared to their contexts in old sentences, which intensifies the degradation of old knowledge while promoting the overfitting of new knowledge.

Duzhen Zhang, Yahan Yu, Xiuyi Chen et al. · 0 citations
#natural language process... Preprint Aug 2026

Language-Statistical Analysis of Neural Audio Codec Tokens Across Architectures, Corpora, and Noise Conditions

This paper analyzes the token statistics of 13 NACs spanning multi-codebook residual vector quantization, single-codebook VQ, and non-VQ designs, evaluated on three corpora under clean, white-noise, and real-world DEMAND-noise conditions to provide architecture-conditioned conventions for applying language-statistical analysis to NAC tokens.

Joonyong Park, Shinnosuke Takamichi, David M. Chan et al. · 1 citation

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.