Skip to content

Category

natural language processing

2,491 papers

#natural language process... Preprint Aug 2026

RENSA: Rich Environment Metadata to Navigate Shared and Distributed Endpoints for Automated Federated SPARQL Query Generation

This work proposes RENSA, a federated SPARQL query generation framework that leverages an extension of SPARQL Builder Metadata (SBM), and demonstrates that RENSA infers class and authority constraints for query variables, enabling the identification of data sources even across heterogeneous endpoints.

Victor Eiti Yamamoto, Hideaki Takeda, Yasunori Yamamoto · 0 citations
#artificial intelligence Preprint Aug 2026

Automated Researchers Can Mitigate Well-characterized Alignment Failures

Using human ideas as the AARs' initial research direction does not improve performance, suggesting current AARs may not need guidance from experienced researchers, and suggests that automating alignment research on well-characterized failures may be practical in the near term.

Yueh-Han Chen, Jia-Xin Wen, J. Kirchner · 0 citations

Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey

Personality-narrative conditioning beat demographic-only conditioning, but neither surpassed simple real-data baselines; Synthetic panels are thus not survey substitutes; their value is diagnostic, with operational use confined to settings lacking real data.

Howard Kim, Keun Tae Cho · 0 citations
#artificial intelligence Review Jul 2026

RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciences

Cross-jurisdiction regulatory divergence detection is introduced: given an FDA requirement and an EMA requirement on the same topic, classify their relationship as AGREE, DIVERGE, or SILENT and three directional observations emerge at pilot scale.

Chu-Chu Wu, Zhi-Ying Zhou, Jing-Zhu Hu et al. · 2 citations
#natural language process... Preprint Aug 2026

Context-Aware Interleaved Batching for WhisperX

By using VAD-derived segment boundaries, the algorithm stabilizes Whisper's text conditioning, allowing us to safely maintain continuous historical context across batched audio segments, all while maintaining high-throughput inference speeds.

Carlos Bain, Max Bain · 0 citations
#natural language process... Preprint Aug 2026

Configurable Semantic Chunking for Biomedical Information Extraction in Retrieval-Augmented Generation

Cross-dataset analysis shows that semantic chunking improves extraction datasets with explicit relation cues, such as GM-CIHT and DDI, while fixed chunking remains competitive or stronger for dense biochemical extraction and binary classification settings such as ChemProt and ADE.

Riya Ahuja, Tim Kacprowski, Roya Shiasi Sardoabi Institute of Data Science in Biomedicine et al. · 0 citations
#natural language process... Preprint Aug 2026

DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening

DIASENTINEL demonstrates a practical framework for reliable, auditable, and privacy-preserving LLM-based clinical decision support for type 2 diabetes mellitus risk screening and guideline-grounded report generation from electronic health records (EHRs).

Yung Wei Shueh, Zhi-Jie Chen, Chiang-Hsuan Hsu et al. · 0 citations
#natural language process... Preprint Aug 2026

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

This work introduces PaperGym, a unified framework that turns each research paper into a complete training environment, and releases the pipeline, the 20,000-instance corpus PaperGym-20k, and the benchmarks PaperGym-Innov and PaperGym-Design.

Yu-Han Wang, Zhengxi Lu, Yuchen Yan et al. · 0 citations
#natural language process... Preprint Aug 2026

Aspire: Can Models Self-Evolve from Vague Goals?

This work introduces ASPIRE, a benchmark for vague-goal-driven self-evolution and shows that vague goals redirect search effort toward goal interpretation, and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals.

Yu-Hao Wu, Jingyuan Zhang, Jia-Jun Shi et al. · 0 citations
#natural language process... Preprint Aug 2026

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

These findings show that recognizing successful actions is insufficient; agents must also transform feedback into executable and transferable policies, and provide a unified framework for diagnosing this process and identifying the bottlenecks that prevent agents from translating interaction experience into reliable self-improvement.

Jia-Jun Shi, Siyang Tao, Yu-Hao Wu et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.