Skip to content

Category

natural language processing

3,089 papers

#natural language process... Preprint Aug 2026

Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation

Two corpus-level metrics based on self-supervised (SSL) speech representations for reference-free forced alignment evaluation are introduced: Phoneme-Cluster Mutual Information (PCMI) and Word Acoustic Consistency Score (WACS).

V.S.D.S.Mahesh Akavarapu, Michael Daniel, Gerhard Jäger · 0 citations
#natural language process... Preprint Aug 2026

Ladders in Chaos: When, How, (and Perhaps Why) Does Test-Time Scaling Improve LLM Machine Translation

This study shows that sequential sampling has a higher performance ceiling, providing a more diverse and effective pool of samples, particularly under smaller sampling budgets, and suggests an explanation of the mechanism through which sequential scaling improves machine translation.

Di Wu, S. Troshin, C. Monz et al. · 0 citations

Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge

Factual question answering (QA) typically assumes a single canonical answer, obscuring whether large language models (LLMs) retain divergent accounts of long-tail facts. To address this gap, we introduce ElephantBench, a closed-book knowledge probe comprising 1,094 questions generated through an auditable graph-based pipeline. The pipeline retrieves related documents from a low-exposure web corpus, identifies naturally occurring disagreements, and converts them into multi-account QA records. Each answer is verified against the originating documents and authoritative public web sources and is then reviewed by human annotators. Across 32 models, even the strongest model recovers both accounts on only 52.4% of questions, while on nearly all remaining questions it recalls one account but omits the other. Scaling model size and inference-time reasoning improve recall but do not eliminate this incompleteness. Corpus analysis further shows that exposure imbalance favors the dominant account, whereas greater minority-side exposure is associated with more complete recall. These findings establish ElephantBench as a reproducible knowledge probe for diagnosing epistemic myopia in parametric memory. More broadly, our graph-based benchmark construction pipeline provides an efficient and scalable way to turn long-tail corpora into source-traceable knowledge probes, supporting efforts to evaluate and advance the epistemic rigour of next-generation LLMs. Code is available at https://github.com/Tencent/ElephantBench.

Zhuo-Shi Pan, Jun-Ru Lu, Yan-Fei Qian et al. · 0 citations
#natural language process... Preprint Aug 2026

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

An RL method tailored for context management is proposed, which uses context and entropy variation to identify critical editing decisions for branch sampling and estimates action-level advantages from all branched trajectories that pass through the corresponding context editing action.

Zhuo-Shi Pan, Qizhi Pei, Jun-Ru Lu et al. · 0 citations
#natural language process... Preprint Aug 2026

A Unified Framework to Elicit Structured Feedback for Interpretable Multi-Trait Essay Scoring

HiFTS is proposed, a unified autoregressive framework that generates hierarchical CoT feedback before predicting trait-level and holistic scores and applies Group Relative Policy Optimization with a composite reward balancing score agreement, calibration, feedback quality, and structural validity.

Shi-Hang Yang, Sanwoo Lee, Ning-Ning Zhao et al. · 0 citations
#natural language process... Preprint Aug 2026

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

CultureConverse is introduced, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains and performance gains from fine-tuning on 27,860 high-quality CultureConverse-DS samples improve in-domain assistance and transfer out-of-domain to cultural MCQ and safety classification benchmarks.

Bryan Chen Zhengyu Tan, Wei-Hua Zheng, Thong T. Doan et al. · 0 citations
#natural language process... Preprint Aug 2026

PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems

This work introduces PersonaForge, a user simulation framework for synthesizing realistic multi-turn user--agent interactions that combines a four-dimensional persona space, SOUL-driven behavioral control calibrated to real-user statistics, and Reverse Deep Construction grounded in authentic seed queries.

Hanglong Lv, Dawei Zhu, Lei Li et al. · 0 citations
#natural language process... Preprint Aug 2026

FinExam-10K: When Retrieval Helps Financial Reasoning?

Professional financial examinations require models to combine domain knowledge, calculation, and judgment, yet no benchmark covers the full CFA and FRM structure under one protocol. We introduce FinExam-10K, to our knowledge the largest reported English benchmark for this setting, with 10,198 expert-reannotated questions spanning CFA Levels I-III and FRM Parts I-II. We release 5,110 questions and sequester 5,088 for a quarterly maintained leaderboard. To separate coverage from local answerability, we report a 10,198-item Full-Coverage Track and a 7,625-item Context-Complete Reasoning Track, which is the primary basis for claims about reasoning from the supplied record. Across 17 models, the best accuracy is 85.29% overall. On the frozen Hard band, the best score is 34.68% on the Full-Coverage Track and 54.57% on the 372 context-complete items. All 17 models share 47 context-complete failures. Function-RAG and FunctionGraph-RAG rescue hundreds of errors but also overturn many correct answers, producing little or negative net gain. A gate trained only on public data decides from the question and initial response when FunctionGraph-RAG should run. On the 5,088 held-out items, the gate invokes FunctionGraph-RAG for 7.9% of questions and improves accuracy from 70.83% to 71.23% (p = .0446).

Yan Lin, Jingyu Sun, Zhong-Liang Guo et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.