Skip to content

Category

natural language processing

2,491 papers

#natural language process... Preprint Aug 2026

Every Token Leaves a Ripple in the Stream of Thought: Eliciting Model-Internal Token Saliency for Chain-of-Thought Compression

Across four reasoning benchmarks and four models, \textsc{MIST} consistently outperforms baseline methods, suggesting that model-internal saliency provides an effective proxy for reasoning-token importance.

Tianyi Zhao, Yinhan He, Wendy Zheng et al. · 0 citations
#natural language process... Preprint Aug 2026

Improving Information Extraction with Learned Queries

This paper shows that another part of the pipeline matters at least as much: the queries used to elicit information extraction, and introduces List of Questions (LoQ), which generates document-specific question sets, and FeedQ, a feedback-driven optimization method that iteratively refines questions against extraction outcomes.

Omar Sharif, S. Vosoughi, Nikhil Singh · 0 citations
#natural language process... Open access Aug 2026

Type-Balanced Contextual Learning for Incremental Named Entity Recognition

This analysis shows that, in new sentences, the contextual associations of tokens representing old entity types exhibit a significantly stronger bias towards new entity types compared to their contexts in old sentences, which intensifies the degradation of old knowledge while promoting the overfitting of new knowledge.

Duzhen Zhang, Yahan Yu, Xiuyi Chen et al. · 0 citations
#natural language process... Preprint Aug 2026

Language-Statistical Analysis of Neural Audio Codec Tokens Across Architectures, Corpora, and Noise Conditions

This paper analyzes the token statistics of 13 NACs spanning multi-codebook residual vector quantization, single-codebook VQ, and non-VQ designs, evaluated on three corpora under clean, white-noise, and real-world DEMAND-noise conditions to provide architecture-conditioned conventions for applying language-statistical analysis to NAC tokens.

Joonyong Park, Shinnosuke Takamichi, David M. Chan et al. · 1 citation
#natural language process... Preprint Aug 2026

Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols

Computational mental health screening using multimodal speech and text has shown great promise. However, existing models often assume all clinical speech protocols carry equivalent evidentiary validity. In reality, heterogeneous protocols, from free interviews to fixed reading tasks, support fundamentally different evidence. Forcing uniform reasoning flattens these boundaries, causing models to hallucinate symptoms from irrelevant text or overclaim support. Even advanced long chain-of-thought LLMs fail to resolve this issue, as free-form reasoning can exacerbate boundary violations. To address this, we reformulate multimodal screening as an evidence-bounded reasoning problem. We introduce the Evidence Package Benchmark, integrating 1,870 packages across six heterogeneous sources with explicit modality masks and evidence permissions. We further propose EviBound, a protocol-aware evidence control framework. Unlike direct LLM prompting, EviBound uses a profile-aware planner to restrict reasoning scope, orchestrates evidence tools via five-way acoustic consensus, and enforces a boundary critic to suppress unsupported claims. Empirical results show EviBound achieves a held-out test Depression AUROC of 0.8658, exceeding the strongest direct omni-modal baseline by +0.0811 AUROC while maintaining zero claim violations. Our work moves beyond unconstrained accuracy toward evidence-consistent, protocol-aware systems for safer clinical NLP research.

Cheng-Yuan Gao, Jiang Wu, Tao Lu et al. · 0 citations
#natural language process... Preprint Open access Sep 2026

Faithfulness Is Not Free: Auditing Offline KV-Cache Quantization in Retrieval-Augmented Generation

Retrieval-augmented generation systems can precompute and store key-value caches of retrieved documents to avoid re-encoding context at every query. Quantizing these caches further reduces storage, but no prior work asks whether compression damages faithfulness, whether responses remain grounded in the retrieved evidence. Faithfulness and accuracy are not equivalent: a model can produce a correct answer that is no longer supported by the context it was given. We evaluate Qwen2.5-7B-Instruct under INT8 and INT4 quantization on RGB and HotpotQA, measuring both accuracy and faithfulness with a hallucination detector, NLI entailment, and an LLM judge. INT8 is near-lossless across both metrics. INT4 reduces accuracy and, more critically, even among answers that remain factually correct, over 90% of faithfulness changes are negative, i.e., accuracy metrics are blind to this regression. The harm grows under noisy retrieval and with more retrieved chunks. Faithfulness must be audited before compressed caches are deployed.

Atta Ul Asad, Ahsan Bilal, Muhammad Ali et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Evaluating and Improving LLM Self-Modeling

To improve self-modeling skill, a scalable synthetic-data pipeline is developed that produces self-modeling training data, and reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks.

Si-Qi Zeng, André Assis, Rowan Wang · 0 citations
#artificial intelligence Preprint Aug 2026

CogEvol: Towards Efficient and Reliable Learning Environment Generation

CogEvol, a family of models trained specifically for Learning Environment Generation: turning a course brief into a finished learning artifact (structured-JSON slides or self-contained interactive HTML pages) in a single pass, lowering the unit cost of AI-native education at scale.

Shangqing Tu, Daniel Zhang-Li, Yucheng Wang et al. · 0 citations
#natural language process... Preprint Aug 2026

Detecting AI Impostors: How Do Middle Schoolers Identify LLM Agents in a Live Collaborative Setting?

LLMs can imitate how people write, which raises concerns about impersonation, trust, and detection in social settings. These concerns are especially important for adolescents, who use generative AI frequently but may struggle to recognize it. We introduce \textit{DoppelBot}, a cooperative social deduction game designed to study how young people detect and respond to AI impersonation. Through studies with middle schoolers, we investigate whether a DoppelBot prompts reflection on privacy and impersonation, how repeated exposure affects AI-detection accuracy as agents become more personalized, and which strategies students use to identify AI doppelg\"angers. We find that students'detection accuracy improves over time, driven by a shift from relying on linguistic cues to leveraging shared social and contextual signals. Students also demonstrated an understanding of AI limitations such as embodiment and reflected on broader issues such as data privacy. To support future research, we release an anonymized dataset of game transcripts and voting behavior.

Dan Schumacher, Pragathi Durga Rajarajan, Haven Kotara et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.