Skip to content

Category

natural language processing

3,231 papers

#natural language process... Preprint Aug 2026

Load-Bearing Context: The Question Damage Score for Evaluating Context Reliance in Linguistic Reasoning

Evaluating three frontier LLMs under instructions to abstain when information is insufficient, it is found they rarely abstain, often continuing to produce correct answers after load-bearing context is removed, and these findings motivate further investigation into context-based reasoning, prior knowledge, memorization, and linguistic inference.

Neh Majmudar, Elena Filatova · 0 citations
#natural language process... Preprint Aug 2026

When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages

This paper initializes byte embeddings directly from the subword representations of a frozen base model, applies a chunk alignment loss to project dynamically grouped byte chunks toward precomputed subword targets, and interleave lightweight part-of-speech supervision to guide boundary detection.

Sanjeev Kumar, Atsuki Yamaguchi, Nikolaos Aletras · 0 citations
#natural language process... Preprint Aug 2026

INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning

INSPIRE is an Internalize-Then-Improve approach combining Reference-Guided Student Internalization (RGSI), which produces high-quality preference candidates under the policy model's own distribution, with a stage-wise rubric preference training strategy that decomposes learning into method-oriented and correctness-oriented stages.

Shuai Wang, Jiayi Kuang, Yinghui Li et al. · 0 citations
#machine learning Preprint Jun 2026

Closing the Operational Gap in Semantic Caching

Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries. Standard practice evaluates these systems using PR-AUC, a metric that only measures how well scores rank and ignores whether they are usable at a fixed threshold. We show this mismatch leads to systematically poor deployment choices, as models with the highest PR-AUC are often the worst in operation. We introduce Precision--Cache Hit Ratio (P-CHR) AUC, a cache-aware metric that measures precision across cache utilization levels, and Operational Retention Rate (ORR), which captures how much offline ranking quality survives at deployment. We decompose the operational gap between offline and deployed quality into a recoverable threshold-utility component and an irreducible structural component fixed by the dataset's positive rate. Our experiments show that the threshold-utility gap is governed by the training objective rather than data scale, and yields only to re-normalizing scores over the candidate pool or changing the training objective. Ultimately, model selection for semantic caching is a threshold-utility problem, not a ranking one, and measuring it is the first step to closing the gap.

Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev et al. · 1 citation

Multilingual Lexical Feature Analysis of Spoken Language for Predicting Major Depression Symptom Severity

Depression symptom severity was associated with five lexical features, including reductions in word count measures, use of first-person plural pronouns and positive word frequency, andLexical features and vector embeddings did improve prediction accuracy beyond baseline models.

A. Tokareva, J. Dineley, Z. Firth et al. · 0 citations
#machine learning Preprint Aug 2026

Trust the Mass: Forced Weights in KV-Cache Eviction

ContourKV, a training-free allocator built from the dropped-mass statistic, wins $93$ of $160$ paired comparisons against that state of the art and loses $22$ at the byte count of the budget-enforcing baselines, and it ties the strongest of them.

J. Shi, Jerry Gu · 0 citations
#machine learning Preprint Aug 2026

Sliding-window beats linear attention

This work shows that Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models, and recommends switching to SWA instead of post-training linear models.

Alexia Jolicoeur-Martineau, R. Sukthanker, Pashmina Cameron et al. · 0 citations
#machine learning Preprint Aug 2026

How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution

Concept-Targeted Attribution (CTA) provides a framework for moving from behavioral probe accuracy to mechanistic explanations of probe performance, enabling more detailed audits of internal concept representations, including safety-critical ones.

V. Palit, Florent Draye, Terry Jingchen Zhang et al. · 0 citations
#machine learning Preprint Jul 2026

Accelerating LLM Inference via Vector Index Based Output Embeddings

This work reformulates the output projection followed by top-k token selection as a maximum inner product search over token embeddings and replaces the dense vocabulary projection with an HNSW-based vector index, suggesting approximate retrieval is a practical alternative to dense output projections in latency-sensitive small-batch decoding.

M. Loretz, Sepp Hochreiter · 0 citations
#machine learning Preprint Aug 2026

Fast Weight Attention for Continual Learning

This framework separates temporal alignment, plasticity, forgetting, and bounded rehearsal in recurrent sequence models, together with numerically stable positive-decay renormalization, to remain competitive in language modeling and improve length extrapolation on variable-digit addition.

Yi-Fan Zhang, Steve Ta, Jasper Zhang et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.