Two factors are identified that predict most of UASL's variability across relations: the mean and dispersion of the linear distance between the related words, and the diversity of the diversity of the syntactic relation's head.
Juan Pablo Vigneaux, Mary Kennedy, Khalil Iskarous et al.· 0 citations
The ability of transformer-based language models to learn k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols, is considered, finding that LLMs trained on them achieve comparable cross-entropy loss regardless of antilocality, but converge more slowly on more antilocal languages.
Andrew McInnerney, Shane Storks, Steven P. Abney et al.· 0 citations
Evaluating three frontier LLMs under instructions to abstain when information is insufficient, it is found they rarely abstain, often continuing to produce correct answers after load-bearing context is removed, and these findings motivate further investigation into context-based reasoning, prior knowledge, memorization, and linguistic inference.
This paper initializes byte embeddings directly from the subword representations of a frozen base model, applies a chunk alignment loss to project dynamically grouped byte chunks toward precomputed subword targets, and interleave lightweight part-of-speech supervision to guide boundary detection.
INSPIRE is an Internalize-Then-Improve approach combining Reference-Guided Student Internalization (RGSI), which produces high-quality preference candidates under the policy model's own distribution, with a stage-wise rubric preference training strategy that decomposes learning into method-oriented and correctness-oriented stages.
Shuai Wang, Jiayi Kuang, Yinghui Li et al.· 0 citations
Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries. Standard practice evaluates these systems using PR-AUC, a metric that only measures how well scores rank and ignores whether they are usable at a fixed threshold. We show this mismatch leads to systematically poor deployment choices, as models with the highest PR-AUC are often the worst in operation. We introduce Precision--Cache Hit Ratio (P-CHR) AUC, a cache-aware metric that measures precision across cache utilization levels, and Operational Retention Rate (ORR), which captures how much offline ranking quality survives at deployment. We decompose the operational gap between offline and deployed quality into a recoverable threshold-utility component and an irreducible structural component fixed by the dataset's positive rate. Our experiments show that the threshold-utility gap is governed by the training objective rather than data scale, and yields only to re-normalizing scores over the candidate pool or changing the training objective. Ultimately, model selection for semantic caching is a threshold-utility problem, not a ranking one, and measuring it is the first step to closing the gap.
Depression symptom severity was associated with five lexical features, including reductions in word count measures, use of first-person plural pronouns and positive word frequency, andLexical features and vector embeddings did improve prediction accuracy beyond baseline models.
A. Tokareva, J. Dineley, Z. Firth et al.· arXiv.org· 0 citations
ContourKV, a training-free allocator built from the dropped-mass statistic, wins $93$ of $160$ paired comparisons against that state of the art and loses $22$ at the byte count of the budget-enforcing baselines, and it ties the strongest of them.
The results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.
This work shows that Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models, and recommends switching to SWA instead of post-training linear models.
Alexia Jolicoeur-Martineau, R. Sukthanker, Pashmina Cameron et al.· 0 citations
The exact DP constant is pin down for the two that carry the practical weight, counterfactual memorization and adaptive extraction, and it is shown that they do not control each other.
Concept-Targeted Attribution (CTA) provides a framework for moving from behavioral probe accuracy to mechanistic explanations of probe performance, enabling more detailed audits of internal concept representations, including safety-critical ones.
V. Palit, Florent Draye, Terry Jingchen Zhang et al.· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.