Skip to content

Category

natural language processing

2,926 papers

#machine learning Preprint Aug 2026

LoGo: Token-Level Dynamic Local-Global Attention

LoGo, a token-level dynamic local-global attention mechanism that uses attention span as a direct proxy for attention budget allocation, is proposed and results suggest that learned token-level span allocation is an effective and scalable way to improve the long-context performance-compute trade-off.

Yuqi Pan, Zheng Li, Bohao Tang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

A benchmark score is a joint property of the model, the evaluation harness, the elicitation budget, the sampled population, and contamination status. Leaderboards publish the model and the score, so capability and leakage stay observationally equivalent. Existing taxonomies classify contamination for automated detection, not the question a reporter faces at publication: given the mitigations already applied, which validity threats remain open? We introduce a taxonomy organized by the mitigation each type defeats -- direct, derivative, temporal, distributional, and acquired -- spanning training-time and evaluation-time leakage. Holding out a private test set closes the first alone. The fifth is acquired during the evaluation itself; because it is a property of one run, it must be recorded with the reported score rather than with the benchmark release. We operationalize it as a four-field disclosure protocol in which"unknown"is a valid entry, released under CC BY 4.0 with a JSON Schema, a validator, and worked examples. Two coders external to the design team applied a pre-registered instrument to 41 documents. Per-variable linear-weighted $\kappa$ runs from 0.00 to 0.35 (median 0.21) over 29 main-pass documents against a single-coder test-retest ceiling of 0.84, collapsing under the class skew the registration anticipated; pooling raises it to 0.46 through chance correction rather than better agreement. Two variables fall below the prevalence-robust threshold registered in advance: strata reporting and the acquired type introduced here. Disagreement concentrates on when a variable applies rather than on what a document states. Elicitation budgets are reported in 13% of documents, and no document addresses all five types. The contribution is the taxonomy, the score-side artifact that follows from it, and a pre-registered measurement of instrument reliability and current disclosure.

Johanna Angulo, Víctor Yeste, H. Espinós-Morató · 0 citations
#machine learning Review Aug 2026

Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human Surrogates

LLMs are increasingly used as human surrogates, often on the premise that richer persona data could make them substitutes or exploratory tools for specific individuals. We test this premise across four datasets covering more than 400,000 participants and more than 6,000 survey items and experimental outcomes. LLMs perform well at the aggregate level: their average responses closely align with average human responses to the same items. But this success largely reflects predicting each item's average human response. Once each item's human mean is removed, LLM predictions explain only 3.05% of the remaining respondent-specific variation, far below the 53.6% human test-retest benchmark. Richer personas, model variants, and fine-tuning do not close this gap. In variance analyses, once item means are removed, the reliable remaining signal is person-by-item. It captures how a respondent departs from the mean on a particular item and is about 8.9x larger than the stable person effect. Persona data encode the respondent, but not this item-specific deviation. LLM responses also compress human response distributions, using less spread, fewer response categories, and distorted distributional shapes. We call this pattern item-mean surrogacy. Current LLM surrogates can approximate item averages, but not the distributions or respondent-specific deviations needed to replace individual humans. We propose four empirical tests for LLM-based human-surrogate claims.

Daehwan Ahn, Chengfeng Mao, Dok-Yun Lee · 0 citations
#artificial intelligence Preprint Aug 2026

Learning Simple Test-Time Environments for LLM Web Agents

This work proposes that LLM web agents can learn simple environment observations at test time, and introduces trial steps for agents to decompose a complex environment observation into sub-modules, and implements a label-free learning method, Test-Time Environment Decomposition (TTED), to adapt agent behaviors with experience during inference.

Jun-Xuan Li, Zijun Liu, Zi-Yi Huang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge

This paper provides a practical, auditable, and application-agnostic framework for validating LLM-augmented recipe data, thereby strengthening the foundations of machine-readable food knowledge infrastructures in the era of LLM-generated content.

Saransh Kumar Gupta, Armaan Shah, Lipika Dey et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Emergent Misalignment Is Not Magical

The EM generalization metric is extended from a scalar distance to a dataset-specific generalization direction, which robustly predicts EM models'evilness under semantics-preserving prompt perturbations including appending random tokens and paraphrasing, where other methods do not reliably generalize.

Ming-Xuan Li, Qirun Dai, Hesi Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended Generation

The inference-time ledger that led here: a model-written schematic recap buys judged document integration and nothing buys development; a verifier written into the stream is imitated, 16.4 fabricated verdict lines per notebook.

Roberto I. Ono · 0 citations
#artificial intelligence Preprint Aug 2026

A rigor-matched audit of periodic-step layer skipping for efficient llm inference: conflayers versus swift, with a supplemental analysis of trained routing alternatives

A rigor-matched, three-seed audit of two periodic-step, search-based methods that make this decision online at inference time and re-evaluate it every few generation steps: a confidence-gated early-exit baseline (ConfLayers) and genuine self-speculative decoding (SWIFT, Xia et al. 2024).

Prateek Kumar Sikdar, Arpan Ghosh · 0 citations
#artificial intelligence Preprint Aug 2026

Test-Time Scaling for Scientific Equation Discovery

This work forms LLM-driven equation discovery as an iterative search process that unifies Best-of-N, sequential refinement, tree search, and evolution-style methods under a common compute-allocation view and finds that search width is the dominant allocation parameter.

Hao-Wei Lin, Hubert Lim, Xiang-Yu Wang et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.