Skip to content

Category

natural language processing

2,394 papers

#natural language process... Preprint Open access Sep 2026

Randomized YaRN Improves Length Generalization for Long-Context Reasoning

Large language models (LLMs) are typically pretrained on short sequences and then extended to work on longer sequences with additional training. However, such LLMs still struggle to further generalize to very long sequences. We propose Randomized YaRN, a training method that improves length generalization by combining YaRN-based positional extrapolation with randomized positional encoding and a length curriculum. During training on short context data, tokens are assigned YaRN positional encodings sampled from a larger position range, exposing the model to out-of-distribution positional representations even on short-context inputs. We evaluate Randomized YaRN on three challenging long-context reasoning benchmarks, BABILong, Multi-Round Coreference Resolution (MRCR), and LongBench v2. When training on data with short context, Randomized YaRN consistently improves reasoning performance on context lengths from 16K to 128K and outperforms standard fine-tuning, with the largest gains appearing at far out-of-distribution lengths. Our results suggest that progressively exposing models to OOD positional distributions provides an effective recipe for generalizable long-context reasoning.

Manas Mehta, Fangcong Yin, Greg Durrett · 0 citations
#natural language process... Preprint Open access Sep 2026

Test-Time Training with Next-Token Prediction

Next-token prediction is the self-supervised signal that trains language models, and every observed prompt token provides the same signal at test time. We study whether this signal can define the inner-loop objective for test-time training (TTT) in pretrained long-context language models. Many TTT architectures require models to be trained with test-time adaptation in mind, limiting their direct applicability to released LLM checkpoints. While recent in-place TTT methods make fast-weight adaptation possible for pretrained LLMs without redesigning the backbone, they leave a central question unresolved: what should each test-time write store? Existing recipes train the fast weight to match a learned local value proxy but they are not directly tied to the self-supervised next-token prediction signal. We introduce Test-Time Training with Next-Token Prediction (TTT-NTP), a drop-in fast-weight adaptation method for pretrained LLMs that instead supervises updates using the model's own next contextual hidden state. This makes each local write follow the same causal computation that supports next-token prediction: the value target is a pointwise linear projection of a single next-position contextual state. On RULER Full-13, averaged over 4k to 32k contexts, TTT-NTP is the only method that consistently improves the released backbone across four models spanning three families and a 0.6-8B size range, by 3.9 points on Llama-3.1-8B, 3.0 on Mistral-7B-v0.3, 4.1 on Qwen3-4B, and 2.9 on Qwen3-0.6B. On the real-world LongBench-v2 long-document QA benchmark, TTT-NTP improves over the base model by 5.6 points on Llama-3.1-8B and 3.7 on Mistral-7B-v0.3, while preserving commonsense and knowledge performance. Our code is publicly available at https://github.com/yancyou/TTT-NTP.

Xuan Ouyang, Zefan Cai, Junjie Hu · 0 citations
#artificial intelligence Review Jun 2026

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

PhysAssistBench is introduced, a benchmark for interactive doctor-patient-EHR assistance that uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality.

T. Du, Peijie Yu, Sihan Shang et al. · 0 citations

Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs

A controlled protocol for evaluating answer stability is introduced: after a model answers a multiple-choice question correctly, it is challenged with a coherent argument for an incorrect option and measured whether the model flips, finding that self-attribution consistently increases flip rates and pooling wrong-answer arguments across models yields stronger adversarial challenges.

Nafiseh Nikeghbal, Amir Hossein Kargaran, Shaghayegh Kolli et al. · 0 citations

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

This survey reviews Multimodal Code Intelligence, covering systems that generate, edit, refine, or reason with code under visually grounded inputs and outputs and organizes benchmarks and methods into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Frameworks.

Xuanle Zhao, Qiushi Sun, Jingyu Xiao et al. · 5 citations

SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks

SciOrch is presented, a framework that trains a lightweight 8B model to orchestrate frontier LLMs for scientific reasoning, and attains the best accuracy on both SGI and SFE with less than half the API cost of typical multi-agent methods.

Jingru Guo, Xiangyuan Xue, Lian Zhang et al. · 0 citations

Are Online Skill and Memory Modules Always Worth Their Tokens? A Budget-Constrained Study of Web Agents

The results suggest that skills and workflow memory can be useful in specific domains, but their apparent gains often vanish against a budget-matched actor, and that run-to-run variance materially affects outcomes and should be reported as a core evaluation criterion for online web agents.

Sina Hajimiri, Masih Aminbeidokhti, J. Dolz et al. · 1 citation

When Similar Means Different: Evaluating LLMs on Arabic-Hebrew Cognates

This work introduces SemCog Bench, a curated benchmark of 1,858 Arabic--Hebrew word pairs with sentence-level annotations for cognate identification and semantic disambiguation and finds that context and scale yield model-dependent gains, while original-script inputs generally perform best.

Junhong Liang, Noor Abo Mokh, Bashar Alhafni · 0 citations

EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge

EvoBrowseComp is introduced, an evolving benchmark of 400 English and 400 Chinese contamination-free complex questions synthesized via live-web traversal that establishes a scalable paradigm for auto-updatable, high-difficulty benchmarking that keeps pace with both evolving world knowledge and advancing agent capabilities.

Yun-Han Wang, Jiaan Wang, Lianzhe Huang et al. · 2 citations

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

This work proposes a post-training alignment method that comprehensively improves the interactivity of full-duplex spoken dialogue models through RL, and addresses the four canonical axes of interactivity: pause handling, turn-taking, backchanneling, and user interruption.

Atsumoto Ohashi, Neil Zeghidour, Alexandre Défossez et al. · 0 citations

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It

QK-Restore consistently restores long-context capability at zero training cost while preserving reasoning performance; for instance, on HypeNet-5B it improves S3@256K from $65.4\%$ to $76.4\%$ while maintaining strong reasoning performance.

Xinyu Zhou, Boyuan Zhu, Yi Xu et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.