Large language models (LLMs) are typically pretrained on short sequences and then extended to work on longer sequences with additional training. However, such LLMs still struggle to further generalize to very long sequences. We propose Randomized YaRN, a training method that improves length generalization by combining YaRN-based positional extrapolation with randomized positional encoding and a length curriculum. During training on short context data, tokens are assigned YaRN positional encodings sampled from a larger position range, exposing the model to out-of-distribution positional representations even on short-context inputs. We evaluate Randomized YaRN on three challenging long-context reasoning benchmarks, BABILong, Multi-Round Coreference Resolution (MRCR), and LongBench v2. When training on data with short context, Randomized YaRN consistently improves reasoning performance on context lengths from 16K to 128K and outperforms standard fine-tuning, with the largest gains appearing at far out-of-distribution lengths. Our results suggest that progressively exposing models to OOD positional distributions provides an effective recipe for generalizable long-context reasoning.
Next-token prediction is the self-supervised signal that trains language models, and every observed prompt token provides the same signal at test time. We study whether this signal can define the inner-loop objective for test-time training (TTT) in pretrained long-context language models. Many TTT architectures require models to be trained with test-time adaptation in mind, limiting their direct applicability to released LLM checkpoints. While recent in-place TTT methods make fast-weight adaptation possible for pretrained LLMs without redesigning the backbone, they leave a central question unresolved: what should each test-time write store? Existing recipes train the fast weight to match a learned local value proxy but they are not directly tied to the self-supervised next-token prediction signal. We introduce Test-Time Training with Next-Token Prediction (TTT-NTP), a drop-in fast-weight adaptation method for pretrained LLMs that instead supervises updates using the model's own next contextual hidden state. This makes each local write follow the same causal computation that supports next-token prediction: the value target is a pointwise linear projection of a single next-position contextual state. On RULER Full-13, averaged over 4k to 32k contexts, TTT-NTP is the only method that consistently improves the released backbone across four models spanning three families and a 0.6-8B size range, by 3.9 points on Llama-3.1-8B, 3.0 on Mistral-7B-v0.3, 4.1 on Qwen3-4B, and 2.9 on Qwen3-0.6B. On the real-world LongBench-v2 long-document QA benchmark, TTT-NTP improves over the base model by 5.6 points on Llama-3.1-8B and 3.7 on Mistral-7B-v0.3, while preserving commonsense and knowledge performance. Our code is publicly available at https://github.com/yancyou/TTT-NTP.
PhysAssistBench is introduced, a benchmark for interactive doctor-patient-EHR assistance that uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality.
T. Du, Peijie Yu, Sihan Shang et al.· arXiv.org· 0 citations
A controlled protocol for evaluating answer stability is introduced: after a model answers a multiple-choice question correctly, it is challenged with a coherent argument for an incorrect option and measured whether the model flips, finding that self-attribution consistently increases flip rates and pooling wrong-answer arguments across models yields stronger adversarial challenges.
Nafiseh Nikeghbal, Amir Hossein Kargaran, Shaghayegh Kolli et al.· arXiv.org· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This survey reviews Multimodal Code Intelligence, covering systems that generate, edit, refine, or reason with code under visually grounded inputs and outputs and organizes benchmarks and methods into four domains: Graphical User Interface, Scientific Visualization, Structured Graphics, and Frontier Tasks and Frameworks.
SciOrch is presented, a framework that trains a lightweight 8B model to orchestrate frontier LLMs for scientific reasoning, and attains the best accuracy on both SGI and SFE with less than half the API cost of typical multi-agent methods.
Jingru Guo, Xiangyuan Xue, Lian Zhang et al.· arXiv.org· 0 citations
The results suggest that skills and workflow memory can be useful in specific domains, but their apparent gains often vanish against a budget-matched actor, and that run-to-run variance materially affects outcomes and should be reported as a core evaluation criterion for online web agents.
Sina Hajimiri, Masih Aminbeidokhti, J. Dolz et al.· arXiv.org· 1 citation
This work introduces SemCog Bench, a curated benchmark of 1,858 Arabic--Hebrew word pairs with sentence-level annotations for cognate identification and semantic disambiguation and finds that context and scale yield model-dependent gains, while original-script inputs generally perform best.
Junhong Liang, Noor Abo Mokh, Bashar Alhafni· arXiv.org· 0 citations
EvoBrowseComp is introduced, an evolving benchmark of 400 English and 400 Chinese contamination-free complex questions synthesized via live-web traversal that establishes a scalable paradigm for auto-updatable, high-difficulty benchmarking that keeps pace with both evolving world knowledge and advancing agent capabilities.
BioDivergence offers a more faithful way to distinguish contextual divergence from direct contradiction and to separate article-level memorization from genuine task learning.
Elias Hossain, S. S. Jennifer, Sabera Akter Bushra et al.· arXiv.org· 0 citations
This work proposes a post-training alignment method that comprehensively improves the interactivity of full-duplex spoken dialogue models through RL, and addresses the four canonical axes of interactivity: pause handling, turn-taking, backchanneling, and user interruption.
Atsumoto Ohashi, Neil Zeghidour, Alexandre Défossez et al.· arXiv.org· 0 citations
QK-Restore consistently restores long-context capability at zero training cost while preserving reasoning performance; for instance, on HypeNet-5B it improves S3@256K from $65.4\%$ to $76.4\%$ while maintaining strong reasoning performance.
Xinyu Zhou, Boyuan Zhu, Yi Xu et al.· arXiv.org· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.