The assumption that attention shows which context the answer depends on is tested on retrieval tasks where the evidence is known exactly, by masking context and measuring whether the answer changes.
It is proved that no estimator computable from the information a deterministic scheme retains is consistent for its own eviction error: evicted values can be altered so that everything retained is unchanged while the true attention-output error grows without bound.
This work analyzesparse autoencoder features across six models and three SAE families and zero-ablate at full layer depth, finding cross-family claims are sensitive to training methodology, not just activation function or scale.
Seonglae Cho, Zekun Wu, Kleyton Da Costa et al.· arXiv.org· 1 citation
Data-Adaptive Lower-Rank Adaptation (DALorRA), a simple and effective variational Bayesian sparse framework that shifts the paradigm of uncertainty quantification from the dense parameter space to the lightweight rank level of low-rank adaptation (LoRA).
Ji-Jie Zhang, Zhenjiang Ren, Quan Zhang et al.· arXiv.org· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
The canonical Crawford-Sobel cheap-talk model is turned into a pre-specified benchmark for LLM honesty under preference misalignment, in which theory supplies an exact oracle.
Student-Centric Answer Sampling (SCAS) is proposed, a framework that selects from verified teacher-generated answers according to their estimated student-centric learning cost and is derived by a token-wise gradient decomposition and used to guide answer selection during training.
Zhengyu Hu, Zheyuan Xiao, Linxin Song et al.· 0 citations
A dynamic per-layer scalar derived by adapting the LARS/LAMB trust-ratio principle to the orthogonalized setting, where the standard denominator candidates---the raw momentum norm or the polar-factor norm---either live in the wrong unit space or carry no update-scale information.
These results expose a rate-granularity trade-off: PairAlign does not uniformly outperform denser tokenizers on every local metric, but provides a lower-rate symbolic interface preserving ordered and relational structure.
Personalized GRPO is introduced, a novel alignment framework that decouples advantage estimation from immediate batch statistics and achieves faster convergence and higher rewards than standard GRPO, thereby enhancing its ability to recover and align with heterogeneous preference signals.
Jialu Wang, Heinrich Peters, A. Butt et al.· arXiv.org· 1 citation
MUSE (Multimodal Unified Safety Evaluation), an open-source, browser-based, run-centric platform for multimodal safety evaluation, demonstrates the value of run-centric, fine-grained evaluation for characterizing multimodal safety behavior beyond a single binary success metric.
This work introduces Constrained GRPO, a Lagrangian-based extension of GRPO for constrained policy optimization, and addresses the coupling induced by reward scalarization by scalarizing standardized advantages rather than rewards.
Roger Girgis, Rodrigue de Schaetzen, Luke Rowe et al.· arXiv.org· 2 citations· ⚡1
It is suggested that diffusion post-training selectively preserves or reorganizes inherited computation according to task structure, rather than uniformly replacing autoregressive mechanisms.
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.