Combining mean-field theory and random matrix theory, a direct link between correlation propagation and the Neural Tangent Kernel (NTK) is established that governs learning in the sequential limit of infinitely wide, infinitely deep networks.
andrea Combette, Nelly Pustelnik, Antoine Venaille· 0 citations
The exact selection time for an isolated cycle of NOTEARS and DAGMA is derived, and a truth-free separation statistic predicts selection time on 320 official NOTEARS/DAGMA trajectories.
Rui Wu, Zongyuan Chen, Hong Xie et al.· 0 citations
This work introduces a problem setting for failure attribution under distribution shift, which enables the models not only to detect out-of-distribution samples, but also to find out the reason for their failure.
This work introduces CrystalGRPO, a CSP-aligned post-training framework that extends existing ODE-to-SDE policy constructions to the joint coordinate--lattice state and provides two operating modes: CrystalGRPO-Q, which prioritizes single-draw recovery, and CrystalGRPO-C, which combines full-trajectory reference regularization with a coverage-aware group advantage to preserve finite-budget target recovery.
Kaixiang Su, Hongfei Xue, Qiang Zhu· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
A framework for designing the reference process that achieves tractable time-dependent conditional distributions and then construct the reference process realizing them as its marginals is proposed, revealing that score matching is not fundamental to diffusion-model training but instead emerges naturally through reversal of the reference process.
It is found that configuration shift consistently erodes CP validity, often driving empirical coverage below the target, and coverage lower bounds are derived that attribute this loss to a discrepancy between calibration and test score distributions.
Yuqicheng Zhu, Jia-Lin Yu, Lin Li et al.· 0 citations
This work empirically validate predictions through exact calculations and controlled stochastic optimization of the worst-case joint-distribution error permitted by excess risk at most $\varepsilon$-identifiability modulus using an $\varepsilon$-identifiability modulus.
SERUM is the first system to produce interpretable process models from unstructured egocentric screen video without manual annotation, opening a scalable pathway for user modeling and behavioral understanding in the wild.
Andy J. Phu, Karin de Langis, James C Mooney et al.· arXiv.org· 0 citations
A2TTA is proposed, an Anchored-and-Agile Test-Time Adaptation framework for evolving traffic sensor networks, which transforms topology-induced forecasting errors into an expandable output calibration problem and separates tem- poral adaptation into persistent global correction and agile context-specific specialization.
Du Yin, Xiachong Lin, Yuejie Tan et al.· arXiv.org· 0 citations
It is found that training on cloned data outperforms raw cross-lingual transfer for depression and anxiety detection on real Japanese speech, suggesting voice cloning is a promising direction for augmenting clinical speech data in low-resource languages.
R. Polle, Owen Parsons, George Fairs et al.· arXiv.org· 0 citations
The assumption that attention shows which context the answer depends on is tested on retrieval tasks where the evidence is known exactly, by masking context and measuring whether the answer changes.
It is proved that no estimator computable from the information a deterministic scheme retains is consistent for its own eviction error: evicted values can be altered so that everything retained is unchanged while the true attention-output error grows without bound.
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026