Together, the results clarify how the input geometry shapes the kernel features and fundamentally impacts its generalization properties and clarify how the input geometry shapes the kernel features and fundamentally impacts its generalization properties.
L. Rizzi, Arie Wortsman Zurich, Bruno Loureiro· 0 citations
The results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.
For finite-dimensional linear inverse problems where the variables are Gaussian, it is well-known that the minimum-mean-square error estimator takes the form of a regularized least-squares data fit. In this chapter, we show that this equivalence extends to a much broader infinite-dimensional setting where generalized splines take the role of linear regressors and generalized Gaussian processes on a nuclear space $S$ are the counterpart of Gaussian random vectors. The scope of this extension is of the same nature as the switch from the classic notion of function to that of a distribution, also known as a"generalized function."Our formalism involves a whitening/regularization operator $L: S\to S'$ whose continuous extension induces a native Hilbert space $H\subset S'$ that plays a central role in our characterization. The presentation is self-contained for the most part and remarkably general and powerful. It allows for the recovery of all known instances of such equivalences; in particular, the methods involving innovations and reproducing-kernel Hilbert spaces developed by Kailath and his students, and the mathematical correspondence between fractional splines and Mandelbrot's fractional Brownian motion (fractals), with the former being the optimal estimators of the latter. It also covers general Bayesian methods for the resolution of infinite-dimensional inverse problems.
This work shows that Sliding Window Attention (SWA) with sinks performs as well or better than post-trained Linear Attention models, and recommends switching to SWA instead of post-training linear models.
Alexia Jolicoeur-Martineau, R. Sukthanker, Pashmina Cameron et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work proposes the first video-language-model post-training technique for mistake detection, which uses a tailored reward function to encourage the model to identify discrepancies between an instruction and the corresponding video, and generalizes especially well to unseen procedures.
Federico Spurio, Olga Zatsarynna, Lars Doorenbos et al.· 0 citations
This work employs the Bures metric as a local preconditioner and uses the mean Uhlmann curvature to develop an achievable-precision aggregation rule that dynamically down-weights unreliable clients and establishes theoretical guarantees by proving a convergence theorem and a variance-dominance proposition.
A framework for this localization problem is developed by assigning to each observation its conditional or marginal contribution across random statistical contexts, which connects resampling diagnostics and data valuation to projection theory and event-level anomaly detection.
State estimation in magnetohydrodynamic flows is critical for real-time monitoring of liquid metal blankets in tokamak fusion reactors. Due to the multiphysics nature of these phenomena, high-fidelity simulations are computationally prohibitive for real-time applications. This work investigates a data- driven Reduced Order Model framework: the Shallow Recurrent Decoder (SHRED) coupled with Principal Component Analysis, to map sparse temperature measurements to the full thermo-hydraulic system's state. The major contribution of this work lies in the two-parameter analysis of a fully three-dimensional domain representative of the DEMO breeding blanket configuration. Here, the flow is subjected to an external magnetic field varying in direction and intensity and is hindered by two cylinders acting as a water-cooling system, which impose a temperature boundary condition on their surfaces. This double-parametric magnetic variation induces nonlinear transitions in the flow dynamics, ranging from chaotic behavior at low magnetic field intensities to laminarized regimes at high intensities, characterized by the formation of asymmetric side layers at an inclination angle of 30 degrees. SHRED reconstruction maintains a mean relative error of approximately 5% for the temperature, pressure, and velocity fields. This accuracy is maintained across both weak and strong magnetic fields, ranging from 0.075 T to 0.300 T, and for inclination angles from 5 to 30 degrees, reflecting its dominant toroidal component. These errors are only slightly larger than the lower error bound dictated by low-rank truncation. The results establish SHRED as a reliable state estimator for complex and realistic engineering applications involving completely unseen parametric scenarios and validate it as an accurate real-time state estimation technique suitable for online monitoring and control of real facilities.
Claudio Scardino, S. Riva, C. Introini et al.· 0 citations
It is shown that, in the sample limit, I-FLOP recovers a DAG in the same interventional Markov equivalence class as the data-generating DAG, where it performs favorably in terms of both performance and run time.
Investigating an explainable DR classification framework using vision foundation models and multiple transfer learning strategies demonstrates that foundation models, particularly DINOv2, can provide strong predictive performance, while LoRA offers a parameter-efficient alternative to full fine-tuning.
This work proposes Explainable Probing of Cross-Domain Sparse Embeddings (EXPOSE), a framework that uses Sparse Autoencoders (SAEs) as an explainable bottleneck to identify and suppress domain-specific components in VFM embeddings.
Anja Witte, M. Lennartz, Jan Baumbach et al.· 0 citations
This work presents a deep learning-based web system for automatic identification of Bangladeshi mango varieties and integrated the model into a Streamlit web application that enables users to upload a mango image and receive a predicted variety with class probabilities.
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.