Skip to content

Category

machine learning

5,133 papers

#machine learning Preprint Aug 2026

RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling

Frozen RNA-type evaluations show that RIBOSPAN learns state-of-the-art RNA representations, with a particularly clear advantage on long RNAs, and emerges as the strongest encoder-only RNA foundation model, achieving state-of-the-art performance in both full-transcript biological property prediction and zero-shot mutation-fitness modeling.

Ziyuan Wang, Bohao Tang, Fei Zhang et al. · 0 citations
#machine learning Preprint Jul 2026

SpecGradFilter: A Spectral Gradient Filtering Framework for Taming Federated Heterogeneity

This work revisits client drift from a novel frequency-domain perspective and uncovers a critical Spectral Bias of Drift: inter-client gradient divergence is predominantly concentrated in low-frequency components which encode client-specific distributional shifts, while high-frequency components representing fine-grained features remain relatively consistent.

Liyang Yuan, Yibo Yang, Dandan Guo et al. · 0 citations
#machine learning Open access Jun 2026

ERP-XTTN: Interpretable Prototype-Guided Cross-Attention for Cross-Subject ERP Classification

OBJECTIVE Interpretable brain-computer interface classifiers that generalize across subjects without calibration remain an open challenge. We evaluated whether prototype-based cross-attention can provide competitive, inherently interpretable event-related potential (ERP) classification across diverse paradigms under deployment-compatible conditions. APPROACH We propose ERP-XTTN (ERP Cross-Attention), a cross-attention architecture that routes input electroencephalographic peaks to fixed difference-wave prototypes via query-key-only cross-attention with no value projection. Classification is based directly on prototype similarity and a separate measure of component amplitude, so that the prototype content contributes to every decision by construction. Prototypes are derived automatically from prominent extrema in the training-fold grand-average difference wave. We evaluated across three public sources (BNCI Horizon 2020, HRI Cursor, and ERP CORE) encompassing eight ERP components (ERN, LRP, ErrP, N170, P300, N2pc, MMN, N400). Evaluations used leave-one-subject-out (LOSO) cross-validation with causal filtering at a three-channel montage, compared against EEGNet, EEG-Deformer, ERP Prototypical Matching Net (EPMN), and xDAWN with Riemannian geometry (xDAWN+RG). MAIN RESULTS At three channels, the mean performance gap between the best baseline and ERP-XTTN was 0.025 area under the receiver operating characteristic curve (AUROC). Prototype interventions confirmed that decisions depend on the physiological content of the prototypes rather than on the routing attention pattern alone. False positives morphologically resembled true positives more than true negatives did across all datasets, indicating classification errors are neurophysiologically explicable. SIGNIFICANCE ERP-XTTN generalizes across diverse ERP morphologies under causal, calibration-free conditions, while retaining competitive performance and decisions that depend directly on physiological prototype content at a three-channel montage. Unlike post-hoc explanation methods for black-box models, the basis of each decision is directly observable in the trained model itself. To our knowledge, this is the first epoch-level LOSO benchmark on ERP CORE.

Charlotte Genevier Wyman, L. Hirshfield · 0 citations

Can Subgraph Explanations Be Weaponized to Steal Graph Neural Networks?

This work presents the first model extraction attack specifically designed for graph classification under strict black-box constraints, which uses model explanation outputs to guide Monte Carlo edge sensitivity estimation toward decision boundaries, with Hoeffding concentration guarantees on estimation accuracy.

Ojas Nimase, Jia-Te Li, Yuelei Zhao et al. · 0 citations

Learned Relay Representations for Forward-Thinking Discrete Diffusion Models

Relay introduces a differentiable per-token channel that passes information between forward passes and is trained via truncated backpropagation through time (BPTT), demonstrating that state-of-the-art DLMs can be explicitly trained to relay latent information forward across decoding steps, advancing the performance-latency Pareto frontier.

Benjamin Rozonoyer, Jacopo Minniti, Dhruvesh Patel et al. · 0 citations

SemEnrich: Self-Supervised Semantic Enrichment of Radiology Reports for Vision-Language Learning

A self-supervised data enrichment method that leverages semantic clustering of report sentences that enrich the findings in the medical reports in the training set by adding positive/neutral observations from different clusters in a self-supervised manner.

Halil Ibrahim Gulluk, Olivier Gevaert · 1 citation

Simplex-to-Euclidean Bijection for Conjugate and Calibrated Multiclass Gaussian Process

This approach uses Aitchison geometry to map simplex-valued class probabilities to an unconstrained Euclidean representation, turning classification into a GP regression problem with fewer latent dimensions than standard multi-class GP classifiers.

Bernardo Williams, Harsha Vardhan Tetali, Arto Klami et al. · 0 citations
#machine learning Open access Apr 2025

Bayesian Experimental Design for Model Discrepancy Calibration: An Auto-Differentiable Ensemble Kalman Inversion Approach

This work clarifies the trade-offs between KL divergence and Wasserstein metrics for the utility function and provides guidelines for selecting suitable criteria in practical BED applications.

Huchen Yang, Xinghao Dong, Jin-Long Wu · 3 citations
#machine learning Conference Open access Mar 2018

Aspiration-based Perturbed Learning Automata

A novel payoff-based learning scheme for distributed optimization in repeatedly-played strategic-form games and it is shown that payoff-dominant Nash equilibria are the only stochastically stable states.

Georgios C. Chasparis · 1 citation

One Model for All: Universal Pre-training for EEG based Emotion Recognition across Heterogeneous Datasets and Paradigms

This work proposes 'One Model for All', a universal pre-training framework for EEG analysis across disparate datasets, and paves the way for more universal, scalable, and effective pre-trained models for diverse EEG analysis tasks.

Xiang Li, You Li, Yazhou Zhang · 0 citations

From tech blogs

See all →
GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.