Sep 2026· IEEE Transactions on Neural Networks and Learning Systems· Vol PP, pp. 1-14· 0 citations
Medicine
TL;DR
In CSIR, a novel method for learning invariant representations is proposed, termed Cauchy-Schwarz invariant representation learning (CSIR), which implements CMI and MI by leveraging the classic Cauchy-Schwarz divergence without the need for extra neural networks.
Abstract
Invariant risk minimization (IRM) has recently emerged as a promising method for invariant representation learning. The effectiveness of IRM relies on the assumption of the cross-domain overlap of invariant features. Some mainstream methods satisfy the cross-domain overlap assumption of invariant features by introducing the principle of the information bottleneck (IB) into IRM to compress feature information while ensuring the maximum correlation between the features and labels. However, these methods rely on extra neural networks to implement IRM and the IB, including conditional mutual information (CMI) and mutual information (MI) penalties. This means that extra neural networks will increase parameter complexity and instability in the network training process. Moreover, these methods presuppose data environment partitioning, rendering them challenging to generalize to continuous domain environments. To solve this problem, we propose a novel method for learning invariant representations, termed Cauchy-Schwarz invariant representation learning (CSIR). In CSIR, we implement CMI and MI by leveraging the classic Cauchy-Schwarz (CS) divergence without the need for extra neural networks. Moreover, the new regularization terms favor continuous random variables, which makes them amenable in a continuous domain environment and eliminates the requirement of presupposed domain environment partitions. We conducted experiments on five different datasets and demonstrated that our approach can learn invariant representations more efficiently.
Divergence-based regularization and Sharpness-Aware Minimization (SAM) are two prominent approaches for improving generalization in deep learning, both motivated by robustness to perturbations. However, their relationship has remained largely unexplored. Building on classical second-order expansions of $f$-divergences,...
Constraint-Aligned Subspace Transformation (CAST), a novel framework that restricts metric adaptation to a low dimensional subspace strictly induced by pairwise constraints, is proposed, and an implicit update formulation that enables distance evaluation in O (r) time, effectively reducing memory complexity to O (nr).
Fei Wang, Le Li, P. Fränti· Proceedings of the 32nd ACM...· 0 citations
Muon emerges as a strong competitor of the AdamW for LLM pretraining, because the matrix-wise update it employs can potentially incur smaller second-order penalty than the once dominating AdamW, which performs coordinate-wise update. However, the spectral flattening procedure in Muon is quite debatable since it discard...
Qiao-Zhe Zhang, Jun Sun, Ying-Zhuang Liu· 0 citations
This work introduces a general kernel-based encoder-decoder framework for operator learning that separates observation, representation, learning, and reconstruction, and develops this framework for multi-input, multi-output operator learning, where operators map between products of potentially distinct function spaces.
Adrien Weihs, Chun-Yang Liao, Jingmin Sun et al.· 0 citations
Instance-wise feature selection (IWFS) identifies informative features for each instance, improving generalization by discarding irrelevant information and enhancing interpretability through personalized explanations. Most IWFS methods adopt a selector--predictor architecture, where a selector generates instance-specif...
Lu Sun, Jun Sakuma· Proceedings of the Thirty-Fi...· 0 citations
Deep visual models typically achieve robustness to geometric transformations through extensive data augmentation or increased model capacity, yet these empirical strategies do not guarantee explicitly equivariant or structurally constrained representations. While group-equivariant CNNs provide principled mechanisms for...
Yao-Xian Yang, Guipeng Lan, Shuai Xiao et al.· IEEE Transactions on Neural...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.