Skip to content
Preprint

Isotropic Embedding Perturbations for Robust Vision Language Encoders

Sep 2026 · 0 citations · 62 references
Computer Science

TL;DR

Aether is introduced, a simple plug-in method that applies diffusion-style random perturbations in the embedding space via controlled alpha-mixing, specifically designed to provide isotropic regularization that remains semantically consistent.

Abstract

Data augmentation is fundamental to training modern deep vision and multimodal models. While individual methods, such as RandAug, CutMix, Mixup, RandErase, and DropPath, offer strong regularization effects, their combined use has saturated in performance due to overlapping functionalities, and aggressive pixel-level manipulations may disrupt delicate cross-modal alignment. This saturation motivates the search for a new augmentation axis within the embedding space rather than the input space. We introduce Aether, a simple plug-in method that applies diffusion-style random perturbations in the embedding space via controlled alpha-mixing, specifically designed to provide isotropic regularization that remains semantically consistent. Inspired by feature-space perturbations in language models and image degradation in generative pretraining, Aether induces mild yet effective perturbations that smooth the representations without compromising the fine-grained structural information required for strong vision-language encoders. Across diverse architectures and across multiple recognition tasks, Aether delivers consistent gains over the advanced recipe combining CutMix, Mixup, DropPath, and RandAug---a level of improvement rarely observed with modern augmentation alternatives. Notably, Aether demonstrates superior effectiveness in multi-modal alignment, succeeding where traditional pixel-space augmentations fail by providing a stable, isotropic regularization signal that respects the integrity of the high-dimensional feature space.

View source

Similar papers

#machine learning Preprint Sep 2026

Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models

Post-training quantization (PTQ) enables efficient deployment of large vision-language models (LVLMs), but is typically calibrated on a small set while expected to generalize across diverse downstream tasks. Although recent PTQ methods for LVLMs incorporate sensitivity signals, they still minimize reconstruction loss w...

Min-Chan Kang, Kyeonghye Park, Seungyeon Sa et al. · 0 citations
Preprint Aug 2026

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

UVU effectively synergizes pixel-level visual perception with semantic-level visual understanding, internalizing visual reconstruction capabilities and unlocking the facilitative role of visual supervision in enhancing understanding in the pre-training stage.

Zhe-Han Kan, Xing-Hua Jiang, Yu-Bo Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HiRAE: Hierarchical Representation Autoencoding with Residual Budgets

Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficul...

Xuan-Yu Zhu, Yan Bai, Yang Shi et al. · 0 citations
Preprint Sep 2026

Efficient Quantization-Aware Distillation with Cross-Modal Alignment for Edge Vision-Language Models

Large-scale vision-language models (VLM) such as CLIP enable strong open-vocabulary reasoning, yet deploying these capabilities on resource-constrained edge devices remains challenging. EdgeVL addresses this problem by distilling CLIP representations into lightweight multi-modal encoders and applying quantization-aware...

Jinwoo Jeon, GyuYeop Do, Yunkyu Lim et al. · 0 citations
Preprint Sep 2026

Masked Swingers: Harnessing Data Augmentation to Advance Autoencoders for Self-Supervised Learning

Self-supervised learning (SSL) removes the need for annotations and makes models that are capable across more domains than supervised learning. The autoencoder SSL framework learns by reconstructing its own input after information loss through a bottleneck or noise injection. Masked autoencoders (MAE) are the most succ...

A. Fuller, Scott C. Lowe, Daniel G. Kyrollos et al. · 0 citations
Open access Sep 2026

Mixture of Lie-Group Kernels for Equivariance-Inspired Feature Learning.

Deep visual models typically achieve robustness to geometric transformations through extensive data augmentation or increased model capacity, yet these empirical strategies do not guarantee explicitly equivariant or structurally constrained representations. While group-equivariant CNNs provide principled mechanisms for...

Yao-Xian Yang, Guipeng Lan, Shuai Xiao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.