This work introduces DR.WILSS, an innovative approach to address catastrophic forgetting in continual learning using diffusion-based generative replay, which leverages language clues to guide the diffusion process, employing self-inpainting and regularization techniques to efficiently produce replay data, aiding the learning process.
Abstract
Weakly supervised class-incremental semantic segmentation (WILSS) aims to train a segmentation model over multiple steps, each introducing new concepts to be learned with only image-level supervision. We introduce DR.WILSS, an innovative approach to address catastrophic forgetting in continual learning using diffusion-based generative replay. Our framework leverages language clues to guide the diffusion process, employing self-inpainting and regularization techniques to efficiently produce replay data, aiding the learning process. By generating high-quality replay data, the information from previously learned classes can be preserved during continual updates, a critical challenge in incremental learning scenarios. To further align the statistics of replay data with those of training samples, we apply LoRAs to the generative model. Experimental results demonstrate state-of-the-art performance across multiple benchmarks and generative architectures, while avoiding storage of training data and the use of additional resource-demanding tools during training. The proposed technique enables an optimal tradeoff between training complexity and inference-time accuracy, making DR.WILSS a promising solution for real-world applications.
DIFFCZSL is proposed, a diffusion-augmented framework that injects generative priors from pre-trained diffusion models into CLIP-based CZSL pipelines and highlights the complementary strengths of generative diffusion representations and discriminative vision-language models for compositional generalization.
Hangyu Tian, Zhen-Qi He, Yanghao Wang et al.· 0 citations
Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-temporal nature. In this...
Enrico Pallotta, Sina Raoufi, Lars Doorenbos et al.· 0 citations
Open-world entity segmentation aims to predict masks for arbitrary objects across domains and at multiple granularities, from parts to whole objects. In this setting, SAM sets a strong standard: trained on SA-1B, comprising 11M images and over 1B carefully annotated masks, it achieves remarkable zero-shot performance....
Christoph Hümmer, Joachim Sicking, Fabian Hüger et al.· 0 citations
Results on the UC Merced (UCM) and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches and provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.
M. L. Mekhalfi, M. M. Al Rahhal, Y. Bazi et al.· IEEE Geoscience and Remote S...· 0 citations
GenView is presented, a controllable framework that augments the diversity of positive views leveraging the power of pretrained generative models while preserving semantics, and an adaptive view generation method that dynamically adjusts the noise level in sampling to ensure the preservation of essential semantic meani...
Xiao-Jie Li, Yibo Yang, Xiang-Tai Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.