Skip to content
Conference

Size-aware contrastive learning for unbiased video scene graph generation

Sep 2026 · International Conference on Computer Vision, Graphics, and Artificial Intelligence (CVGAI 2026) · 0 citations

Abstract

Video Scene Graph Generation (VidSGG) aims to parse subject-predicate-object triplets from videos, a cornerstone for high-level video understanding. However, existing methods are plagued by severe predicate imbalance: a few frequent predicates (e.g., looking at) dominate the training distribution, leading to heavily biased predictions. To address this, we propose SACL-UP (Size-Aware Contrastive Learning with Uncertainty-aware Penalty). Our framework adopts a decoupled two-stage training strategy: Stage 1 pre-trains the model on head predicates to learn robust spatio-temporal semantics; Stage 2 integrates a frequency-adaptive contrastive loss to dynamically calibrate gradients, supported by a hybrid memory bank for hard negative mining. Experiments on Action Genome demonstrate that SACL-UP significantly outperforms state-of-the-art methods, effectively bridging the performance gap between head and tail classes.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.