Size-aware contrastive learning for unbiased video scene graph generation
Abstract
Video Scene Graph Generation (VidSGG) aims to parse subject-predicate-object triplets from videos, a cornerstone for high-level video understanding. However, existing methods are plagued by severe predicate imbalance: a few frequent predicates (e.g., looking at) dominate the training distribution, leading to heavily biased predictions. To address this, we propose SACL-UP (Size-Aware Contrastive Learning with Uncertainty-aware Penalty). Our framework adopts a decoupled two-stage training strategy: Stage 1 pre-trains the model on head predicates to learn robust spatio-temporal semantics; Stage 2 integrates a frequency-adaptive contrastive loss to dynamically calibrate gradients, supported by a hybrid memory bank for hard negative mining. Experiments on Action Genome demonstrate that SACL-UP significantly outperforms state-of-the-art methods, effectively bridging the performance gap between head and tail classes.