Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 51 references
TL;DR
A frequency-aware elastic prototype boundary learning framework, termed SGE-Net, that combines visual and semantic features with dynamic gating to provide more reliable relation features for the above boundary learning and introduces elastic boundary-aware distance calibration during inference.
Abstract
Scene graph generation (SGG) addresses the task of detecting objects in an image and predicting the relationships among them. Although prototype-based methods have recently achieved clear progress on long-tailed SGG, fine-grained low-frequency predicates remain difficult to recognize because their relation features often exhibit larger intra-class variation and more dispersed distributions, making them easily confused with semantically similar high-frequency coarse-grained predicates under a unified prototype-matching rule. To alleviate this issue, we propose a frequency-aware elastic prototype boundary learning framework, termed SGE-Net. Under fixed relation prototypes, the framework learns relation-category-specific boundary scales through explicit frequency compensation and frequency-adaptive virtual sampling, so that relation prediction can exploit not only prototype-center matching but also category-dependent decision-boundary information. During inference, we further introduce elastic boundary-aware distance calibration, enabling the boundary information learned during training to better distinguish relation categories that are easily confused under prototype matching. In addition, we combine visual and semantic features with dynamic gating to provide more reliable relation features for the above boundary learning. Experiments and analyses on Visual Genome and Open Images V6 demonstrate that the proposed method achieves consistent gains in both long-tailed relation prediction and overall evaluation metrics.
This work proposes AMCA, an unbiased SGG framework integrating adaptive multi-prototype learning with cross-modal alignment, which achieves consistently competitive performance across multiple SGG tasks, with particularly strong improvements on the unbiased mR@K metric.
Jinhao Fan, Yuanhao Xi, Chuanping Hu et al.· Journal of King Saud Univers...· 0 citations
: Lightweight semantic segmentation remains challenging because compact backbones often weaken feature discriminability and lose fine-grained boundary details. In DeepLabV3 + -style encoder-decoder architectures, the direct fusion of high-level semantic features and low-level spatial features may introduce semantic-spa...
Wang Zhang, Lanlan Li, Jiayi Xing et al.· Computers, Materials & C...· 0 citations
Confidence-gated relational distillation is proposed, an exemplar-free teacher–student framework that combines feature-level relation preservation with semantic-level background correction that provides an effective balance between old-class retention and novel-class acquisition without introducing replay data or separ...
Lei Wang, Rong-Xiang Liu· Applied Sciences· 0 citations
This work presents a novel self-supervised architecture centered on graph prototype learning that sets a new state-of-the-art on the ARMM dataset with an accuracy of 95.70%, substantiating the efficacy and transferability of prototype-guided self-supervised learning for skeleton-based action representation.
Zhijie Xu, Hong-Wei Chen, Xia Li· International Journal of Mac...· 0 citations
Video Scene Graph Generation (VidSGG) aims to parse subject-predicate-object triplets from videos, a cornerstone for high-level video understanding. However, existing methods are plagued by severe predicate imbalance: a few frequent predicates (e.g., looking at) dominate the training distribution, leading to heavily bi...
Sheng-Hao Li· International Conference on...· 0 citations
Compact object detectors are suitable for resource-constrained visual perception, but their limited representation capacity creates an accuracy gap relative to large models. Conventional detector distillation often relies on prediction-level supervision or a single feature-alignment target, such as response, distributi...
Youngjae Cheong, Jhonghyun An· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.