Shot-Conditioned Vision-Language Adaptation for Effective Harmful Content Detection from Online Short Videos
This work proposes SVLA, a novel π -adaptive strategy to dynamically estimate shot-level anomaly density, replacing rigid selection with calibrated supervision, and employs a shot-conditioned temporal encoder to respect video hierarchy and adopt a dual-path contextual adapter to resolve semantic ambiguity.
Shuai Xu, Zao Qiu, Xuelin Zhu et al.
· 0 citations