Skip to content

Author

Türkay Yildirim

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Spatial Knowledge Distillation in Video Models via Vision-Language Guided Zero-Shot Pretraining

Video action recognition tasks require large-scale labeled data to achieve high performance when pretraining is not used. However, annotating video data is costly and time-consuming. To reduce this dependency on labeled data, transfer learning approaches are commonly employed. Vision–language models, which learn generalizable visual representations from large-scale data, are effective at capturing spatial information and thus serve as suitable teacher models for spatial knowledge distillation. In this work, we propose a knowledge distillation-based pretraining approach that leverages zero-shot predictions from vision–language models to initialize video models. In this framework, the resulting soft class distributions are used as supervisory signals and transferred to the student video model. Experimental results on the UCF101 and HMDB51 datasets demonstrate that the proposed method provides an effective weight initialization strategy and yields consistent performance improvements.

A. Çelik, Türkay Yildirim · 0 citations