Preprint
Jul 2026
Self-Supervised Visual Representation Learning: Pretrain-Finetuning or Joint Training?
This work systematically investigate whether jointly optimizing the self-supervised and supervised objectives during training provides a better alternative, and finds that JT consistently improves data and training efficiency while being robust in low-label settings, while PFT is more reliable in more specialized domains.
Nusrat Munia, Tyler Ward, Nishat Nayla et al.
· 0 citations