Preprint
Jul 2026
SigLIP-HD by Fine-to-Coarse Supervision
This work studies an interesting problem: how to achieve fine visual perception under lower cost without larger images, and builds this framework on the advanced SigLIP 2 model, which consistently delivers stronger results than the baseline model, especially on OCR-related tasks.
Lihe Yang, Zhen Zhao, Hengshuang Zhao
· 0 citations