Skip to content

Author

Jike Zhong

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Across 95K instances at four grid densities (4$\times$4 to 16$\times$16), we find that \textbf{zero-shot VLMs largely lack geometric reasoning}: only one of five frontier models (GPT-5.5) exceeds random baseline on 4$\times$4 puzzles, while all others perform at chance level. While supervised fine-tuning achieves $>$97\% on 4$\times$4, \textbf{all models collapse on larger grids}: GPT-5.5 drops from 70\% to near-random on 8$\times$8, and even fine-tuned models fall below 5\% on 12$\times$12. This ``scaling cliff''suggests current architectures cannot maintain consistent constraint satisfaction as the number of pieces increases. \ours{} establishes scalable geometric reasoning as an open challenge for vision-language models.

Shawn Li, Wei Yang, Jike Zhong et al. · 0 citations

Bridging the Granularity Gap: Object-Centric Masking for Contextual Visual Learning

This work proposes to model objects as a stronger semantic unit for visual prediction, encouraging the encoder to learn the global context and semantics among visual elements, and shows that an object-centric objective reduces pixel-averaging shortcuts and yields more globally coherent and context-consistent representations.

Jike Zhong · 0 citations
Preprint Jul 2026

Agentic Context Learning with Self-Discovered Specification

The results suggest context learning hinges on not only content acquisition but also specification acquisition, and designs a deliberately simple intervention PSCI (private specification-contract induction) which extracts local specifications and enforces them through adversarial checking and repair.

Jike Zhong, Ming Li, Yuxiang Lai et al. · 1 citation