Skip to content

Author

Zhiyuan Shi

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations natu...

Weiliang Chen, Haowen Sun, Jun Gao et al. · 2 citations
Preprint Aug 2026

GEB-Bench: Abstract Structures Told in Many Voices

Evaluating twelve open and proprietary models, it is found that abstraction failure is lawful, and the central finding is a gap between recognition and cross-voice mapping: models identify a structure within one voice far better than they carry it across voices; every model pays this tax, and mapping strong enough to n...

Tong Zhang, Zhiyuan Shi, Yun Peng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.