Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

PhiZero: A World Model Built Around Physical Language

We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans'ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.

Shuyao Shang, Yuqi Wang, Ruopeng Gao et al. · 1 citation
Jun 2026

Bridging Video Understanding and Generation in a Unified Framework

Vega is a unified framework that bridges video understanding and generation and employs a hybrid architecture combining autoregressive (AR) prediction with diffusion-based rendering, providing a structured representation that guides the diffusion module in rendering dense, high-resolution video frames.

Yuqi Wang, Runyi Li, Ruoyu Feng et al. · 1 citation