Skip to content

Author

Yi-Hao Liu

We have 7 of 49 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Towards Physics-Faithful Generation of Scientific Diagrams

Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams, and curate and structurally annotate 4.3 million physics images, carry expert-level annotation, and adapt a unified multimodal backbone.

Ming-Hui Zhang, Jinxin Shi, Yi-Fan Chang et al. · 0 citations
Preprint Jul 2026

From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space

NapMem is introduced, a framework for learning to use long-term user memory as a structured action space rather than passively retrieved context, and suggests that long-term user memory benefits from coupling structured storage with a learned policy for using memory at the appropriate granularity.

Yue Xu, Yutao Sun, Yihao Liu et al. · 3 citations
Jul 2026

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

This work introduces Skill Self-Play (Skill-SP), a co-evolutionary framework comprising a proposer, a solver, and a dynamic skill controller that effectively bridges the gap between structured verification and open-ended exploration.

Siyuan Huang, Pengyu Cheng, Haotian Liu et al. · 2 citations
#artificial intelligence Preprint Aug 2026

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

This work formalizes probabilistic alignment as a distributional criterion for world models and introduces PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics, and introduces PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions...

Yuandong Pu, Le Zhuo, Sayak Paul et al. · 0 citations
Jul 2026

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Evaluating representative proprietary and open-source multimodal models, it is found that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks.

Siyu Yan, Zhuoran Yan, Haiying Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.