Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams, and curate and structurally annotate 4.3 million physics images, carry expert-level annotation, and adapt a unified multimodal backbone.
Ming-Hui Zhang, Jinxin Shi, Yi-Fan Chang et al.· 0 citations
WorldRover turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.
Specialize-and-Merge Online Policy Distillation (SMOPD) is proposed, a two-stage training method for multi-reward optimization that outperforms GDPO across 1.5B, 3B and 7B backbones.
Wen Wang, Jia-Hua Bao, Tu Yongsiqi et al.· 1 citation
NapMem is introduced, a framework for learning to use long-term user memory as a structured action space rather than passively retrieved context, and suggests that long-term user memory benefits from coupling structured storage with a learned policy for using memory at the appropriate granularity.
This work introduces Skill Self-Play (Skill-SP), a co-evolutionary framework comprising a proposer, a solver, and a dynamic skill controller that effectively bridges the gap between structured verification and open-ended exploration.
Siyuan Huang, Pengyu Cheng, Haotian Liu et al.· arXiv.org· 2 citations
This work formalizes probabilistic alignment as a distributional criterion for world models and introduces PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics, and introduces PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions...
Yuandong Pu, Le Zhuo, Sayak Paul et al.· 0 citations
Evaluating representative proprietary and open-source multimodal models, it is found that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks.