Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent world state and enforcing programmable rules over extended interactions. We introduce Programmable World Model, a framework that decouples world-state evolution from visu...
Zheng-Hui Huang, Gui-Xu Lin, Jia-Cheng Lin et al.· 0 citations
Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redundant compute, view inconsistencies, and poor scalability. We propose MASS (Multiplayer world models with Authoritative Shared State) to resolve this limitation. Inspired...
Ziqi Cai, Si-Qi Yang, Yimu Wang et al.· 0 citations
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics....
Mu-Yao Niu, Ji-Xuan He, Rui-Han Yu et al.· 0 citations
Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that generate frames from text/control-hints prompts, AlayaRenderer preserves scene structure without altering the underlying world dynamics. This demonstrates an alternative...
Guixu Lin, Zheng-Hui Huang, Siqi Yang et al.· arXiv.org· 0 citations
Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligence. Recent video generative models provide a data-driven route toward this goal by predicting future observations conditioned on user actions, and are increasingly rega...
Zhen Li, Zian Meng, Shuwei Shi et al.· arXiv.org· 2 citations
This work proposes Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process to support RL post-training.
P. Zhou, Hesong Wang, Zhengfeiyang Zhang et al.· 0 citations
OrthoPilot is a clinical artificial intelligence system powered by a large language model (LLM) that integrates hospital data streams with authoritative external knowledge for continuous musculoskeletal care and autonomously retrieves real-time imaging, laboratory, pathology and order data and translates evolving patie...
AlayaWorld enables open-ended real-time interaction, allowing users to freely navigate and perform diverse actions such as combat, spell casting, and monster summoning, and the framework unifies the complete development-from data preparation model architecture, model training, inference acceleration, and deployment-wit...
AlayaWorld Team, Kaipeng Zhang, Chuanhao Li et al.· arXiv.org· 1 citation
WorldRover turns long-horizon world exploration into a scalable data-generation problem, providing supervision for models that must build, maintain, and revisit coherent representations of an explorable world.
This work introduces Surprise Forcing, a training-free framework that treats both limitations as online resource-allocation problems and improves long-horizon consistency and visual quality while retaining real-time streaming throughput.
Sekai2 is introduced, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling, and Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant tem...
Kang He, Wen-Shuo Peng, Zi-Hui Gao et al.· 2 citations· ⚡1
GROVE is introduced, a training-free framework that supports both behaviors with one memory grown causally from a continuous video stream, and achieves the best results among the compared methods.
Sitong Gong, Caixin Kang, Tianyu Yan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.