Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only the representation across seven tokenizat...
Junhao Chen, Ming-Jin Chen, Jingjia Mao et al.· 1 citation
The new version of AlayaWorld substantially revise how conditioning signals are represented and integrated into the model, replacing the previous depth-warping-based spatial memory with a streaming 3D point-cache renderer.
AlayaWorld Team Kaipeng Zhang, Chuanhao Li, Y. Zhan et al.· 1 citation· ⚡1
AlayaWorld enables open-ended real-time interaction, allowing users to freely navigate and perform diverse actions such as combat, spell casting, and monster summoning, and the framework unifies the complete development-from data preparation model architecture, model training, inference acceleration, and deployment-wit...
AlayaWorld Team, Kaipeng Zhang, Chuanhao Li et al.· arXiv.org· 1 citation
D, a reference-guided renderer that extends Wan2.2 camera control from Plucker rays alone to a joint camera-plus-geometry interface and projects a neural 4D G-buffer from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior, supporting tracking+world-po...
Junhao Chen, Mingjin Chen, Henghaofan Zhang et al.· 0 citations
AlayaWorld is presented, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p and introduces a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 s...
AlayaWorld Team Kaipeng Zhang, Chuanhao Li, Y. Zhan et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.