We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation. The StepAudio Music Tokenizer represents audio as a 50-Hz stream from a 65536-entry single codebook, using semantically informed self-supervised and multi-t...
Chengli Feng, Zhi-Yue Wu, Jia-Hao Song et al.· 0 citations
Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradig...
This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence.
Zhiqin Yang, Jing-Wen Fu, Yu-Han Liu et al.· 1 citation
This work proposes ReBind, a systematic framework that introduces semantic instructions with embedded reference tokens as the intermediate representation for multi-reference image-conditioned video editing and develops ReBind-Instruct, a specialized MLLM that learns to establish explicit bindings between visual attribu...
Xin-Yu Liu, Shi-Hao Li, Weihong Lin et al.· arXiv.org· 2 citations
Experimental results indicate that fully automated data curation combined with imbalance-aware training yields non-trivial improvements, but models still struggle to capture fine-grained acoustic features, indicating a gap between surface-level alignment and deep musical comprehension.
Ziya Zhou, Shangda Wu, Shenyang Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.