Skip to content

Author

Bo-Yang Wang

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

StepAudio 3 Music Technical Report

We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation. The StepAudio Music Tokenizer represents audio as a 50-Hz stream from a 65536-entry single codebook, using semantically informed self-supervised and multi-t...

Chengli Feng, Zhi-Yue Wu, Jia-Hao Song et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis

We present VGGT-Diff, a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. Existing novel view synthesis (NVS) methods face a fundamental trade-off: reconstruction-based approaches preserve observed geometry but struggle to synthesize unseen regions, while diffusion-based methods provide s...

Kang-Jie Chen, Xiang-Yu Li, Dong-Bin Zhang et al. · 0 citations
Preprint Sep 2026

StepAudio 3 Gen Technical Report

We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator tha...

Bin Lin, Bo Zhao, Bo-Yang Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.