4DHumanDiff is presented, a diffusion framework that directly generates dynamic humans represented by 4D Gaussian Splatting from text prompts and achieves better temporal and multi-view consistency, and reduces inference time by more than 10x.
Abstract
Generating high-quality 360-degree dynamic human assets from text prompts is challenging. Existing methods usually synthesize monocular or multi-view videos first and then fit a 4D representation, which is expensive and often causes incomplete geometry or view-inconsistent renderings. We present 4DHumanDiff, a diffusion framework that directly generates dynamic humans represented by 4D Gaussian Splatting (4DGS) from text prompts. By modeling the structured 4D representation space end-to-end, 4DHumanDiff avoids video pre-generation and per-scene reconstruction, making it better suited for view-consistent and temporally coherent asset generation. The model uses a 3D U-Net backbone with temporal attention for motion-aware generation. We further construct a large-scale text-to-4DGS dataset with 60,000 high-quality pairs, and introduce 2D regularization and training-free 4D interpolation to improve rendering quality and motion smoothness. Experiments show that 4DHumanDiff generates consistent 360-degree dynamic humans within one minute, achieves better temporal and multi-view consistency, and reduces inference time by more than 10x.
TGRHuman, a novel approach for generating realistic 3D humans from text that decouples geometry and texture generation to alleviate the issues commonly encountered in NeRF-based methods, outperforms existing text-to-3D human methods in both geometry and texture quality.
Muxin Zhang, Chao-Hui Yu, Yuanwang Yang et al.· Fundamental Research· 0 citations
Block3D is proposed, a block-wise diffusion framework that partitions the discrete shape-token sequence into contiguous blocks, generates the blocks autoregressively, and jointly denoises all tokens within the current block and introduces confidence-guided intra-block correction, which revises low-confidence tokens bef...
Bowen Cui, Weijie Wang, Zeyu Zhang et al.· 1 citation
DiGS-Avatar is proposed, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design, and introduces a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion stude...
This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequ...
Dynamic visual data, such as monocular videos of moving humans and changing environments, contain rich spatio-temporal information but are often large, redundant, and difficult to represent in a compact and structured form. 4D Gaussian Splatting has recently emerged as an explicit representation for dynamic scene recon...
Junita Sari, Nicolas Slat, Di Jiang et al.· 2026 12th International Conf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.