Aug 2026· Visual Computing for Industry, Biomedicine, and Art· Vol 9· 0 citations· 48 references
Medicine
TL;DR
An end-to-end hierarchical framework for text-to-3D scene generation that synergistically integrates state-of-the-art components for video synthesis and mesh reconstruction is introduced, offering a powerful solution for applications, such as virtual reality and digital twins.
Abstract
Rapid advancements in generative models have enabled substantial progress in text-to-3D scene synthesis. However, existing approaches often lack hierarchical structure and geometric consistency, limiting their use in complex, editable scene creation. To address this issue, an end-to-end hierarchical framework for text-to-3D scene generation is introduced. The main contribution of this study lies not in advancing a single generative model but in the design of a framework that synergistically integrates state-of-the-art components for video synthesis and mesh reconstruction. By decoupling scene semantics from dynamic camera trajectory control, the proposed system efficiently produces structured object-level editable three-dimensional scenes from natural language. Comprehensive experiments showed that the proposed integrated approach achieved superior scene quality, consistency, and editing flexibility compared with existing methods, offering a powerful solution for applications, such as virtual reality and digital twins.
The Hash-Atlas network is proposed, which reformulates 3D scene editing as operations on 2D atlas images, thereby achieving a workflow decoupling of the 2D editing and 3D reconstruction processes.
Shuangkang Fang, Yu-Feng Wang, Yi-Hsuan Tsai et al.· 1 citation
3D scene generation has rapidly evolved, significantly promoting the innovation of content creation. In this context, interaction techniques serve as a pivotal bridge connecting user intent with the generative models, thereby enabling precise control, real-time feedback and personalized customization of complex 3D scen...
Yuqi Li, Si-Wei Meng, Chuan-Guang Yang et al.· Proceedings of the Thirty-Fi...· 33 citations· ⚡1
This approach develops an improved text-to-image editing system that enables users to apply sequential edits while preserving previous alterations, with an added option to undo edits when necessary.
S. R, S. K., S. Harish et al.· Proceedings of the 1st Inter...· 0 citations
This narrative review syn thesizes representative works published between 2023 and 2025 across four interconnected sub-areas: text-to-3D/4D generation, 3D scene editing, dynamic scene animation, and motion and video generation.
Panoramic image generation is increasingly important for immersive applications such as virtual reality, augmented reality, and 3D content creation. Unlike perspective images, panoramic images represent a viewer-centered $360^\circ$ surrounding space, where directional expressions such as left, right, front, and behind...
De-Rui Li, Qian Qiao, Yuhao Sun et al.· 0 citations
A modular AI visual effects generation framework tailored to film and television production pipelines that achieves collaborative geometry-texture modeling through joint optimization of NeRF and GAN, and resolves stage-wise representation inconsistency via cross-module feature alignment.
P. Yan, Y.-B. Ma, H.-L. Li et al.· Advanced Electromagnetics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.