Skip to content
Open access

Text-to-hierarchical three-dimensional scene generation: a new approach for layered three-dimensional modeling from natural language

Aug 2026 · Visual Computing for Industry, Biomedicine, and Art · Vol 9 · 0 citations · 48 references
Medicine

TL;DR

An end-to-end hierarchical framework for text-to-3D scene generation that synergistically integrates state-of-the-art components for video synthesis and mesh reconstruction is introduced, offering a powerful solution for applications, such as virtual reality and digital twins.

Abstract

Rapid advancements in generative models have enabled substantial progress in text-to-3D scene synthesis. However, existing approaches often lack hierarchical structure and geometric consistency, limiting their use in complex, editable scene creation. To address this issue, an end-to-end hierarchical framework for text-to-3D scene generation is introduced. The main contribution of this study lies not in advancing a single generative model but in the design of a framework that synergistically integrates state-of-the-art components for video synthesis and mesh reconstruction. By decoupling scene semantics from dynamic camera trajectory control, the proposed system efficiently produces structured object-level editable three-dimensional scenes from natural language. Comprehensive experiments showed that the proposed integrated approach achieved superior scene quality, consistency, and editing flexibility compared with existing methods, offering a powerful solution for applications, such as virtual reality and digital twins.

Read PDF

Similar papers

Conference Open access Sep 2026

A Comprehensive Survey of Interaction Techniques in 3D Scene Generation

3D scene generation has rapidly evolved, significantly promoting the innovation of content creation. In this context, interaction techniques serve as a pivotal bridge connecting user intent with the generative models, thereby enabling precise control, real-time feedback and personalized customization of complex 3D scen...

Yuqi Li, Si-Wei Meng, Chuan-Guang Yang et al. · 33 citations · ⚡1
Conference Open access 2025

Enhanced Text-to-Image Editing with Multi-Step Control and Explainability

This approach develops an improved text-to-image editing system that enables users to apply sequential edits while preserving previous alterations, with an added option to undo edits when necessary.

S. R, S. K., S. Harish et al. · 0 citations
#diffusion models Conference Open access Sep 2026

AI-Driven 3D/4D Content Generation, Editing, and Animation Using Diffusion Models and Gaussian Splatting: A Narrative Review

This narrative review syn thesizes representative works published between 2023 and 2025 across four interconnected sub-areas: text-to-3D/4D generation, 3D scene editing, dynamic scene animation, and motion and video generation.

Yan-Ni Liu · 0 citations
Preprint Aug 2026

Bridging Language and Spherical Space: Object-Centric Control for Text-to-Panorama Generation

Panoramic image generation is increasingly important for immersive applications such as virtual reality, augmented reality, and 3D content creation. Unlike perspective images, panoramic images represent a viewer-centered $360^\circ$ surrounding space, where directional expressions such as left, right, front, and behind...

De-Rui Li, Qian Qiao, Yuhao Sun et al. · 0 citations
Review Open access Aug 2026

Research on an AI Visual Effects Generation Framework for Virtual Digital Humans in Film and Television Production

A modular AI visual effects generation framework tailored to film and television production pipelines that achieves collaborative geometry-texture modeling through joint optimization of NeRF and GAN, and resolves stage-wise representation inconsistency via cross-module feature alignment.

P. Yan, Y.-B. Ma, H.-L. Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.