A human-in-the-loop GenAI-assisted framework for producing immersive 3D visualization prototypes rather than historically verified reconstructions is proposed, which integrates multi-view image generation, knowledge-informed review, single-image-to-3D generation, topology inspection, and perceptual calibration.
Abstract
Immersive visualization can support interpretation of architectural heritage in historical paintings, yet translating 2D pictorial evidence into navigable 3D scenes remains challenging. Conventional workflows rely on physical survey data, while direct generative AI (GenAI) may produce structural hallucinations and lack historical constraints. This study proposes a human-in-the-loop GenAI-assisted framework for producing immersive 3D visualization prototypes rather than historically verified reconstructions. It integrates multi-view image generation, knowledge-informed review, single-image-to-3D generation, topology inspection, and perceptual calibration. Four fragments from the Northern Song Dynasty painting Along the River During the Qingming Festival were examined as a single-case proof of concept. Across three tested model pairs, raw AI assets were generated in approximately 3–4 min and were suitable for distant-background use; close-up visualization required 1–2 h of refinement, while basic structural editability required 4–5 h of post-processing, reducing the initial time advantage. A mixed-methods study with nine domain experts and 30 non-expert participants used the UES-SF, an adapted VisAWI, and semi-structured interviews analyzed through inductive thematic analysis. All eight subscale scores exceeded their neutral midpoints after Bonferroni correction (all adjusted p<0.001), indicating favorable perceptions of the guided experience. Interviews suggested potential for spatial exploration, museum interpretation, and education. However, geometric discontinuities, detail loss, color deviation, and historical-semantic errors remained, requiring expert review and manual correction. Transferability beyond this artwork and architectural tradition remains untested.
In independent game development, creating high-quality 3D assets remains a critical bottleneck, as traditional workflows require specialized skills that hinder non-expert designers. While Generative AI has democratized 2D concept art, translating these visions into game-ready 3D assets is technically demanding. To address this, we propose an automated pipeline that leverages semantic and geometric feature extraction to synthesize 3D meshes and PBR materials from AI-generated 2D images. Our system orchestrates a workflow bridging text-to-image diffusion models with advanced computer vision modules. It employs monocular depth estimation and intrinsic decomposition to interpret geometric structures and isolate material properties (albedo, roughness, metallic) from multi-view consistent concepts. These features are algorithmically processed to produce finalized .glb assets. A comparative user study with indie developers demonstrates that this pipeline reduces asset production time by approximately 80% compared to traditional modeling tools. By shifting the user’s role from manual vertex manipulation to high-level semantic curation, this research validates a novel workflow that empowers designers to rapidly populate immersive worlds, streamlining the future of interactive media design.
Jie Hu, Jinyu Li, Zixia Wang et al.· AHFE International· 0 citations
Scientific visualization is changing from passive observation to active, AI-assisted collaboration. While Extended Reality (XR) has proven valuable for comprehending dense 3D arrays, traditional VR applications are typically deployed in rigid, single-purpose, and monolithic architectures. In this paper, we present the evolution of ASCRIBE-XR: a virtual reality platform backed by remote computation that has been re-engineered into a dynamic, service-oriented ecosystem. We introduce three core innovations that make immersive data analysis easier, faster, and more flexible when using multimodal scientific imaging. First, a lightweight Python REST interface decouples XR logic from the rendering engine, enabling real-time, programmable scene customization and on-demand data generation. Second, we present a Specimen Catalog architecture that lets the platform pivot between radically different disciplines, ranging from archaeological heterogeneous concrete and fuel-cell membranes to the root system of a bioenergy grass, by describing each dataset through portable metadata rather than hard-coded application logic. Finally, we introduce a prompt-driven layer powered by the Claude Agent SDK, allowing researchers to generate, segment, and manipulate volumetric and mesh data through natural language dialogue within the virtual space. For example, applying foundation models such as the Segment Anything Model (SAM) to perform zero-shot segmentation on demand. By bridging human intent with remote computation, ASCRIBE-XR relaxes the constraints of conventional visualization tools, offering a highly adaptable, conversational platform for scientific discovery with human auditing.
Ronald Pandolfi, L. Weidner, J. Sethian et al.· Journal of Imaging· 0 citations
This research demonstrates how AI transcends mere visual generation to become a new pathway for the dynamic preservation and revitalization of cultural heritage.
Liwen Fan, Aojie Feng, Yuxuan Li et al.· Creativity & Cognition· 0 citations
Abstract. The Borobudur Temple’s Hidden Foot (Karmawibhangga layer), comprising 160 bas-relief panels sealed behind protective stone cladding since the early 20th century, presents a compelling challenge for digital cultural heritage practice. This paper presents a digital twin reconstruction and immersive virtual reality (VR) system that provides interactive access to this physically inaccessible archaeological space. The system integrates three pre-existing datasets: UAV photogrammetric point clouds of the outer temple structure, deep-learning-based 3D relief reconstructions derived from archival photographs (Pan et al., 2022), and semantic segmentation results of relief surfaces (Ji et al., 2023). Within a Unity 2022.3.22f1-based VR environment optimised for the HTC VIVE Pro Eye head-mounted display, these datasets are unified into a navigable Hidden Foot passage. A novel lighting-driven Level of Detail (LoD) strategy—where rendering complexity is governed by the illumination radius of a virtual torch—achieves 65–70 FPS on the HMD while preserving visual fidelity within the illuminated zone. Eye-gaze-triggered textual annotations support contextual iconographic interpretation, while interactive toggling between photorealistic and semantically labelled views enables both casual exploration and scholarly analysis. A user study (n = 10) employing a System Usability Scale-adapted questionnaire yielded mean satisfaction ratings exceeding 4.5/5 across all evaluation dimensions. The system demonstrates a scalable methodology for transforming inaccessible heritage layers into immersive, annotated digital environments.
Satoshi Takatori, Jiajun Gao, Liang Li et al.· The International Archives o...· 0 citations
This study proposes a four-stage closed-loop model—semantic analysis, form mapping, rendering, and interaction evaluation—enhanced by a dual-loop mechanism combining cultural feedback and experience optimization. Three national-level ICH elements—Nuo opera masks, Qiang flute tunes, and lacquerware patterns—are explored. A 15,000-frame action library and 3D scanning data are created, achieving frame-by-frame enhancement via Deformable Convolution, reducing error to 0.42%. A blind test with 340 participants shows increases in immersion (0.64), cultural resonance (0.51), and narrative consistency (13.7%) over the industry baseline. Regression analysis confirms a strong link between semantic depth and audience immersion. The results show that 3D animation-driven reconstruction enhances interaction and provides a measurable framework for ICH expression, offering insights for cultural inheritance and digital industry development.
Yunfeng Hu, Hao Liu· International Journal of Agr...· 0 citations
Generative AI lets anyone create rich visual content in seconds, yet translating that content into a physically fabricable artifact still demands manual decomposition, occlusion repair, and structural verification that most tools leave entirely to the user. We present FabDreamer, an image-to-physical system that carries an image to fabrication-ready SVGs through three stages with deliberately staged AI initiative: (1) AI leads decomposition into depth-ordered layers, (2) assists on demand during creative editing with realtime 3D preview, and (3) advises on structural integrity before export. We instantiate this workflow for layered laser-cut art and evaluate it through three rounds including a formative analysis, an early prototype user evaluation (N=13), and a cross-domain practitioner study with specialists from 6 fabrication domains (N=6). Our findings show that physical awareness during design opens creative opportunities beyond error prevention, that practitioners appropriate the system's generic geometric operations for their own domains, and that the fabrication agent covers geometry-readable constraints while domain knowledge remains with the maker.
Chenfeng Gao, Zeya Chen, Anjie Yang et al.· 0 citations