Skip to content
Conference

A Multi-Modal Context-Aware Algorithm for Immersive Japanese Translation Screen via LiDAR-Driven Spatial Reconstruction and Visual Semantic Rendering

Aug 2026 · International Conference Computational Vision and Bio Inspired Computing · pp. 74-80 · 0 citations · 24 references

Abstract

To address the issues of spatial positioning drift, unstable semantic overlay, and insufficient real-time rendering efficiency in complex dynamic scenes of immersive Japanese translation screens in augmented reality environments, which limit the development potential of related applications, this paper proposes a multimodal context-aware algorithm (LMCAT) based on LiDAR-driven spatial reconstruction and visual semantic rendering. This method constructs a three-layer unified computational framework that collaboratively optimizes geometric space, visual semantics, and communication links. In the spatial perception layer, a continuous spatial potential field modeling mechanism is introduced. This mechanism achieves continuous reconstruction of sparse point clouds through density-adaptive topological radius and spherical shell topological convolution operators, establishing a topologically consistent 3D spatial physical field. In the visual semantic layer, a spatial-semantic dual-gated attention overlay model (SSDG-AoA) is designed. This paradigm utilizes visibility gating and semantic relevance gating to jointly optimize the spatial projection position, transparency distribution, and edge consistency of subtitles. This enables dynamic fusion of translated content and the real scene. In the communication layer, an exponentially weighted dynamic key synchronization mechanism (EW-DKS) is proposed. This mechanism embeds spatial reconstruction features into the encryption parameter generation process and improves the anti-interference capability during remote translation data transmission. The experimental platform was built based on 52,400 frames of data from three scenarios: Office-AR-LiDAR, Urban-Motion-LiDAR, and Crowd-Occlusion Set. It was also compared with methods such as VIO+Plane Overlay, GSNeRF-like Rendering, NeRF-Semantic Overlay, and SparsePoint Transformer Rendering. Experimental results show that in spatial reconstruction tasks, the proposed method achieves RMSEs of 1.82 cm, 3.07 cm, and 4.21 cm in Office-AR, Urban-Motion, and Crowd-Occlusion scenarios, respectively. Even with 40% point cloud loss, the reconstruction error is only 3.56 cm, demonstrating strong robustness. In visual semantic rendering tasks, the PSNR reaches 29.07 dB, SSIM reaches 0.90, and LPIPS is reduced to 0.19. Ablation experiments show that the complete SSDG-AoA structure reduces the caption misalignment rate from 12.3 % to 3.8 %. At the same time, the system's average end-to-end latency is 41.2 ms, and the P95 latency is 63.5 ms.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.