Jul 2026· 2026 IEEE 27th China Conference on System Simulation Technology and its Applications (CCSSTA)· pp. 446-451· 0 citations· 28 references
Abstract
To address the problems of insufficient global structural consistency and local texture blurring in novel view synthesis under single-view conditions, this paper proposes a three-dimensional novel view synthesis method based on the fusion of a Vision Transformer and dual residual branches. The proposed method employs a Vision Transformer (ViT) to extract global features and capture long-range dependencies through a self-attention mechanism. Meanwhile, two complementary local branches are constructed. The RESFB module is designed based on Fast Fourier Convolution to fuse spatial-domain and frequency-domain information, while the RESTiedSE module introduces a TiedSE attention mechanism into the Res2Net framework to adaptively enhance key channel responses. The global and local features are fused in a multi-scale manner and combined with the NeRF volume rendering paradigm to generate novel views. Experiments conducted on the SRN-Chair and SRN-Car subsets of the ShapeNet dataset demonstrate the effectiveness of the proposed method. The results show that the proposed model achieves state-of-the-art SSIM and competitive PSNR, and effectively improves the clarity and structural fidelity of single-view novel view synthesis.
A dense matching network based on a Transformer and multi-scale feature fusion, called Task-aware Multi-Scale Matching Network (TMSMNet) is proposed, which outperforms mainstream methods such as RAFT-Stereo on the D1-all metric of KITTI- 2015 and demonstrates good generalization and robustness.
Shi-Xiong Liu· ITM Web of Conferences· 0 citations
A dual representation-based LFVS method that employs deformable convolutional and Deep Residual Channel Attention (DRCA) networks that achieves state-of-the-art performance on synthetic and real-world LF benchmarks.
Muhammad Zubair, Paulo J. L. Nunes, Caroline Conti et al.· IEEE Open Journal of Signal...· 0 citations
A novel image inpainting framework based on a Multi-Scale Parallel Dense Connection Network (MSPDCN) with holistically nested edge detection first employed to extract structural priors and estimate edge information of missing regions, which provides guidance for subsequent reconstruction and alleviates boundary blurrin...
Jie Wang, Li-Yuan Zhang, Yi-Bo Deng et al.· 電腦學刊· 0 citations
It is demonstrated that integrating spatial and frequency-domain representations through a dual-branch Vision Transformer architecture enhances photo aesthetic assessment performance.
A transformer-based channel attention block improves the discriminative capability of fused features in low-texture regions and enhances global consistency in a lightweight deep stitching framework that integrates multi-scale feature fusion with attention-enhanced matching.
Aiming at the problems of large parameters and high computational complexity in deep learning-based image super-resolution networks, this paper proposes a lightweight super-resolution network that fuses multi-attention mechanism and Blueprint Separable Convolution (BSConv). BSConv is introduced to improve performance w...
Yi-Yan Huang, Lin Guo· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.