Aug 2026· Journal of X-Ray Science and Technology· pp.
8953996261476606
· 0 citations· 21 references
Medicine
TL;DR
The experimental results demonstrate that MHSG-Model has a good performance on preserving detail information and removing artifacts, which show the potential to be applied in clinical sparse view spectral CT reconstruction.
Abstract
ObjectiveSparse View Computed Tomography (SVCT) is an effective way to reduce radiation dose. However, missing information of projection leads to noise and artifacts. Current SVCT reconstruction networks cannot effectively utilize the correlation information of images from different energies, and linear interpolation in the projection domain brings more secondary artifacts. A dual-domain approach involving data complementation and correction is proposed as a potential solution in this work.ApproachWe introduce the MHSG-Model for sparse-view spectral CT, designed for integrated processing in both the projection and image domains. This model consists of a Projection Domain Generative Model (PDGM) and an Image Domain Correction Module (IDCM). The PDGM employs a generative model integrated with a selective State Space Model (SSM) to recover missing projection data, reducing secondary artifacts and computational complexity. The IDCM utilizes a novel cross-attention mechanism, incorporating a multi-scale attention for high-energy images and a high-frequency attention for low-energy images, to further enhance reconstruction quality.Main resultsThe MHSG-Model was launched and validated on an abdominal slice simulation dataset and a sparse view spectral CT dataset of the AAPM DL Challenge. The proposed network achieved superior performance metrics, with an RMSE of 0.0111, MAE of 15.91, SSIM of 0.9674, and PSNR of 39.73 dB. The experimental results demonstrate that MHSG-Model has a good performance on preserving detail information and removing artifacts, which show the potential to be applied in clinical sparse view spectral CT reconstruction.
Introduction Sparse-view computed tomography (CT) reconstruction is crucial for clinical diagnostics, as reducing radiation exposure is essential to minimize risks to patients. Existing dual-domain reconstruction methods leverage both image and projection domains but often process them sequentially, overlooking their implicit correlations. Methods To address this limitation, we propose Cross-Domain TransNet, a Transformer-based dual-domain framework for sparse-view CT reconstruction. The proposed model captures long-range dependencies within each domain and integrates image and sinogram representations through a hybrid self-attention mechanism. In addition, a Convolution Fusion Layer (CFL) is introduced to enhance feature interactions and facilitate more effective utilization of dual-domain information. Results Extensive experiments on the NIH-AAPM dataset demonstrate the superior performance and generalization capability of the proposed method under various sparse-view settings. The results show that Cross-Domain TransNet consistently improves reconstruction quality, effectively suppresses noise, and reduces artifacts, outperforming both conventional reconstruction algorithms and state-of-the-art deep learning approaches. Conclusion Cross-Domain TransNet provides an effective and robust solution for sparse-view CT reconstruction. By fully exploiting complementary information from both image and projection domains, the proposed framework enhances diagnostic image quality while supporting radiation dose reduction.
Junling Wang, Chunhua Zou, Hongjie Yang et al.· Frontiers in Nuclear Medicin...· 0 citations
This work designs an implicit pixel-wise learnable step size to adapt to the spatial gradient heterogeneity of CT images and develops a cross-prompt guiding mechanism to enable inter-domain prompt interaction, which facilitates efficient prompt generation and enhances the convergence stability of the model.
Wenchao Du, Qiao Mu, Huanhuan Cui et al.· IEEE Transactions on Medical...· 0 citations
In industrial computed tomography for defect detection in 18650 lithium-ion batteries, streak-like artifacts caused by sparse-view projection sampling severely hinder the accurate identification of subtle structural defects. This paper proposes a hybrid-domain CT reconstruction algorithm with projection inpainting based on a joint CNN-Transformer architecture. The proposed method integrates projection-domain restoration with image-domain optimization to construct an end-to-end dual-domain collaborative reconstruction framework. Specifically, a hybrid-domain consistent restoration module is designed to leverage filtered back-projection priors to guide the convolutional network in performing an initial completion of sparse-view projection data. In addition, a multi-level Transformer structure is introduced to model global correlations across projection views, thereby accurately correcting projection deviations caused by sparse sampling and noise. The overall framework enables an accurate mapping from undersampled sinograms to high-fidelity CT images. Simulation and experimental results demonstrate that the proposed method effectively suppresses artifacts under sparse-view projection sampling condition and outperforms present methods in key metrics such as RMSE, PSNR, and SSIM. In particular, it shows superior performance in edge preservation and structural recovery for pixel-level defects, highlighting its strong potential for high-performance industrial CT defect inspection.
Zihao Liu, Chenglong Wang, Zhengxin Li et al.· International Conference on...· 0 citations
Cone-beam computed tomography (CBCT) with sparse projection views offers reduced radiation dose and faster scans but introduces severe streak artifacts and spatial coverage gaps. We address these challenges within a unified framework. First, we replace conventional UNet/ResNet encoders with TransUNet, a hybrid CNN–Transformer architecture that jointly models local details and long-range spatial context. It is adapted to CBCT reconstruction by concatenating multi-scale feature maps and introducing a lightweight attenuation-prediction head. Trans-CBCT outperforms the best baseline by 1.17 dB in PSNR and by 0.0163 in SSIM on LUNA16 with only six projection views. Second, we incorporate a neighbor-aware Point Transformer with explicit 3D positional encodings and a neighbor-aware attention module aggregating information from each point’s k-nearest spatial neighbors to enforce volumetric coherence. The resulting Trans2-CBCT achieves an additional 0.63 dB increase in PSNR and 0.0117 increase in SSIM over Trans-CBCT. In experiments with 6-10 views, Trans-CBCT and Trans2-CBCT consistently outperform all prior methods in both PSNR and SSIM on LUNA16. On the ToothFairy dataset, Trans2-CBCT leads in five of the six measurements, outperforming all baselines in PSNR. These results highlight the effectiveness of combining hybrid CNN–Transformer features with geometry-aware point-based reasoning for sparse-view CBCT reconstruction.
Minmin Yang, Yunhui Zhu, Huantao Ren et al.· Italian National Conference...· 0 citations
Sparse-view CT reconstruction aims to synthesize volumetric CT images from a limited number of X-ray projections, reducing radiation dose while maintaining diagnostic quality. However, the substantial cross-modal gap between 2D X-rays and 3D CT volumes presents significant challenges, including uneven distribution of information across different views, artifact issues, and excessive computational costs. Therefore, in this paper, we propose MAC-DiffCT, a multi-scale adaptive conditional diffusion model designed for accurate and efficient CT reconstruction from biplanar X-rays. Our approach first extracts multi-scale 2D features from multi-view X-rays using a UNet-based encoder. A novel Bi-Directional Cross-Attention (BDC-Att) module adaptively fuses features by assigning spatially varying weights to each view. We then introduce a Multi-Scale Feature Sampling (MS-FS) module that projects 3D coordinates onto 2D planes, sample features across scales, and integrates them via a multi-layer perceptron to form a latent structural representation. This 3D structural feature serves as a condition for a latent-space conditional diffusion model, which reconstructs high-quality CT volumes with enhanced anatomical fidelity. An additional signed distance function (SDF) loss is applied to promote structural consistency in the 3D space. Experimental results on both public and private chest datasets demonstrate that MAC-DiffCT consistently outperforms existing methods, achieving the highest reconstruction accuracy with PSNR of 27.68 and 25.43 dB, SSIM of 0.8223 and 0.7598, and the lowest LPIPS of 0.0992 and 0.1322. Downstream evaluation via lung segmentation and an interpretability study further highlight the transparent, explainable, and anatomically grounded nature of our model. MAC-DiffCT offers a low-radiation alternative to conventional CT, especially for vulnerable patients requiring repeated imaging and intraoperative scenarios where CT use is constrained.
Chang Li, Jiao Meng, John Moraros et al.· IEEE journal of biomedical a...· 0 citations