This paper proposes K-NeAS, a unified and scalable architecture for automated, multi-material surface reconstruction that replaces independent material networks with a shared latent backbone and introduces a fully differentiable $K$-material sequential soft selector to model an arbitrary number of overlapping tissues.
Abstract
Computed Tomography (CT) carries significant ionizing radiation risks, driving the need for sparse-view reconstruction. Implicit scene representations (ISRs) address this by recovering continuous volumetric attenuation fields directly from sparse projections, and recent geometry-aware extensions jointly model surface geometry alongside attenuation to improve fidelity and enable clean tissue segmentation without manual thresholding. However, these methods remain limited by manually tuned attenuation bounds and rigid two-material constraints. This paper proposes $K$-NeAS, a unified and scalable architecture for automated, multi-material surface reconstruction. We replace independent material networks with a shared latent backbone and introduce a fully differentiable $K$-material sequential soft selector to model an arbitrary number of overlapping tissues. To eliminate manual tuning, we automate attenuation bounding using a Gaussian Mixture Model (GMM) and implement a scheduled auxiliary floater loss to mitigate geometric hallucinations common under extreme sparsity. Evaluated across four clinical Cone-Beam CT (CBCT) datasets, $K$-NeAS successfully scales to arbitrary material counts, achieving superior 3D volumetric fidelity at $K=3$ materials on complex multi-tissue regions such as the Abdomen ($33.28\text{ dB}$ 3D PSNR vs. $31.40\text{ dB}$ single-material NeAS baseline, a $+1.88\text{ dB}$ improvement). Furthermore, our model exhibits enhanced robustness under sparse-sampling conditions, outperforming baseline 3D PSNR by up to $1.17\text{ dB}$ under 5- and 10-view constraints.
HiGDiff is proposed, a feed-forward hierarchical Gaussian diffusion framework that decomposes reconstruction both spatially and from structure to detail in three distinct CT benchmark datasets.
This work designs an implicit pixel-wise learnable step size to adapt to the spatial gradient heterogeneity of CT images and develops a cross-prompt guiding mechanism to enable inter-domain prompt interaction, which facilitates efficient prompt generation and enhances the convergence stability of the model.
Wenchao Du, Qiao Mu, Huanhuan Cui et al.· IEEE Transactions on Medical...· 0 citations
Computed tomography (CT) throughput is limited by scan time, which grows with both the number of projections acquired and the detector integration time for each. Reconstructing high-quality volumes from sparse-view or low-dose measurements therefore depends on an informative prior, typically a neural network trained for one specific scan setting and retrained whenever the modality, geometry, or material changes. We investigate whether a single diffusion model trained across several imaging domains can instead serve as a prior for many CT problems simultaneously. We evaluate the proposed method using the same frozen model on three datasets that differ in modality, beam geometry, material, and degradation type, spanning flaw analysis in additively manufactured metal parts imaged with cone-beam X-ray CT and concrete microstructure imaged with parallel-beam neutron CT. Our proposed method out-performs analytic reconstructions in all three cases, providing a step toward a reusable foundation prior for heterogeneous CT reconstruction problems.
Haley Duba-Sullivan, Patxi Fernandez-Zelaia, Obaidullah Rahman et al.· 0 citations
Experiments on AAPM and DeepLesion under multiple sparse-view and noise settings show that CG-GLORE achieves strong quantitative performance, stable convergence, lower noise power, and improved visual fidelity compared with representative reconstruction methods.
Tran Xuan Hieu Le, D. C. Bui, V. Le et al.· 0 citations
Sparse-view CT reconstruction aims to synthesize volumetric CT images from a limited number of X-ray projections, reducing radiation dose while maintaining diagnostic quality. However, the substantial cross-modal gap between 2D X-rays and 3D CT volumes presents significant challenges, including uneven distribution of information across different views, artifact issues, and excessive computational costs. Therefore, in this paper, we propose MAC-DiffCT, a multi-scale adaptive conditional diffusion model designed for accurate and efficient CT reconstruction from biplanar X-rays. Our approach first extracts multi-scale 2D features from multi-view X-rays using a UNet-based encoder. A novel Bi-Directional Cross-Attention (BDC-Att) module adaptively fuses features by assigning spatially varying weights to each view. We then introduce a Multi-Scale Feature Sampling (MS-FS) module that projects 3D coordinates onto 2D planes, sample features across scales, and integrates them via a multi-layer perceptron to form a latent structural representation. This 3D structural feature serves as a condition for a latent-space conditional diffusion model, which reconstructs high-quality CT volumes with enhanced anatomical fidelity. An additional signed distance function (SDF) loss is applied to promote structural consistency in the 3D space. Experimental results on both public and private chest datasets demonstrate that MAC-DiffCT consistently outperforms existing methods, achieving the highest reconstruction accuracy with PSNR of 27.68 and 25.43 dB, SSIM of 0.8223 and 0.7598, and the lowest LPIPS of 0.0992 and 0.1322. Downstream evaluation via lung segmentation and an interpretability study further highlight the transparent, explainable, and anatomically grounded nature of our model. MAC-DiffCT offers a low-radiation alternative to conventional CT, especially for vulnerable patients requiring repeated imaging and intraoperative scenarios where CT use is constrained.
Chang Li, Jiao Meng, John Moraros et al.· IEEE journal of biomedical a...· 0 citations
Sparse-view computed tomography (CT) reconstruction aims to recover high-quality CT volumes from a limited number of X-ray projection images, thereby reducing radiation exposure during image acquisition. However, this problem is inherently ill-posed because each projection provides only indirect line-integral supervision, and different attenuation distributions can explain similar sparse measurements. Existing analytic and iterative methods often suffer from streak artifacts and unstable solutions, while supervised learning-based methods require paired training data and may generalize poorly across anatomical regions or acquisition settings. Neural Radiance Field (NeRF)-based methods have recently shown promise by representing the attenuation field as a continuous coordinate-based function optimized directly from projection images. Nevertheless, these methods mainly enforce projection consistency and do not explicitly use volume-domain uncertainty to guide subsequent reconstruction. In this work, we propose EpiC-NeRF, a CT-specific closed-loop framework that actively feeds estimated epistemic uncertainty back into sparse-view reconstruction. EpiC-NeRF adapts evidential uncertainty estimation and aggregation to the X-ray CT line-integral formulation and maintains the resulting spatial uncertainty in a persistent three-dimensional Epistemic Grid Map. The accumulated uncertainty is used by Epistemic-Adaptive Layer Normalization to modulate intermediate features and by dual active sampling to guide ray- and point-level sample allocation. The newly estimated uncertainty then updates the grid map and guides subsequent optimization iterations, forming a unified feedback loop between uncertainty estimation and CT reconstruction. Experiments on four CT volume datasets demonstrate that EpiC-NeRF achieves improved reconstruction fidelity over existing analytic, iterative, and neural implicit reconstruction methods.
Donghyuk Choo, Haill An, Younhyun Jung· Mathematics· 0 citations