This paper proposes K-NeAS, a unified and scalable architecture for automated, multi-material surface reconstruction that replaces independent material networks with a shared latent backbone and introduces a fully differentiable $K$-material sequential soft selector to model an arbitrary number of overlapping tissues.
Abstract
Computed Tomography (CT) carries significant ionizing radiation risks, driving the need for sparse-view reconstruction. Implicit scene representations (ISRs) address this by recovering continuous volumetric attenuation fields directly from sparse projections, and recent geometry-aware extensions jointly model surface geometry alongside attenuation to improve fidelity and enable clean tissue segmentation without manual thresholding. However, these methods remain limited by manually tuned attenuation bounds and rigid two-material constraints. This paper proposes $K$-NeAS, a unified and scalable architecture for automated, multi-material surface reconstruction. We replace independent material networks with a shared latent backbone and introduce a fully differentiable $K$-material sequential soft selector to model an arbitrary number of overlapping tissues. To eliminate manual tuning, we automate attenuation bounding using a Gaussian Mixture Model (GMM) and implement a scheduled auxiliary floater loss to mitigate geometric hallucinations common under extreme sparsity. Evaluated across four clinical Cone-Beam CT (CBCT) datasets, $K$-NeAS successfully scales to arbitrary material counts, achieving superior 3D volumetric fidelity at $K=3$ materials on complex multi-tissue regions such as the Abdomen ($33.28\text{ dB}$ 3D PSNR vs. $31.40\text{ dB}$ single-material NeAS baseline, a $+1.88\text{ dB}$ improvement). Furthermore, our model exhibits enhanced robustness under sparse-sampling conditions, outperforming baseline 3D PSNR by up to $1.17\text{ dB}$ under 5- and 10-view constraints.
HiGDiff is proposed, a feed-forward hierarchical Gaussian diffusion framework that decomposes reconstruction both spatially and from structure to detail in three distinct CT benchmark datasets.
This work designs an implicit pixel-wise learnable step size to adapt to the spatial gradient heterogeneity of CT images and develops a cross-prompt guiding mechanism to enable inter-domain prompt interaction, which facilitates efficient prompt generation and enhances the convergence stability of the model.
Wenchao Du, Qiao Mu, Huanhuan Cui et al.· IEEE Transactions on Medical...· 0 citations
Computed tomography (CT) throughput is limited by scan time, which grows with both the number of projections acquired and the detector integration time for each. Reconstructing high-quality volumes from sparse-view or low-dose measurements therefore depends on an informative prior, typically a neural network trained for one specific scan setting and retrained whenever the modality, geometry, or material changes. We investigate whether a single diffusion model trained across several imaging domains can instead serve as a prior for many CT problems simultaneously. We evaluate the proposed method using the same frozen model on three datasets that differ in modality, beam geometry, material, and degradation type, spanning flaw analysis in additively manufactured metal parts imaged with cone-beam X-ray CT and concrete microstructure imaged with parallel-beam neutron CT. Our proposed method out-performs analytic reconstructions in all three cases, providing a step toward a reusable foundation prior for heterogeneous CT reconstruction problems.
Haley Duba-Sullivan, Patxi Fernandez-Zelaia, Obaidullah Rahman et al.· 0 citations
Sparse-view computed tomography (CT) reduces radiation dose by acquiring fewer projection views, but the resulting inverse problem is highly ill-posed and often produces severe streak artifacts. Existing deep reconstruction methods have achieved promising performance, yet many rely on first-order updates or large regularization networks, which can be less effective in ill-conditioned settings. We propose \textbf{CG-GLORE}, a compact deep unrolling framework inspired by second-order optimization for sparse-view CT reconstruction. Each unrolled stage uses a CG-solved linear system based on a structured Hessian surrogate: it retains the physics-induced curvature of the data-fidelity term while using an identity approximation for the learned regularization term. Thus, the method is second-order-inspired rather than an exact Newton method for the full learned objective. To model image priors, we design a Global-Local Regularization Network (GLORE), which combines convolutional local feature extraction with a Long-Range Dependency Representation module based on sparse patchification and Nystr\"{o}m attention. This design captures anatomical details and non-local dependencies while maintaining practical complexity. Experiments on AAPM and DeepLesion under multiple sparse-view and noise settings show that CG-GLORE achieves strong quantitative performance, stable convergence, lower noise power, and improved visual fidelity compared with representative reconstruction methods.
Tran Xuan Hieu Le, D. C. Bui, V. Le et al.· 0 citations
Sparse-view computed tomography (CT) reconstruction aims to recover high-quality CT volumes from a limited number of X-ray projection images, thereby reducing radiation exposure during image acquisition. However, this problem is inherently ill-posed because each projection provides only indirect line-integral supervision, and different attenuation distributions can explain similar sparse measurements. Existing analytic and iterative methods often suffer from streak artifacts and unstable solutions, while supervised learning-based methods require paired training data and may generalize poorly across anatomical regions or acquisition settings. Neural Radiance Field (NeRF)-based methods have recently shown promise by representing the attenuation field as a continuous coordinate-based function optimized directly from projection images. Nevertheless, these methods mainly enforce projection consistency and do not explicitly use volume-domain uncertainty to guide subsequent reconstruction. In this work, we propose EpiC-NeRF, a CT-specific closed-loop framework that actively feeds estimated epistemic uncertainty back into sparse-view reconstruction. EpiC-NeRF adapts evidential uncertainty estimation and aggregation to the X-ray CT line-integral formulation and maintains the resulting spatial uncertainty in a persistent three-dimensional Epistemic Grid Map. The accumulated uncertainty is used by Epistemic-Adaptive Layer Normalization to modulate intermediate features and by dual active sampling to guide ray- and point-level sample allocation. The newly estimated uncertainty then updates the grid map and guides subsequent optimization iterations, forming a unified feedback loop between uncertainty estimation and CT reconstruction. Experiments on four CT volume datasets demonstrate that EpiC-NeRF achieves improved reconstruction fidelity over existing analytic, iterative, and neural implicit reconstruction methods.
Donghyuk Choo, Haill An, Younhyun Jung· Mathematics· 0 citations