Polyp segmentation in colonoscopy images plays a pivotal role in computer-aided medical diagnosis and the early prevention of colorectal cancer. However, existing methods often suffer from performance degradation when confronted with extreme polyp scale variation and polyp boundary ambiguity. To address these challenges, we propose the Staged Global-to-Local Cross-Scale Fusion Network (SGLF-Net), which adopts a novel staged global-to-local learning paradigm to progressively refine segmentation from coarse global semantics to fine-grained local details. Specifically, the Global Semantic Perception Stage integrates a Swin Transformer Encoder and a Dynamic Attentive Decoder (DAD) to construct comprehensive multi-scale contextual representations. The Local Detail Refinement Stage employs an Edge-aware Dynamic Attentive Decoder (E-DAD) to enhance structural fidelity and boundary precision through explicit edge-guided supervision. Furthermore, we introduce the Cross Spatial-Scale Feature Aggregation and Reconstitution (CSSAR) module, equipped with hybrid attention mechanisms, to facilitate efficient semantic structural interaction between the two cascaded stages. Extensive experiments on five public benchmark datasets demonstrate that SGLF-Net consistently outperforms state-of-the-art methods in both segmentation accuracy and boundary preservation.
Tan Guo, Wen-Han Zhang, Fu-Lin Luo et al.· IEEE journal of biomedical a...· 0 citations
Low-dose computed tomography (LDCT) measurements contain mixed Poisson-Gaussian noise. However, most self-supervised methods rely on generic image statistics and do not explicitly model this noise, which may limit their ability to effectively suppress realistic LDCT noise. To address this issue, we propose a physics-driven framework with cross-domain iteration for self-supervised LDCT denoising. The proposed framework proceeds in three main steps. First, a learned sinogram prior and the LDCT noise model guide posterior inference of photon counts, enabling separation of the Poisson and Gaussian components. Second, the separated Poisson and Gaussian components are respectively processed by binomial thinning and Gaussian data thinning to construct two branches, and residual scaling matches each branch's noise level to that of the observation, yielding a training pair with approximately independent noise realizations from one low-dose measurement. Finally, the pair is used to train an image-domain network whose forward-projected outputs update the prior. Through cross-domain iteration, the prior and the training pair are progressively refined while maintaining consistency with CT acquisition physics. Experiments on simulated data from AAPM, LIDC-IDRI, and LoDoPaB-CT and on real LDCT data show consistent gains over the evaluated self-supervised baselines across dose levels, with performance comparable to the evaluated supervised baseline.
Xian-Lei Han, Shaoyu Wang, Jiancheng Fang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.