Accurate medical image segmentation plays a vital role in clinical diagnostics by facilitating the precise delineation of anatomical structures and pathological regions. However, the performance of existing segmentation methods is often constrained by the scarcity of high-quality annotated datasets, as manual labeling is both labor-intensive and reliant on domain-specific expertise. To address this limitation without requiring additional annotations, we propose a novel multimodal segmentation framework that leverages medical text annotations as an auxiliary modality to complement visual information. In particular, our approach introduces a learnable encoding strategy for joint distribution modeling of image and text, which enables discriminative fusion and effectively suppresses cross-modal redundancy. Moreover, we innovatively design a frequency-domain prompt encoder based on the discrete wavelet transform (DWT) to capture multi-frequency features, thereby significantly enhancing the model’s ability to delineate fine-grained boundaries. Overall, our framework integrates cross-attention for effective cross-modal interaction, employs joint distribution modeling to enable discriminative and redundancy-reduced multimodal fusion, and incorporates auxiliary supervision to strengthen the learning of task-relevant features. Extensive experiments on nine public datasets across three clinical tasks—including cell, lung infection, and polyp segmentation—demonstrate that our method achieves competitive segmentation performance while maintaining favorable computational efficiency. Comprehensive ablation studies and feature distribution visualizations further validate the effectiveness and robustness of our proposed components. The code will be made publicly available at https://github.com/chenpeng052/MDFP
Chao Huang, Peng Chen, Jie Wen et al.· IEEE Transactions on Image P...· 0 citations
Graph Contrastive Learning (GCL) is a popular self-supervised learning (SSL) technique. However, mainstream GCL methods usually favor single fine-grained random augmentation schemes, which will destroy the structural integrity of the graph, and they largely ignore the topology of the graph structure, that is, multi-granularity characteristics. Graphs are typically composed of homogeneous regions with varying granularities, where nodes within a region exhibit strong homogeneous properties. However, most of the real graphs are heterogeneous, and in the local regions of heterogeneous graphs, interconnected nodes may have similar semantic information, even if they do not belong to the same class. To this end, we propose a new multi-granularity graph contrastive learning framework via granular-ball (GBGCL) to explore the potential on heterogeneous graphs. Specifically, we develop an adaptive granular-ball augmentation strategy that identifies multi-granularity homogeneous regions in the heterogeneous graph, and treat nodes in the same granular-ball as positive pairs, while nodes in different granular-balls as negative pairs. In addition, we integrate the feature information of the nodes in original graph and semantic information of similar nodes in the feature space. Node representations are obtained through joint optimization of losses. Experiments on heterogeneous graphs demonstrate the unique advantages of our framework.
Shuyin Xia, Guan Wang, Cheng Tan et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.