Medical Image Segmentation Based on the CLIP and SAM Base Models
Abstract
. Medical image segmentation is a key technology for achieving precise medical care. Traditional methods rely on a large amount of labeled data and have limited generalization capabilities. Visual foundation models represented by CLIP and SAM have obtained general capabilities through large-scale pre-training, providing a new paradigm for medical segmentation. CLIP achieves semantic-guided segmentation through image-text alignment, while SAM achieves general segmentation through prompt interaction. Both can effectively reduce the reliance on labeled data. This paper systematically reviews the applications of these two types of models in medical segmentation: first, analyzes the necessity of their applications; second, separately summarizes the technical routes and adaptation progress of semantic-guided methods based on CLIP and prompt-interaction methods based on SAM; finally, discusses their advantages and challenges. The review shows that these models significantly improve data efficiency and cross-domain generalization capabilities, but still need further exploration in medical specificity adaptation and fine-grained segmentation.