Concept-Enhanced Multi-Scale Cross-Modal Alignment for Medical Visual Representation Learning.
Medical Vision-Language Pre-training (Med VLP) on paired medical images and reports has emerged as a promising direction for learning visual representations. However, current alignment approaches remain insufficient for learning fine-grained pathological details, largely due to the inherent difficulty of tokenizing rep...