Jul 2026· Signal Processing and Communications Applications Conference· pp. 1-4· 0 citations· 41 references
Abstract
Vision-language models (VLMs) have significantly advanced open-vocabulary image understanding by learning aligned representations from large-scale image-text datasets. Despite their zero-shot generalization capabilities, adapting these foundation models to specific application domains remains challenging. Full fine-tuning is often infeasible due to computational costs and the risk of overfitting when labeled data are limited. Parameter-efficient fine-tuning (PEFT) approaches promise to address this issue by updating only a small set of parameters while keeping the pre-trained encoders largely frozen. However, existing PEFT strategies often exhibit insufficient spatial awareness, rendering them suboptimal for medical imaging, where subtle visual differences can lead to distinct clinical diagnoses. To address this challenge, in this study we propose a novel PEFT technique that leverages natural-image pretrained DINOv3’s attention maps to enforce spatial alignment.
This survey comprehensively evaluates the underlying mechanisms, inherent strengths, and specific weaknesses of each FSIC methods into three categories: Metric Learning, Optimization/Meta-Learning, and Transfer Learning with Large Model Fine-Tuning.
Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely knowwhichinternal units encode clinical findings orwherethat information lives in the representation. We first study this on a 3D chest vision-language model (Pillar-0) by probing its frozen vision embeddin...
F. Nooralahzadeh, L. Bogensperger, C. Bluethgen et al.· arXiv.org· 0 citations
Recent advances in vision-language models (VLMs) have shown remarkable performance in medical image classification tasks. However, applying VLMs to fetal cardiac ultrasound (FCU) remains challenging due to compound distribution shifts, including covariate shifts caused by cross-center heterogeneity and semantic shifts...
Zi-Yi Liu, Wei-Hu Song, Yu-Peng Ma et al.· Proceedings of the Thirty-Fi...· 0 citations
Foundation models for medical image segmentation, like prompt-based MedSAM, generalize well across domains and modalities, often in zero or few-shot setups. However, their performance depends on the quality of prompts and the adaptation of the models to custom datasets. This work systematically examines how MedSAM gene...
Marko Haralović, Sounic Akkaraju, Carlo Baretta et al.· 1 citation· ⚡1
TopKSigLIP outperforms existing open-source mammography and general medical VLMs on both internal and external benchmarks on density assessment, BI-RADS classification, finding subtyping, and cancer prediction under zero-shot evaluation.
Y. Jeon, Beatrice Brown-Mulry, R. Isaac et al.· 0 citations
Medical image segmentation is a key component of computer-aided diagnosis and treatment planning. Despite substantial progress in deep learning–based models, most existing approaches depend heavily on large annotated datasets and often fail to generalize across heterogeneous clinical environments, limiting their deploy...
V. Nguyen, Hoang Quan Luong, Phuc Ngoc Pham· IEEE International Conferenc...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.