Cross-modal bias in medical vision-language models: a pipeline-aware framework for mechanisms, evaluation, and mitigation
Medical vision-language models encode images and clinical text in a shared representation. Across radiology and ophthalmology, their diagnostic performance now approaches that of specialist clinicians. The mechanism behind that performance is also the source of a problem that has gone largely unexamined. These models a...