This work proposes Co-occurrence Weighted Adaptation (CoWA), which leverages disease co-occurrence patterns as a reliability signal for adaptation, enabling adaptation to rely more on consistent predictions while reducing the impact of noisy ones.
Abstract
Medical imaging models often degrade when deployed at new clinical sites due to differences in imaging equipment, protocols, and patient populations. Test-time adaptation (TTA) addresses this by updating a pretrained model using only unlabeled target data, without access to source data. However, existing TTA methods were designed for single-label classification on natural image benchmarks, minimizing entropy uniformly across all samples without considering label dependencies. This overlooks a key property of multi-label medical imaging: pathologies do not occur independently but exhibit structured co-occurrence patterns. In this work, we propose Co-occurrence Weighted Adaptation (CoWA), which leverages disease co-occurrence patterns as a reliability signal for adaptation. CoWA estimates label co-occurrence structure from model predictions and downweights samples that deviate from expected patterns, enabling adaptation to rely more on consistent predictions while reducing the impact of noisy ones. We evaluate CoWA on chest X-ray benchmarks under domain shifts and demonstrate consistent improvements over established baselines.
Recent vision-language models for chest X-ray understanding are largely built on image-report alignment and therefore rely heavily on MIMIC-CXR as the dominant pretraining source. While effective at scale, this paradigm underexplores an important alternative source of supervision: a range of existing multi-label classi...
Haozhe Luo, Zi-Yu Zhou, Shelley Zixin Shu et al.· 0 citations
Test-time adaptation (TTA) has become a practical way to adapt deployed models to unlabeled target data, a setting that is especially relevant in computational pathology where staining, scanner, and cohort shifts are routine. While most TTA methods are evaluated by their effect on accuracy, clinical use also depends on...
R. G. Bahumanya, M. HarshithV., S. N. Gowda et al.· 0 citations
This work systematically investigates how evaluation-reference choices affect model performance and ranking in both pathology classification and image quality assessment (IQA), and shows that for supervised image classifiers, changing the label source leads to substantial differences not only in performance estimates b...
Panagiotis Fytas, Ian Selby, C. Karner et al.· arXiv.org· 0 citations
This work proposes a label-free selection criterion built on SUDO, a framework for evaluating clinical AI systems without ground-truth annotations, and shows that AURCC can be used to rank a variety of vision-language models on chest X-ray classification across three inter-hospital shift scenarios, under zero-shot and...
J. Larrea, L. Mansilla, Enzo Ferrante· 0 citations
Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present findings are left unreported due to the omission of subtle findings. For example, prior studies show that cardiomegaly may be omitted from ICU chest...
Y. Kobayashi, P. Ramesh, Muhammad Ahmed Chaudhry et al.· 0 citations
A semantic attention-driven retrieval framework based on a lightweight Meta-Domain Adaptive Segmentation Network (MDA-SN) with an adaptive data normalization strategy to enhance infection detection in cross-dataset analysis and achieves real-time execution.
Muhammad Owais, Taimur Hassan, Naqash Afzal et al.· Scientific Reports· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.