Deep learning has demonstrated strong performance in medical imaging. However, its limited interpretability remains a major barrier to clinical trust and safe deployment. This limitation is particularly relevant in multi-label classification, where quality control methods are still underdeveloped and commonly rely only on model outputs, without incorporating gradient-level information that may better reflect prediction reliability. In this study, we propose a quality control framework for multi-label medical image classification that improves both reliability and interpretability. The framework includes a graph-based class-distinctiveness method that analyzes saliency-derived information to identify unreliable predictions, as well as a retrieval-based extension that provides case-based explanations for flagged outputs. The proposed methods were evaluated on the CheXpert dataset and compared with established output-based quality control approaches. Robustness was assessed using bootstrapped test sets, and differences in ranking across bootstrap samples were analyzed using the Wilcoxon signed-rank test. The proposed framework outperformed baseline methods, achieving a higher mean F1 score (0.574 vs. 0.563), while the best-performing variant showed higher sensitivity (0.752 vs. 0.700). In bootstrapped analyses, it achieved better mean ranks than the baseline approaches. Averaged across input-image noise levels of 0.001-0.005, under IxG-based evaluation, the proposed framework showed improvements in bootstrapped F1 over the baseline, with the retrieval-based variant achieving a 21.92% improvement and the corresponding non-retrieval variant achieving a 13.53% improvement. Clinicians further evaluated the retrieved examples to determine their relevance for interpreting flagged predictions. These findings indicate that gradient-level and graph-based analysis can enhance the effectiveness, transparency, and clinical applicability of quality control in multi-label medical image classification.
Shelley Zixin Shu, Aurélie Pahud de Mortanges, A. Poellinger et al.· Journal of imaging informati...· 0 citations
Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution Whole Slide Images (WSIs), limiting their generalization across arbitrary resolutions. Gigapixel WSIs inherently contain diagnostic patterns at multiple scales, including cellular morphologies, tissue architectures, and global context, mirroring how expert pathologists examine WSIs. We introduce Multi-Resolution Pyramid Transformer (MRPT), a model that hierarchically aggregates multi-resolution information from cellular to tissue and WSI levels. MRPT employs a biologically meaningful Consecutive Cross-Resolution Attention (CCRA) mechanism to capture scale-independent interactions and enforces multi-resolution semantic consistency by aligning embeddings across resolutions, yielding robust and generalizable WSI representations. Pre-trained in a multi-resolution self-supervised manner on 624M patches, 2.4M regions, and 36K WSIs, MRPT learns rich coarse-to-fine histopathology features. Extensive experiments on 34 diverse datasets show that MRPT surpasses recent foundation models and Multimodal Large Language Models (MLLMs) in cancer subtype classification, tissue phenotyping, and Visual Question Answering (VQA) for WSI understanding.
B. Alawode, Moshira Abdalla, Dwarikanath Mahapatra et al.· 0 citations
Medical vision-language models (MVLMs) promise broad zero-shot generalization, yet their reliability collapses when confronted with unseen modalities and domains, precisely where clinical robustness matters most. To address this gap, we revisit test-time modality generalization from the perspective of Mixture-of-Experts (MoE) and ask: can experts route-and-adapt without any optimization during inference? We identify a fundamental specialization-generalization dilemma at test time, where blindly aggregating modality experts dilutes modality-specific knowledge, while selecting one highly confident expert risks mismatch under shift. To address this, we propose MoBE: a fully optimization-free framework that performs dynamic expert selection and adaptation at test time. MoBE combines entropy-guided dynamic routing in MoE settings with expert-wise Bayesian adaptation, enabling experts to update their confidence and adapt online without gradient updates. Without parametric updates, MoBE augments a static MVLM with test-time routing and online statistics, achieving average accuracy gains of +4.72, +7.17, and +4.3 over state-of-the-art TTA methods across seen, unseen, and heterogeneous medical benchmarks, highlighting the effectiveness of training-free expert adaptation for robust modality generalization.
Raza Imam, Darakshan Rashid, Yutong Xie et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.