Skip to content

Author

Shelley Zixin Shu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

INFORMER-Interpretability-Founded Monitoring of Medical Image Deep Learning Models: Application to Chest X-ray Pathologies.

Deep learning has demonstrated strong performance in medical imaging. However, its limited interpretability remains a major barrier to clinical trust and safe deployment. This limitation is particularly relevant in multi-label classification, where quality control methods are still underdeveloped and commonly rely only on model outputs, without incorporating gradient-level information that may better reflect prediction reliability. In this study, we propose a quality control framework for multi-label medical image classification that improves both reliability and interpretability. The framework includes a graph-based class-distinctiveness method that analyzes saliency-derived information to identify unreliable predictions, as well as a retrieval-based extension that provides case-based explanations for flagged outputs. The proposed methods were evaluated on the CheXpert dataset and compared with established output-based quality control approaches. Robustness was assessed using bootstrapped test sets, and differences in ranking across bootstrap samples were analyzed using the Wilcoxon signed-rank test. The proposed framework outperformed baseline methods, achieving a higher mean F1 score (0.574 vs. 0.563), while the best-performing variant showed higher sensitivity (0.752 vs. 0.700). In bootstrapped analyses, it achieved better mean ranks than the baseline approaches. Averaged across input-image noise levels of 0.001-0.005, under IxG-based evaluation, the proposed framework showed improvements in bootstrapped F1 over the baseline, with the retrieval-based variant achieving a 21.92% improvement and the corresponding non-retrieval variant achieving a 13.53% improvement. Clinicians further evaluated the retrieved examples to determine their relevance for interpreting flagged predictions. These findings indicate that gradient-level and graph-based analysis can enhance the effectiveness, transparency, and clinical applicability of quality control in multi-label medical image classification.

Shelley Zixin Shu, Aurélie Pahud de Mortanges, A. Poellinger et al. · 0 citations
Preprint Aug 2026

HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts

Recent vision-language models for chest X-ray understanding are largely built on image-report alignment and therefore rely heavily on MIMIC-CXR as the dominant pretraining source. While effective at scale, this paradigm underexplores an important alternative source of supervision: a range of existing multi-label classification datasets, which provide cleaner and more explicit disease signals than free-text reports, and can offer broader pathology coverage when combined across sources. However, learning from such heterogeneous datasets is nontrivial, as differences in label ontologies, annotation protocols, acquisition pipelines, and report styles can cause models to entangle clinical semantics with dataset identity, leading to poor transfer despite increased scale. In this work, we revisit radiology VLM construction from the perspective of harmonized multi-source learning. We propose HarMoE, a dataset-aware mixture-of-experts framework that learns shared cross-dataset medical semantics while confining source-specific variation to lightweight residual experts in deeper decoder layers. To further exploit clean supervision from labeled datasets, we train in a unified disease vocabulary with masked multi-dataset supervision, enabling the model to leverage complementary annotations without introducing false negatives. Experiments on large-scale chest X-ray benchmarks show that HarMoE consistently improves zero-shot classification, out-of-distribution transfer, and grounding over strong baselines. Our results suggest that building robust radiology VLMs requires moving beyond single-source image-report alignment toward structured knowledge construction from heterogeneous datasets with cleaner supervision and broader coverage. Code and the 873k harmonized dataset will be released at https://github.com/Roypic/harmoe.

Haozhe Luo, Zi-Yu Zhou, Shelley Zixin Shu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.