Skip to content

Author

Dongxiao Zhu

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Cross-fraction prior learning for scalable organ-at-risk segmentation in abdominal MR-guided radiotherapy.

BACKGROUND Manual organ-at-risk (OAR) delineation takes 20-40 min per case, a major bottleneck within the 50-90 min treatment window of abdominal MR-guided adaptive radiotherapy (MRgRT). Most deep learning systems adopt single-fraction approaches that discard valuable temporal context from prior treatment fractions. PURPOSE This study develops AdaptSeg, a scalable framework leveraging cross-fraction anatomical priors to substantially improve OAR segmentation without per-patient retraining. METHODS We implemented a dual-path neural architecture conditioning current fraction segmentation on paired image-mask information from supporting fractions. AdaptSeg was instantiated with convolutional (3D UNet) and transformer-based (SwinUNETR) backbones. Evaluation used 104 pancreatic cancer patients across 520 treatment fractions for four abdominal organs (colon, duodenum, small bowel, stomach), with patient-level splitting: 72 training, 10 validation, 22 test patients. Performance metrics included Dice Similarity Coefficient (DSC), 95th percentile Hausdorff Distance (HD95), and Average Symmetric Surface Distance (ASSD); paired comparisons used two-sided Wilcoxon signed-rank tests with Benjamini-Hochberg correction, and 95% bootstrap confidence intervals for the means. RESULTS Cross-fraction priors improved segmentation performance for both tested backbones. The 3D UNet achieved 87.22% mean DSC versus 83.78% baseline, while SwinUNETR reached 85.19% versus 82.49% baseline. For highly deformable organs, improvements included up to 7.0 percentage point DSC gains (small bowel: 77.8% to 84.8%, p < 0.001 ) and 62% boundary error reduction (colon HD95: 23.13 to 8.74 mm, p < 0.001 ). Compared to nine state-of-the-art methods, AdaptSeg achieved the best overall performance with substantial improvements in mean DSC (4.1%), HD95 (43%), and ASSD (39%) over the strongest baseline. All variants maintained computationally feasible inference under 1.6 s per case. Temporal prior selection showed a backbone-dependent preference: the CNN favored sequential priors, and the transformer additionally benefited from randomized support sampling during training (sequential support is used at inference). CONCLUSIONS Cross-fraction anatomical priors improved OAR segmentation for both tested backbone families, indicating that temporal context is an underutilized resource in fractionated radiotherapy. AdaptSeg provides a scalable, computationally feasible framework for accelerating MRgRT workflows without per-patient adaptation, with sub-1.6 s inference compatible with the time constraints of online adaptive treatment.

Chengyin Li, D. Rusu, Rafi Ibn Sultan et al. · 0 citations
Preprint Aug 2026

MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation

Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text-guided segmentation methods within the Vision-Language Model (VLM) paradigm often use language only as a late conditioning signal, limiting its influence on visual representation learning. We introduce MedPlex (Medical Plexus of Vision and Language), an end-to-end VLM framework that makes text guidance a continuous, clinically grounded component of segmentation learning. Through Bi-Fusion (Bidirectional Fusion), visual and textual representations evolve jointly across the encoding hierarchy. MedPlex further introduces class-level and region-level concept alignment to organize the shared representation at complementary granularities. Class-level alignment anchors each anatomical target to an aggregated clinical concept profile, while region-level alignment preserves individual concepts, such as shape, location, appearance, and texture, through class-specific visual evidence. In this way, language provides structured supervision throughout the encoder rather than serving only as a late-stage cue. MedPlex achieves state-of-the-art performance across CT and MR benchmarks for multi-organ, cardiac substructure, and tumor segmentation, including settings with real free-text clinical supervision. Code: https://github.com/rafiibnsultan/MedPlex.

R. Sultan, Hui Zhu, Chengyin Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models

The results show that PURGE consistently reduces hallucinations and spurious-correlation-driven errors while maintaining or improving overall performance in most evaluated settings, providing both a reusable evaluation protocol and an effective mitigation framework for more reliable LVLMs.

Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.