BACKGROUND
Manual organ-at-risk (OAR) delineation takes 20-40 min per case, a major bottleneck within the 50-90 min treatment window of abdominal MR-guided adaptive radiotherapy (MRgRT). Most deep learning systems adopt single-fraction approaches that discard valuable temporal context from prior treatment fractions.
PURPOSE
This study develops AdaptSeg, a scalable framework leveraging cross-fraction anatomical priors to substantially improve OAR segmentation without per-patient retraining.
METHODS
We implemented a dual-path neural architecture conditioning current fraction segmentation on paired image-mask information from supporting fractions. AdaptSeg was instantiated with convolutional (3D UNet) and transformer-based (SwinUNETR) backbones. Evaluation used 104 pancreatic cancer patients across 520 treatment fractions for four abdominal organs (colon, duodenum, small bowel, stomach), with patient-level splitting: 72 training, 10 validation, 22 test patients. Performance metrics included Dice Similarity Coefficient (DSC), 95th percentile Hausdorff Distance (HD95), and Average Symmetric Surface Distance (ASSD); paired comparisons used two-sided Wilcoxon signed-rank tests with Benjamini-Hochberg correction, and 95% bootstrap confidence intervals for the means.
RESULTS
Cross-fraction priors improved segmentation performance for both tested backbones. The 3D UNet achieved 87.22% mean DSC versus 83.78% baseline, while SwinUNETR reached 85.19% versus 82.49% baseline. For highly deformable organs, improvements included up to 7.0 percentage point DSC gains (small bowel: 77.8% to 84.8%, p < 0.001 ) and 62% boundary error reduction (colon HD95: 23.13 to 8.74 mm, p < 0.001 ). Compared to nine state-of-the-art methods, AdaptSeg achieved the best overall performance with substantial improvements in mean DSC (4.1%), HD95 (43%), and ASSD (39%) over the strongest baseline. All variants maintained computationally feasible inference under 1.6 s per case. Temporal prior selection showed a backbone-dependent preference: the CNN favored sequential priors, and the transformer additionally benefited from randomized support sampling during training (sequential support is used at inference).
CONCLUSIONS
Cross-fraction anatomical priors improved OAR segmentation for both tested backbone families, indicating that temporal context is an underutilized resource in fractionated radiotherapy. AdaptSeg provides a scalable, computationally feasible framework for accelerating MRgRT workflows without per-patient adaptation, with sub-1.6 s inference compatible with the time constraints of online adaptive treatment.
Chengyin Li, D. Rusu, Rafi Ibn Sultan et al.· Medical Physics (Lancaster)· 0 citations
Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text-guided segmentation methods within the Vision-Language Model (VLM) paradigm often use language only as a late conditioning signal, limiting its influence on visual representation learning. We introduce MedPlex (Medical Plexus of Vision and Language), an end-to-end VLM framework that makes text guidance a continuous, clinically grounded component of segmentation learning. Through Bi-Fusion (Bidirectional Fusion), visual and textual representations evolve jointly across the encoding hierarchy. MedPlex further introduces class-level and region-level concept alignment to organize the shared representation at complementary granularities. Class-level alignment anchors each anatomical target to an aggregated clinical concept profile, while region-level alignment preserves individual concepts, such as shape, location, appearance, and texture, through class-specific visual evidence. In this way, language provides structured supervision throughout the encoder rather than serving only as a late-stage cue. MedPlex achieves state-of-the-art performance across CT and MR benchmarks for multi-organ, cardiac substructure, and tumor segmentation, including settings with real free-text clinical supervision. Code: https://github.com/rafiibnsultan/MedPlex.
R. Sultan, Hui Zhu, Chengyin Li et al.· 0 citations
The results show that PURGE consistently reduces hallucinations and spurious-correlation-driven errors while maintaining or improving overall performance in most evaluated settings, providing both a reusable evaluation protocol and an effective mitigation framework for more reliable LVLMs.
Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan et al.· 0 citations
The framework, INFUSE, first stabilizes visual and textual representations around perturbation-averaged and ground-truth anchors, then aligns the stabilized representations across modalities with bidirectional contrastive objectives.
Aditi Sarker, Rafi Ibn Sultan, Hui Zhu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.