DisenTextSeg: Disentangling Textual Information from Clinical Reports for Multi-Modal Medical Image Segmentation
Abstract
Accurate segmentation of multi-modal magnetic resonance imaging (MRI) is central to clinical diagnosis, treatment planning, and prognostic evaluation of tumors. Existing methods predominantly rely on imaging data alone, neglecting the rich semantic information embedded in clinical reports. Although recent vision-language fusion approaches have begun to incorporate textual features, they share two common limitations: they adopt a one-report-per-case design that assigns a single unified text to all four MRI modalities, and they fuse text and image features indiscriminately without distinguishing shared anatomical location descriptions from modality-specific imaging findings, leading to feature redundancy and degraded segmentation accuracy on subtle sub-regions. We propose DisenTextSeg, a novel framework that explicitly disentangles shared and modality-specific features from per-modality clinical reports and integrates them into a diffusion-based segmentation network. Shared anatomical features are extracted via learnable queries and multi-head cross-attention, while modality-unique features are recovered through orthogonal subtraction. The disentangled features are injected separately into the dual encoders of the segmentation backbone, providing complementary guidance from global anatomical context and modality-specific appearance cues. Additionally, region-specific prompt masks derived from the MNI152 brain atlas are automatically generated from the report text and incorporated as spatial priors. Evaluated on TextBraTS23 and a private cervical cancer dataset, DisenTextSeg achieves average Dice scores of 89.37% and 59.83%, respectively, outperforming competitive baselines and demonstrating consistent cross-domain performance.