Improving GI Polyp Segmentation with a Balanced-Mix-Driven
Gastrointestinal (GI) endoscopy is a cornerstone diagnostic procedure for detecting inflammatory diseases, polyps, and early-stage cancers. Recent advances in deep learning have significantly improved automated endoscopic image analysis; however, their performance remains limited by scarce annotations, severe class imbalance, and poor generalization across diverse imaging conditions. Moreover, jointly learning classification and segmentation poses additional challenges due to task imbalance and the high annotation cost of pixel-level labels. To address these limitations, we propose a Balanced-Mix-Driven framework that leverages 99,417 unlabeled images from the HyperKvasir dataset through Self-Supervised Learning (SSL)-based pretraining. Our core contribution, Balanced-Mix, is an interpolation strategy that progressively shifts from coarse to fine-grained mixing during pretraining, preventing trivial representation learning. Experimental results on the Kvasir-SEG dataset demonstrate that our method achieves a Dice score of $\mathbf{9 2 . 4 0} \boldsymbol{\%}$ and an mIoU of $\mathbf{8 7 . 0 1 \%}$, outperforming established baselines such as UNet++ and ResUNet++. This validates the effectiveness of curriculum-based self-supervised learning for dense medical prediction tasks.