Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· 0 citations· 33 references
TL;DR
A simple and efficient self-supervised pre-training framework for 3D medical images based on a two-fold patch-wise perturbation strategy, requiring substantially less memory, computation, and training time than the state-of-the-art pre-training pipelines.
Abstract
Self-supervised pre-training has become a key paradigm for reducing annotation costs in 3D medical imaging, yet many recent approaches rely on complex objectives or incur substantial computational overhead. We propose a simple and efficient self-supervised pre-training framework for 3D medical images based on a two-fold patch-wise perturbation strategy. The method applies Bernoulli patch masking and discrete rotations, and trains a shared encoder with a three-head objective for reconstruction, perturbation localization, and rotation prediction. This design encourages spatially aware and transferable representations while remaining computationally lightweight. Experiments across diverse segmentation and classification benchmarks, including modality-shift scenarios, demonstrate consistent improvements over general self-supervised baselines and competitive or superior performance compared to recent medical self-supervised methods, while requiring substantially less memory, computation, and training time than the state-of-the-art pre-training pipelines.
This work proposes XPos3R, a generalizable pose regression method that eliminates preoperative preparation, and introduces an asymmetric encoder-decoder architecture that improves cross-modal feature alignment while maintaining computational efficiency.
Shi-Yan Su, Ruyi Zha, Hong-Dong Li et al.· 0 citations
Few-shot medical image segmentation relies on dense, boundary-sensitive prototype matching, yet common pre-training objectives mainly optimize global alignment or reconstruction, creating an objective gap that hurts boundary delineation and increases adaptation cost. This raises the question: how to pre-train represent...
Shou-Peng Chen, Yi-Ming Miao, Li-Mei Peng et al.· Proceedings of the Thirty-Fi...· 0 citations
Medical image segmentation is still challenging, especially in semi-supervised scenarios where limited annotations are expected to support both accurate boundary delineation and coherent anatomical structures. We propose LR-GCF, a Local-Rotation-Driven Global Consistency Framework that couples strong local geometric pe...
Zhen Yang, Dong-Shuai Zhang, Yun-Liang Qi et al.· Proceedings of the Thirty-Fi...· 0 citations
Experimental results show that the approach outperforms the state-of-the-art on the Pancreas-CT dataset by a large margin and enables rapid transfer learning from 2D-pretrained models to 3D medical tasks with few labeled data, making it especially valuable for rare disease diagnosis.
Ruize Shi, Xiao-Yan Li, Bo-Yue Wang et al.· International Conference on...· 0 citations
B-MIM is introduced, a modification of the iBOT objective that stochastically reduces global semantic alignment to prioritize local patch reconstruction and suggests that reducing global semantic pressure during pretraining enhances generalization to intricate anatomical structures.
S. González, Karen Sanchez, J. M. Saavedra et al.· 0 citations