Findings establish joint-embedding predictive generation as a promising direction for 3D medical image synthesis and encourage further research in this direction.
Abstract
Joint-embedding predictive architectures (JEPAs) have primarily been developed for self-supervised representation learning. Denoising JEPA (D-JEPA) recently demonstrated strong generative capabilities on natural images, yet the applicability to 3D medical imaging remains unexplored. Building on the D-JEPA framework, we present Med-D-JEPA, a systematic adaptation and evaluation of joint-embedding predictive generation for 3D brain MRI. Med-D-JEPA operates on continuous latent tokens produced by a 3D KL-regularized adversarial variational autoencoder, and combines masked context prediction, representation-level alignment, per-token diffusion, and iterative next-set-of-token sampling. We evaluate unconditional and class-conditional generation quality on BraTS2019 and OASIS-1 datasets; downstream classification utility; and preliminary whole-tumor segmentation on BraTS2020. Across different generation settings, Med-D-JEPA achieves superior or competitive performance compared to several strong baselines on fidelity and diversity metrics. Compared to training with real samples, Med-D-JEPA-based synthetic pretraining improves classification AUC from 0.63 to 0.85 on BraTS2019 and from 0.78 to 0.87 on OASIS-1. In the segmentation study, pretraining on Med-D-JEPA samples improves Dice from 0.74 to 0.80 and reduces HD95 from 13.40 to 9.56 mm. These findings establish joint-embedding predictive generation as a promising direction for 3D medical image synthesis and encourage further research in this direction.
Rad-JEPA 3D, a joint-embedding predictive framework that learns volumetric CT representations by predicting the latent features of a complete scan from a masked view, is introduced and Hidden States Orthogonal Regularization (HSOR) is proposed, which aligns student-teacher hidden states and reduces feature redundancy throughout the encoder.
Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers. Although representation alignment has improved efficiency in general computer vision, its role within the highly specialized endoscopic image space remains unclear. We introduce REVEAL (Representation-driven Endoscopic Visual Embedding Alignment), the largest generative foundation model for endoscopy to date, trained on GastroNet-5M (GN-5M), a multicenter dataset of 5 million endoscopic frames. Instead of depending on out-of-domain priors, REVEAL employs encoders pretrained directly on the endoscopic distribution to align diffusion latents with domain-specific visual features, preserving fine textures and intricate anatomical structures. Beyond image generation, REVEAL also serves as a powerful feature extractor; in multiple benchmarks, it delivers performance that is competitive with, and in several cases exceeds, endoscopic foundation models such as EndoViT and Endo-FM, specifically tuned for classification tasks, while demonstrating strong representation robustness under realistic imaging corruptions. REVEAL produces high-fidelity images and maintains robust structural coherence in latent-space edits such as inpainting and outpainting. This high-capacity backbone lowers the computational threshold for building specialized clinical tools, offering an open, versatile foundation for conditional synthesis, segmentation, and out-of-distribution detection in future intelligent gastroenterology systems.
Francisco Caetano, T. Jaspers, Haiko Middeljans et al.· 0 citations
Introduction Inter-scanner variability in magnetic resonance imaging (MRI) adversely affects the diagnostic and prognostic quality of scans and necessitates the development of models that are robust to domain shift arising from the unseen scanner data. A review of recent advances in domain adaptation and domain generalization showed that the efficacy of strategies involving modifications or constraints on the latent space appears to be contingent upon the level and/or depth of supervision during model training. Methods We propose an adaptive multi-stage domain unlearning (ADMU) technique to improve robustness to unseen scanner domains. Building on the state-of-the-art segmentation framework nnU-Net, we employ deep supervision at deep encoder stages by applying domain classifier unlearning, sequentially across these stages to reduce domain-discriminative latent features. Following the self-configurable approach of nnU-Net, the auxiliary feedback loop implements an adaptive backpropagation schedule for unlearning. Experiments were conducted on four public datasets (one for training, three for testing) to benchmark white-matter lesion segmentation methods. Five benchmark models and/or strategies, spanning passive to active domain-robust training strategies, were tested, and five state-of-the-art methods were compared. Results AMDU demonstrated consistent, robust and improved cross-dataset segmentation performance on three test sets versus baseline nn-Unet variants. The advantage of AMDU was in enhanced lesion sensitivity with balanced false detections, resulting in good overall segmentation quality, as measured by segmentation overlap and relative lesion volume error. Compared to continuous domain unlearning, the adaptive scheduling balanced the adverse impact of unlearning onto the main segmentation task. Intensity-based preprocessing was found to be detrimental to segmentation performance. Discussion The proposed AMDU strategy was shown to be complementary to data augmentation. Demonstrated for white-matter lesion segmentation it relied only the FLAIR modality, simplifying preprocessing to spatial normalization to brain atlas, with no intensity harmonization, for best cross-dataset segmentation performance. The source code is available at https://github.com/Pubec/nnunetv2-unlearning.
Domen Preloznik, Ž. Špiclin· Frontiers in Medicine· 0 citations
Results demonstrate that Adaptive-CGAN can improve diagnostic performance, synthetic image quality, computational efficiency, and environmental sustainability in AI-assisted healthcare.
W. Saber, A. El-Baz, R. Rizk· Cluster Computing· 0 citations
Vision foundation models are increasingly used as reusable encoders in medical image computing, yet their high-dimensional spatial embeddings are difficult to inspect beyond downstream task performance or global dimensionality reduction. We propose position-prompted PCA (P3CA), an encoder-agnostic method for local probing of channel-rich spatial tensors. Given a user-selected spatial prompt, P3CA estimates the feature normalization and dominant covariance directions within that region, then applies the resulting projection to the full tensor to visualize where locally informative directions are expressed. This produces a region-conditioned representation lens without modifying the encoder, retraining, or requiring task-specific labels. We implement P3CA in EmbedVision, an interactive 3D Slicer-based workflow, and evaluate it across natural images, colorectal pathology foundation-model embeddings, and spatial transcriptomic tensors. Across these settings, prompted projections reveal local structure suppressed by global PCA, improve prompt-matched pathology discrimination from frozen three-dimensional projections, and support comparison between learned and measured spatial representations.
A. Jamzad, Dilakshan Srikanthan, F. Akbarifar et al.· 0 citations
Efficient Point Masked Autoencoders (EP-MAE), a new framework designed to significantly reduce the training cost of 3D self-supervised pre-training while maintaining strong representation quality, and provides a scalable and effective foundation for future 3D neural network models is presented.
Jian Zhu, Jiale Zhao, Cheng Lin et al.· Neural Networks· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.