Skip to content
Preprint

Beyond Representation Learning: A Systematic Study of Joint-Embedding Predictive Generation for 3D Brain MRI

Aug 2026 · 0 citations · 60 references
Computer Science

TL;DR

Findings establish joint-embedding predictive generation as a promising direction for 3D medical image synthesis and encourage further research in this direction.

Abstract

Joint-embedding predictive architectures (JEPAs) have primarily been developed for self-supervised representation learning. Denoising JEPA (D-JEPA) recently demonstrated strong generative capabilities on natural images, yet the applicability to 3D medical imaging remains unexplored. Building on the D-JEPA framework, we present Med-D-JEPA, a systematic adaptation and evaluation of joint-embedding predictive generation for 3D brain MRI. Med-D-JEPA operates on continuous latent tokens produced by a 3D KL-regularized adversarial variational autoencoder, and combines masked context prediction, representation-level alignment, per-token diffusion, and iterative next-set-of-token sampling. We evaluate unconditional and class-conditional generation quality on BraTS2019 and OASIS-1 datasets; downstream classification utility; and preliminary whole-tumor segmentation on BraTS2020. Across different generation settings, Med-D-JEPA achieves superior or competitive performance compared to several strong baselines on fidelity and diversity metrics. Compared to training with real samples, Med-D-JEPA-based synthetic pretraining improves classification AUC from 0.63 to 0.85 on BraTS2019 and from 0.78 to 0.87 on OASIS-1. In the segmentation study, pretraining on Med-D-JEPA samples improves Dice from 0.74 to 0.80 and reduces HD95 from 13.40 to 9.56 mm. These findings establish joint-embedding predictive generation as a promising direction for 3D medical image synthesis and encourage further research in this direction.

View source

Similar papers

Jul 2026

Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography

Rad-JEPA 3D, a joint-embedding predictive framework that learns volumetric CT representations by predicting the latent features of a complete scan from a masked view, is introduced and Hidden States Orthogonal Regularization (HSOR) is proposed, which aligns student-teacher hidden states and reduces feature redundancy throughout the encoder.

Quoc-Huy Trinh, Minh-Van Nguyen, Ulas Bagci · 0 citations
Preprint Aug 2026

Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation

Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers. Although representation alignment has improved efficiency in general computer vision, its role within the highly specialized endoscopic image space remains unclear. We introduce REVEAL (Representation-driven Endoscopic Visual Embedding Alignment), the largest generative foundation model for endoscopy to date, trained on GastroNet-5M (GN-5M), a multicenter dataset of 5 million endoscopic frames. Instead of depending on out-of-domain priors, REVEAL employs encoders pretrained directly on the endoscopic distribution to align diffusion latents with domain-specific visual features, preserving fine textures and intricate anatomical structures. Beyond image generation, REVEAL also serves as a powerful feature extractor; in multiple benchmarks, it delivers performance that is competitive with, and in several cases exceeds, endoscopic foundation models such as EndoViT and Endo-FM, specifically tuned for classification tasks, while demonstrating strong representation robustness under realistic imaging corruptions. REVEAL produces high-fidelity images and maintains robust structural coherence in latent-space edits such as inpainting and outpainting. This high-capacity backbone lowers the computational threshold for building specialized clinical tools, offering an open, versatile foundation for conditional synthesis, segmentation, and out-of-distribution detection in future intelligent gastroenterology systems.

Francisco Caetano, T. Jaspers, Haiko Middeljans et al. · 0 citations
Review Open access Jul 2026

Adaptive multi-stage domain unlearning for white-matter lesion segmentation

Introduction Inter-scanner variability in magnetic resonance imaging (MRI) adversely affects the diagnostic and prognostic quality of scans and necessitates the development of models that are robust to domain shift arising from the unseen scanner data. A review of recent advances in domain adaptation and domain generalization showed that the efficacy of strategies involving modifications or constraints on the latent space appears to be contingent upon the level and/or depth of supervision during model training. Methods We propose an adaptive multi-stage domain unlearning (ADMU) technique to improve robustness to unseen scanner domains. Building on the state-of-the-art segmentation framework nnU-Net, we employ deep supervision at deep encoder stages by applying domain classifier unlearning, sequentially across these stages to reduce domain-discriminative latent features. Following the self-configurable approach of nnU-Net, the auxiliary feedback loop implements an adaptive backpropagation schedule for unlearning. Experiments were conducted on four public datasets (one for training, three for testing) to benchmark white-matter lesion segmentation methods. Five benchmark models and/or strategies, spanning passive to active domain-robust training strategies, were tested, and five state-of-the-art methods were compared. Results AMDU demonstrated consistent, robust and improved cross-dataset segmentation performance on three test sets versus baseline nn-Unet variants. The advantage of AMDU was in enhanced lesion sensitivity with balanced false detections, resulting in good overall segmentation quality, as measured by segmentation overlap and relative lesion volume error. Compared to continuous domain unlearning, the adaptive scheduling balanced the adverse impact of unlearning onto the main segmentation task. Intensity-based preprocessing was found to be detrimental to segmentation performance. Discussion The proposed AMDU strategy was shown to be complementary to data augmentation. Demonstrated for white-matter lesion segmentation it relied only the FLAIR modality, simplifying preprocessing to spatial normalization to brain atlas, with no intensity harmonization, for best cross-dataset segmentation performance. The source code is available at https://github.com/Pubec/nnunetv2-unlearning.

Domen Preloznik, Ž. Špiclin · 0 citations
Preprint Aug 2026

P3CA: Encoder-Agnostic Interpretation of Vision Foundation Model Embeddings via Spatial Probing

Vision foundation models are increasingly used as reusable encoders in medical image computing, yet their high-dimensional spatial embeddings are difficult to inspect beyond downstream task performance or global dimensionality reduction. We propose position-prompted PCA (P3CA), an encoder-agnostic method for local probing of channel-rich spatial tensors. Given a user-selected spatial prompt, P3CA estimates the feature normalization and dominant covariance directions within that region, then applies the resulting projection to the full tensor to visualize where locally informative directions are expressed. This produces a region-conditioned representation lens without modifying the encoder, retraining, or requiring task-specific labels. We implement P3CA in EmbedVision, an interactive 3D Slicer-based workflow, and evaluate it across natural images, colorectal pathology foundation-model embeddings, and spatial transcriptomic tensors. Across these settings, prompted projections reveal local structure suppressed by global PCA, improve prompt-matched pathology discrimination from frozen three-dimensional projections, and support comparison between learned and measured spatial representations.

A. Jamzad, Dilakshan Srikanthan, F. Akbarifar et al. · 0 citations
Aug 2026

EP-MAE: A resource-efficient masked autoencoding framework for 3D neural representation learning.

Efficient Point Masked Autoencoders (EP-MAE), a new framework designed to significantly reduce the training cost of 3D self-supervised pre-training while maintaining strong representation quality, and provides a scalable and effective foundation for future 3D neural network models is presented.

Jian Zhu, Jiale Zhao, Cheng Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.