Under realistic dark-protein deployment conditions where relational modalities are unavailable, ProtJEPA significantly outperforms naive concatenation of remaining modalities and cross-domain evaluations on drug–target interaction and disorder prediction confirm transfer beyond training modalities.
Abstract
Over 99.9% of known protein sequences lack experimentally validated functional annotations. We present ProtJEPA, a multimodal Joint-Embedding Predictive Architecture that trains a sequence-only student encoder to predict joint embeddings spanning ten biological modalities—sequence, structure, knowledge graph, protein interactions, literature, localization, tissue expression, GO function, anatomy, and disorder—requiring only sequence at inference. The key innovation is target whitening, which eliminates severe anisotropy in joint targets (mean cosine 0.984 to 0.086) and prevents representation collapse without covariance regularization. On 1,828 held-out dark proteins with zero primary Pfam family overlap with training, ProtJEPA achieves 58.07% Hit@10 on zero-shot GO retrieval (+2.80 pp, p = 0.020), 69.99% enzyme class accuracy (+9.64 pp, p < 0.001), and +11.87 pp subcellular localization at 1% labels (p < 0.001). Under realistic dark-protein deployment conditions where relational modalities are unavailable, ProtJEPA significantly outperforms naive concatenation of remaining modalities. Cross-domain evaluations on drug–target interaction and disorder prediction confirm transfer beyond training modalities, with the T1-only < ESMC < ProtJEPA ordering replicated across six independent tasks. Ablations establish that Phase 1 aggregator pretraining and target whitening are each independently load-bearing.
FuncSeek is described, a contrastive learning model which utilizes three diverse, complementary PLMs: ESM2 (to model evolutionary co-variation), ProstT5 (for bilingual sequence and structure embeddings), and ProteinBERT (for functional semantic similarities) that each capture a different aspect of protein biology: evol...
Leendert J. Cloete, Hugh G. Patterton· bioRxiv· 0 citations
Attention weight analysis shows the model identifies biologically relevant signals, including C-terminal membrane regions and N-terminal mitochondrial sequences, confirming that ESM-2 embeddings enable effective performance with simplified architectures.
Johaimen Omar· International journal of lif...· 0 citations
Over eight diverse protein foundational models trained on 550,120 SwissProt proteins with AlphaFold structures, enriched embeddings improved zero-shot remote homology retrieval, increasing Precision@10 and MRR by up to 0.13 and 0.11, respectively.
Gabriel Bianchin de Oliveira, Fahad Saeed· bioRxiv· 0 citations
Overall, DrugPLMFormer provides a reproducible, leakage-aware framework for retrospective sequence-based druggability screening and target prioritization, while prospective validation and experimental confirmation remain necessary before operational deployment.
Z. Kafi, Khosro Rezaee, Hossein Eslami· Journal of King Saud Univers...· 0 citations
This work introduces an automated and scalable method for interpreting SAE features in ESM-2 by using geometrically inspired features of the protein $\text{C}_\alpha$ backbone, providing a robust method of annotating proteins activated within SAE neurons at a residue level, providing a bridge between mechanistic interp...
S. Setlur, Djordje Mihajlovic, Darrick Lee· 0 citations
RGLLA-PPIS, a novel multimodal prediction model that integrates retrieval-augmented learning and residual GNNs for PPIS identification, outperforms several state-of-the-art baselines in both accuracy and robustness and demonstrates its potential to guide real-world protein engineering tasks.
Jia Mi, Ya-Wen Liu, Chong Chu et al.· IEEE Transactions on Neural...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.