Multiple sequence alignment (MSA) Pairformer is presented, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains.
Abstract
Protein-protein interactions underlie biological complexity, and modeling their coevolution is essential for characterizing and engineering molecular assemblies. While protein and genomic language models have excelled at modeling individual proteins, extending these capabilities to protein complexes remains challenging. We present multiple sequence alignment (MSA) Pairformer, a protein language model that builds on AlphaFold2/3's bidirectional refinement between sequence and pairwise residue representations to accurately model the evolution of protein-protein interactions, despite training exclusively on individual chains. MSA Pairformer achieves nearly 3-fold improvement over existing methods in predicting protein-protein interface contacts and better distinguishes binding from non-binding sequences. A learned attention mechanism selectively weights sequences by their inferred evolutionary relevance, enabling discovery of subfamily-specific contacts. On single-protein benchmarks, it achieves state-of-the-art contact prediction and strong variant effect prediction using only 111 million parameters, over two orders of magnitude smaller than frontier models. These results offer an evolutionarily grounded, computationally efficient alternative to the scaling paradigm.
HyBind-NN is developed, a multimodal graph neural network that integrates protein language models (PLMs) with 3D structural and dynamic datasets to predict protein–protein and protein–peptide affinity, and it is demonstrated that combining ESM-2 sequence embeddings with precise 3D Voronoi spatial geometry enables accurate affinity predictions across diverse structural datasets.
E. A. Bogdanova, A. Chernukhin, Alexey K. Shaytan· International Journal of Mol...· 0 citations
This mini-review summarizes recent developments in devising and applying protein language models for biological sequences, emphasizing viral protein analysis, and outlines a road map for the potential application of LLMs in empowering virology research and pathogen surveillance.
Tianyi Fei, Siqi Li, Ziyue Yang et al.· Briefings in Bioinformatics· 0 citations
This work argues that explicitly modeling at the functional motif level provides both mechanistic insight into sequence-function relationships and interpretable control over protein generation, an important step toward compositional design of novel protein functions.
Boon How Low, W. Goh, Boyang Li et al.· Proceedings of the 32nd ACM...· 0 citations
This framework provides a clearer understanding of how methodological shifts have shaped the capabilities, limitations, and practical roles of recent models.
Wengan He, Yongsheng Luo, Lihong Jiang et al.· 0 citations
Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution and transcriptional deployment across the global ocean.
Prot-ΔΔG is introduced, a purely sequence-based deep learning framework that integrates large-scale pre-trained protein language models with a BiGRU-DBRNN encoder that effectively captures evolutionary and context-dependent patterns without relying on structural inputs.
Han Zhou, Yuxiang Wang, Xiumin Shi et al.· Journal of Computer-Aided Mo...· 0 citations