Jul 2026· IEEE journal of biomedical and health informatics· Vol PP· 0 citations
Medicine
TL;DR
MultiMol is a multimodal framework that integrates SMILES sequences, molecular images, molecular graphs, and 3D conformations through tailored pre-training tasks and a scalable fusion mechanism, and consistently outperforms existing methods in cyclic peptide permeability prediction.
Abstract
Cyclic peptide drugs show great potential in antiviral, antibacterial, anticancer, and immunomodulatory therapies, yet accurate prediction of their membrane permeability remains challenging. Existing approaches based on SMILES, molecular graphs, or 3D structures have inherent limitations: SMILES lack spatial information, graphs inadequately capture stereochemistry, and 3D methods are sensitive to conformational variability. Moreover, current multimodal fusion strategies often fail to effectively integrate heterogeneous molecular information. To address these challenges, we propose MultiMol, a multimodal framework that integrates SMILES sequences, molecular images, molecular graphs, and 3D conformations through tailored pre-training tasks and a scalable fusion mechanism. Experiments show that MultiMol consistently outperforms existing methods in cyclic peptide permeability prediction. Visualization and interpretability analyses further demonstrate its strong feature extraction and generalization capabilities. MultiMol also prioritizes promising KRAS-targeting cyclic peptides, supporting its practical utility in virtual screening. The code is available at https://github.com/chaoxiuxiu/multi-mol.
Cyclic peptides are a promising therapeutic modality, offering the potential to target challenging intracellular protein-protein interactions involved in cancer and other diseases. However, their clinical utility is frequently restricted by poor membrane permeability. While deep learning offers new methodologies to predict permeability, current models are limited by a reliance on 2D molecular representations that fail to capture the conformational flexibility inherent to macrocycles. Existing 3D resources also lack physics-based sampling of conformational dynamics across solvent environments that are critical for membrane permeability. To bridge this gap, we present CycPeptMPDB-4D, a comprehensive dataset comprising atomistic molecular dynamics trajectories for 5,160 structurally diverse cyclic peptides, including unnatural, N-methylated, and D-residues in circle and lariat topologies. Each peptide was simulated using the AMBER14SB force field in both explicit water and hexane environments for 50 nanoseconds to generate conformational ensembles in aqueous and membrane-mimicking phases. The trajectories capture the “chameleon-like” property, evidenced by markedly reduced conformational flexibility and polar surface area in the hydrophobic phase. Technical validation demonstrates that the simulated ensembles are in high agreement with experimental NMR data, covering NMR conformers within an RMSD of 1.6 Å. The dataset provides clustered ensembles, representative structures, and specialized descriptors such as desolvation free energy. This resource is designed to facilitate the development of deep learning models that incorporate 3D or 4D (trajectory- or ensemble-based) information to improve the prediction of cyclic peptide membrane permeability.
Wei Liu, Nguyen Hung Pham, Chandra S. Verma et al.· Scientific Data· 0 citations
Cyclic peptides are promising therapeutic agents, but their clinical translation is often limited by poor membrane permeability arising from complex, solvent-dependent conformational ensembles. Here, we present a modeling framework inspired by molecular chameleon, integrating sequence and structural information by explicitly representing solvent-dependent ensembles in aqueous and nonpolar environments. The model achieves strong predictive performance (MAE = 0.29, R = 0.85, R2 = 0.70 on CycPeptMPDB) under random split and retains robust performance under scaffold split evaluation (MAE = 0.34, Pearson R = 0.68). Beyond molecule-level prediction, analysis of one-residue analog-pair in testing set shows that the model captures experimentally meaningful ΔPAMPA trends, supporting its utility for cyclic peptide lead optimization. Model interpretation highlights polarity shielding captured by ΔPSA3D as a key mechanistically interpretable feature associated with permeability. These results demonstrate that permeability emerges from ensemble reorganization rather than a single conformational transition, establishing a generalizable framework for cyclic peptide design and ADMET prediction.
Sicheng Wen, Yang Wang, Yue Qian· Journal of Chemical Informat...· 0 citations
PeptiVerse is a unified platform that leverages large foundation models to predict diverse peptide developability properties from both amino acid sequences and SMILES representations, enabling accessible, scalable analysis for peptide drug design.
Accurate prediction of drug-target binding affinity (DTA) is a key task in virtual screening. However, current computational methods face a key challenge: sequence-based approaches often fail to capture critical spatial information, while structure-based models rely on computationally expensive 3D coordinates, which restrict their scalability. To address this issue, we propose StructuraDTA, a novel multimodal framework that adopts an implicit structure modeling strategy. Instead of using static protein folding data, our method encodes drug molecular graphs via Graph Isomorphism Networks (GINs) to capture fine-grained topological features. Meanwhile, we optimize protein representations by integrating probabilistic structural priors into a pretrained language model, which effectively simulates thermodynamic conformational flexibility without relying on explicit 3D structural data. A bidirectional cross-attention mechanism is then used to dynamically align these heterogeneous feature modalities. Comprehensive evaluations on the Davis and KIBA benchmark datasets show that StructuraDTA stably outperforms state of-the-art comparison methods. Importantly, the model exhibits strong robustness in cold-start scenarios, and can accurately predict binding affinities for previously unseen drugs and targets. By retaining the predictive performance of structure based models while maintaining the high inference efficiency of sequence-based methods, we provide an accurate and scalable solution to accelerate genome-scale drug discovery research.
Junlin Xu, Ye Yuan, Menglong Hu et al.· IEEE journal of biomedical a...· 0 citations
This review systematically examines the key methodological innovations, including peptide representation learning, multi-modal fusion strategies, multi-label learning paradigms, and emerging predictive frameworks empowered by deep neural architectures and ProtLM-based embeddings, and summarizes the practical applications of these models in peptide database mining, functional mechanism interpretation, and mutation effect prediction.
This work proposes a DTA prediction method, PSDTA, which integrates physicochemical properties into the initial feature representations and explicitly incorporates structural information on amino acids, thereby avoiding the risk of information leakage caused by directly using coordinates as features and enhancing the model's generalization capability.
Shuang Wang, Mao Li, Peifu Han et al.· Journal of Chemical Informat...· 0 citations