Skip to content
Open access

CrossMol: Cross-Modal Mask-Predict Pre-training For 3D Molecular Data.

Sep 2026 · Bioinformatics · 0 citations
Medicine

Abstract

MOTIVATION Self-supervised pre-training models for molecular data have demonstrated notable results across many downstream tasks. The inherent multimodal properties of molecules have also motivated efforts to capture information from different modalities. However, current multimodal molecular pre-training models usually treat these modalities as equal and independent, despite differences in their information content. Three-dimensional (3D) molecular structures generally contain finer-grained information than the simplified molecular-input line-entry system (SMILES), which primarily captures higher-level semantic information, such as molecular topology. RESULTS We designed a cross-modal mask-predict pre-training model, CrossMol, to capture semantic associations between modalities with unequal information volumes. The model completes missing 3D structure using higher-level semantic information from another modality, such as SMILES. This allows it to learn cross-modal associations and better understand fine-grained 3D structural information. We also introduce a reweighted distance prediction loss to improve the modelling of short-range structural information. Experiments show that CrossMol achieves large performance gains on multiple downstream molecular tasks, attaining state-of-the-art results. AVAILABILITY AND IMPLEMENTATION The source code and training data are available at https://github.com/zhengkangjie/crossmol. The code is archived at https://doi.org/10.5281/zenodo.19653204. SUPPLEMENTARY INFORMATION Supplementary material contains the task-specific hyperparameters.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.