Skip to content
Preprint

Molexar: A Unified Multimodal Molecular Foundation Model for Drug Design

Jun 2026 · 0 citations · 70 references
Biology

TL;DR

The results establish Molexar as a practical unified foundation for computational chemistry and drug-design workflows and establish Molexar as a practical unified foundation for computational chemistry and drug-design workflows.

Abstract

Molecular generation is a central challenge in drug discovery, requiring models that explore vast chemical space while satisfying diverse design constraints. We present Molexar, a unified multimodal molecular foundation model built on Fragment-SELFIES, a robust, fragment-aware molecular language with validity-preserving decoding and explicit fragment structure. A pretrained autoregressive decoder learns the Fragment-SELFIES syntax and molecular distribution; supervised fine-tuning (SFT) then trains the same decoder on condition-molecule pairs spanning scalar molecular properties, pharmacophore fingerprints, protein sequences, and binding pockets, injecting each condition by in-place replacement of value-token embeddings so that all generation modes share one autoregressive path. Molexar achieves strong efficiency at a small parameter count while matching or exceeding larger models. The pretrained model reaches 100% validity and high drug-likeness in unconditional and fragment-constrained generation; the SFT model follows single- and multi-property instructions and remains competitive on target-conditioned generation on the CrossDocked2020 test set. On MolGenBench, Molexar further generates molecules with favorable safety and potency. These results establish Molexar as a practical unified foundation for computational chemistry and drug-design workflows.

View source

Similar papers

Open access Aug 2026

PockLigGPT: Pocket-Sequence-Conditioned Molecular Generation with GPTs and RL

PockLigGPT achieves competitive docking-oriented performance under a standardized evaluation protocol while maintaining chemical plausibility, favorable physicochemical profiles, and Lipinski-based drug-likeness.

Pablo Varas Pardo, Guillermo Marcos-Ayuso, Eugenia Ulzurrun et al. · 0 citations
Jun 2026

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

A new baseline method is developed identifying the critical components to solve the NMO tasks, including a novel representation for modeling structural constraints and a domain-agnostic pretraining strategy to eliminate pharmaceutical dataset bias.

Matthias Blaschke, Daniel Kienzle, Zsuzsanna Koczor-Benda et al. · 0 citations
Open access Jul 2026

Multimodal feature fusion for molecular property classification.

This study provides a large-scale empirical evaluation of multimodal feature fusion for molecular property classification by systematically integrating SMILES-based chemical language representations with fingerprint-based structural descriptors across 60 benchmark datasets.

Jing Liu, Li Xue, Yin Wang et al. · 0 citations
Preprint Jul 2026

Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

Vilya-1 is introduced, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability.

Vilya Research Pascal Sturmfels, M. Salem, Naozumi Hiranuma et al. · 1 citation
Open access Jul 2026

StructureSAFE: A structure-aware chemical language model for unified hit identification and lead optimization

Structure-based generative models (SBGMs) hold great promises for accelerating drug discovery by enabling target-aware molecular design. However, existing approaches face fundamental challenges: three-dimensional graph-based models can explicitly incorporate protein structural information but often generate chemically implausible molecules due to limited training data, while chemical language models (CLMs) produce chemically plausible molecules but struggle to effectively leverage three-dimensional structural information for structure-conditioned generation and hard to incorporate lead optimization functionality due to the nature of SMILES string. Here, we present StructureSAFE, a structure-aware chemical language model that resolves this trade-off by integrating protein structural and evolutionary encoders with the SAFE molecular representation via pretraining and finetuning training scheme, enabling both de novo hit identification and a comprehensive suite of lead optimization subtasks within a unified framework. Comprehensive benchmarking on the MolGenBench dataset demonstrates that StructureSAFE achieves state-of-the-art (SOTA) performance across multiple metrics, with particularly pronounced improvements in chemical plausibility relative to graph-based models lacking pretraining. Evaluation on a rigorously constructed held-out test set further confirms its ability to generate drug-like, synthetically accessible molecules with competitive predicted binding affinities for previously unseen targets on both hit identification and lead optimization setting. In silico case studies across four therapeutically relevant targets validate its capacity to generate chemically plausible molecules that recapitulate key binding interactions of known high-affinity ligands while proposing novel interactions for potential better affinity and exploring previously unknown regions of chemical space. Taking together, StructureSAFE represents a versatile and practical tool to provide high-quality candidate molecules for augmenting medicinal chemistry workflows in both hit identification and lead optimization campaigns.

Bo Yang, Ke Xu, Chijian Xiang et al. · 0 citations
Review Jul 2026

Beyond SBDD: Geometric Deep Learning in Polypharmacology and Multi-target Drug Design

This review elucidates the paradigm shift in drug discovery from serendipitous exploration to rational, structure-driven polypharmacological molecular engineering, thereby providing a clear, structured guide for navigating the complexities of next-generation therapeutics.

Tianming Han, Zhijie Pan, Wenchi Ge et al. · 0 citations