Skip to content
Preprint

Vilya-1: An all-atom foundation model for macrocycle structure prediction and design

Jul 2026 · 1 citation · 61 references
Computer Science Biology

TL;DR

Vilya-1 is introduced, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability.

Abstract

Macrocyclic peptides are an increasingly important therapeutic modality, but existing computational methods for modeling their structures and properties are limited in scope and do not generalize well across the synthetically accessible chemical space. In this work, we introduce Vilya-1, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability. Vilya-1 operates on a uniform all-atom representation and is trained on heterogeneous structural datasets spanning diverse topologies and chemical classes. Across a broad set of macrocycles composed of canonical and non-canonical residues, Vilya-1 substantially improves geometric accuracy relative to physics-based methods, co-folding networks, and deep-learning conformer generators, while maintaining broad chemical coverage that extends to small molecules. Vilya-1 also supports generative applications, enabling the design of novel macrocycles with tailored chemical, structural, and property profiles. Together, these capabilities establish Vilya-1 as a foundation model for accelerating the development of next-generation macrocycle therapeutics.

View source

Similar papers

Preprint Jul 2026

Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2

Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.

Vilya Research Pascal Sturmfels, Naozumi Hiranuma, M. Salem et al. · 0 citations
Open access Jul 2026

HELM-BERT: Topology-Aware Representations for Chemically Modified Peptides

Chemically modified and macrocyclic peptides are increasingly important therapeutics, yet current molecular representation models do not natively represent chemical modification and covalent topology in a unified way. Atom-level strings obscure macrocyclic connectivity, whereas protein sequence models cannot encode noncanonical residues and explicit cross-links. Here we pretrain an encoder-only transformer directly on Hierarchical Editing Language for Macromolecules (HELM) notation, which specifies monomer identity and connectivity. In this work, we show that the resulting representations achieve best mean performance in cyclic peptide membrane permeability prediction (random split R 2 = 0.668; retaining best mean performance under a Murcko scaffold split), exceeding external pretrained SMILES-based encoders. An architecture-matched SMILES control narrowed the HELM–SMILES gap under full finetuning, whereas HELM-BERT retained clearer advantages in frozen-representation settings. HELM-BERT also preserves HELM-specified macrocyclic topology in a linearly accessible form and supports competitive peptide–protein interaction prediction across complementary Propedia and ChEMBL benchmarks. More broadly, these findings suggest that pretraining directly on notation that makes structural constraints explicit offers a transferable strategy for biomolecular modalities that fall between small-molecule chemistry and protein sequence.

Seungeon Lee, Takuto Koyama, Itsuki Maeda et al. · 0 citations
Open access Jul 2026

Mavchen-1: A Conformational Ensemble Platform for Protein–Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.

Ryan Varghese, Pooja Tiwary, Krishil Oswal · 0 citations
Jul 2026

Derisking Affinity Optimization for Macrocycles and Cyclic Peptides: High-Precision Free Energy Simulations across Five Diverse Targets.

Macrocycles and cyclic peptides represent a compelling therapeutic modality for engaging historically challenging biological targets, yet their inherent structural complexity and laborious synthetic demands pose significant hurdles in drug discovery. Given the high resource investment required for macrocyclic synthesis, the integration of rigorous, physics-based computational methods provides a critical advantage for accurately prioritizing design candidates. While Free Energy Perturbation (FEP) has emerged as a transformative method for affinity prediction, its application to complex, beyond-rule-of-five (bRo5) macrocycles demands specialized enhanced sampling protocols and careful consideration of receptor conformational states to effectively navigate their multidimensional conformational landscapes. Here, we present a retrospective validation of the FEP+ framework across five diverse macrocyclic and cyclic peptide inhibitor series: KRAS, PCSK9, MCL-1, JAK2, and Cyclin A/B. Encompassing over 230 unique peptidic and nonpeptidic analogues, our analysis of this consolidated data set demonstrates robust predictive accuracy (global pairwise RMSEΔΔG = 1.06 kcal/mol) and reveals critical insights into the complex interplay between ligand preorganization, hydration dynamics, and binding energetics. The results demonstrate reliable intraseries rank-ordering and robust absolute accuracy across an experimental dynamic range exceeding 10 kcal/mol (>7 orders of magnitude in binding affinity), establishing FEP+ as an effective computational method for derisking and accelerating the discovery of clinically viable bRo5 therapeutics.

Ernest Awoonor-Williams, Alexandre Beautrait, Loukas Petridis et al. · 0 citations
Review Open access Aug 2026

Diffusion-Based Protein Structure Design: Geometric Modelling, Validation Strategies, and Thermodynamic Challenges

This review focuses on coordinate- and residue-frame-based diffusion approaches for generating protein structures, paying particular attention to geometric equivariance, conditioning strategies, all-atom modelling and interaction-aware design.

Wen-Ran Li, Xavier F. Cadet, David Medina-Ortiz et al. · 0 citations
Open access Aug 2026

Structure-free, site-resolved contrastive learning extends small-molecule discovery beyond the reach of structure-based modeling

Virtual screening asks which molecules, among an enormous space of drug-like chemistry, are worth synthesizing and testing against a protein target. Most modern methods answer this question by building and scoring an explicit three-dimensional pose through molecular docking, or the co-folding models that now approach experimental accuracy. Building these poses presumes a well-defined pocket. However, the non-orthosteric, cryptic, and intrinsically disordered sites where unexplored ligandability lies offer no such pocket to explore. Here we present Ptarmigan-1, a contrastive model that co-embeds the residues of a protein with candidate small molecules in a shared latent space, from sequence and two-dimensional chemistry alone, and without ever constructing a pose. Freed from the requirement for protein structures, Ptarmigan-1 trains directly on chemoproteomic and bioactivity data of mixed resolution, scores a compound in ten milliseconds rather than the tens of seconds a co-folding model demands, and resolves each prediction to the residues a compound engages. On well-folded, orthosteric targets it performs comparably to a collection of co-folding and docking models, and on covalent, cryptic, and disordered sites it matches or exceeds them. Ptarmigan-1 localizes reversible and covalent inhibitors to the pockets they engage, even for targets withheld from training, and screens the entire human proteome against a library of 3.4 billion compounds in under a day. By decoupling molecular recognition from structure, Ptarmigan-1 recasts virtual screening as a nearest-neighbor query in a rich latent space shared by protein residues and the compounds that bind them.

William E. Fondrie, Daniele Canzani, Lillian Tatka et al. · 1 citation