Generative artificial intelligence (AI) holds transformative potential for drug discovery, yet existing architectures typically operate in open loops without experimental feedback. Here we introduce rapid compound directed optimization (RCDO), a closed-loop reinforcement learning framework that accelerates the optimization process by bridging dry-lab computation with wet-lab feedback. RCDO couples a three-dimensional structure-guided generative model with a multi-level reward system updated after each design cycle using experimental measurements from all synthesized compounds, including inactive or developability-failed compounds. By continuously aligning the generative model with accumulated wet-lab measurements, RCDO substantially compresses optimization timelines. We evaluated RCDO through retrospective benchmarking against historical optimization trajectories and prospective wet-lab campaigns targeting ROR1, NLRP3, and NSD3. Across prospective evaluations, RCDO rapidly resolved key optimization bottlenecks within two to three design cycles: improving the oral exposure of an ROR1 inhibitor by 40-fold while maintaining antitumor efficacy, reducing CYP2C19 inhibition of an NLRP3 antagonist by 20-fold while preserving inflammasome activity, and boosting the binding affinity of an NSD3 hit by 18-fold. By directly coupling wet-lab feedback to generative learning, RCDO establishes an efficient platform for compound directed optimization, transforming AI-driven drug discovery from static generation into continuous experimental adaptation.
Lead optimization in drug discovery requires iteratively refining molecular candidates while preserving structural similarity to the original compound. Since each evaluation is costly, sample efficiency, the ability to achieve strong performance with limited oracle calls, becomes critical. Existing methods, from geneti...
Ziqing Wang, Yibo Wen, William Pattie et al.· Proceedings of the 32nd ACM...· 0 citations
Abstract As a critical step in drug discovery, molecule optimization aims to improve specific properties of lead compounds. Inspired by isosteres (i.e. molecular substructures with similar reactive electron shells), existing AI-based methods have been proposed for molecule optimization. However, they often rely on avai...
Zhen-Yi Wu, Peng-Cheng Zhao, Jia-Ning Li et al.· Briefings in Bioinformatics· 0 citations
Most computationally predicted materials are never synthesized because conventional synthesis optimization is slow, expertise-dependent, and iterative. Here we present a closed-loop framework that automates this expert workflow by placing human tacit knowledge in the loop through a large language model (LLM) that disti...
Fang Sheng, Steven B. Torrisi, Amanda A. Volk et al.· 0 citations
This work introduces Elite-Weighted Supervised Fine-tuning (EW-SFT), which uses reward to guide elite selection of high-scoring molecules, and updates the model by its own pretraining loss on that set, and consistently outperforms the corresponding native optimizers.
Shiyun Wa, Yifei Wang, A. G. Green et al.· 0 citations
InfRL (Inference-time Reinforcement Learning) offers a practical and domain-agnostic approach to harness reinforcement learning during inference, bridging the gap between static prompting and computationally intensive parameter-level fine-tuning.
Sikun Guo, Amir Hassan Shariatmadari, Jiuqi Wang et al.· Proceedings of the 32nd ACM...· 0 citations
Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the s...
Hong-Yi He, Zheng-Wen Lin, Xiao Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.