POLO: Preference-Guided Multi-Turn Reinforcement Learning for Sample-Efficient Lead Optimization
Lead optimization in drug discovery requires iteratively refining molecular candidates while preserving structural similarity to the original compound. Since each evaluation is costly, sample efficiency, the ability to achieve strong performance with limited oracle calls, becomes critical. Existing methods, from geneti...