Similar papers
Machine Learning-Assisted Evolution of Broadly Functional Enzyme Libraries
Results indicate that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope, and suggest that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope.
Combining Stability-Centered Atomistic Design with Machine Learning for Targeted Enzyme Optimization
A machine-learning-assisted enzyme-engineering (MLEE) workflow that adds substrate-specific functional information to htFuncLib through an initial screening and sequencing round that may bypass the need for transition-state models and reduce the effort required for obtaining high-activity variants.
An enzyme-specific protein language model for catalytic property prediction
Enzymes drive cellular metabolism, yet predicting catalytic properties from amino acid sequences remains challenging. Existing protein language models (PLMs) provide powerful general-purpose representations but are often inefficient for high-throughput screening and insufficiently adapted to enzyme-specific tasks. Here, we propose EnzGFM, an enzyme-specific PLM based on a Mamba-Transformer hybrid architecture with hierarchical pre-training to capture enzyme-specific patterns. Across enzyme property prediction benchmarks, EnzGFM consistently outperforms Transformer-based PLMs with 2–5-fold acceleration, achieving relative improvements of 16.67% in kinetic parameter prediction, 15.69% in enzyme-reaction mapping, 13.19% in EC number classification, and 20.04% in mutation effect assessment. Building on EnzGFM, we develop EnzGFM-Agent, an enzyme-focused agentic pipeline. Experimental validation further suggests that EnzGFM-Agent can enrich beneficial variants within small candidate pools. Together, these results demonstrate that EnzGFM captures enzyme-specific sequence-function patterns, while EnzGFM-Agent translates these predictions into experimentally actionable candidates and can help reduce wet-lab screening burden for practical enzyme engineering. Enzyme function prediction from amino acid sequences remains a central challenge in computational biology, despite recent advances in protein language models. This manuscript introduces EnzGFM, an enzyme-specific hybrid model that improves both accuracy and efficiency across multiple prediction tasks and, together with the EnzGFM-Agent pipeline, demonstrates the ability to identify experimentally validated beneficial variants while reducing screening effort.
Enzyme Engineering: From Classical Strategies to AI-Driven Biocatalyst Design.
This review examines enzyme engineering from classical methods to AI-assisted biocatalyst development, highlighting key advances, challenges, and emerging trends in autonomous laboratories, sustainable biocatalysis, and computational protein design.
Coevolution-informed Bayesian optimization for sample-efficient protein design
This work introduces ALSEBO (Active Learning Sequence Exploration via Bayesian Optimization), which couples a generative latent sequence landscape to Bayesian optimization and featurizes candidates with direct-coupling-analysis (DCA) coevolutionary statistics.
AUKAT: Conditional VAE-Driven Augmentation and Neural Modeling of Enzyme Turnover Numbers
Accurate prediction of enzyme turnover numbers (kcat) is essential for applications in systems biology, metabolic engineering, and drug discovery, yet remains challenging due to the limited availability and uneven distribution of experimental data. Here, we present AUKAT, an integrated framework that combines conditional generative modeling with deep neural prediction to improve kcat estimation. A conditional variational autoencoder generates synthetic training instances in embedding space, followed by a selection pipeline that retains samples with strong agreement across independent evaluators, thereby ensuring data reliability. A hybrid convolutional neural network and transformer-based architecture is then used to predict kcat from substrate, enzyme functional, and species embeddings. Incorporating synthetic data improved predictive performance for both random forest and neural network models in five-fold cross-validation, with larger gains observed for the neural network architecture. Benchmarking against DLKcat demonstrated comparable predictive accuracy on the standard test set, while evaluation on stricter unseen subsets indicated improved generalization for low-similarity substrates and enzymes. Feature importance analysis further showed that AUKAT leverages substrate, enzyme functional, and species information in a more balanced manner rather than relying predominantly on a single feature source. In addition, AUKAT-human, a specialized model trained using a pre-training and fine-tuning strategy, achieved improved prediction accuracy for human enzyme kinetics. Overall, AUKAT provides a scalable approach for enzyme kinetics prediction and offers a practical solution to data scarcity in biochemical modeling.