Enzymes frequently exhibit promiscuous activity beyond their native roles, providing starting-points for new functions. Finding these promiscuous enzymes, especially for non-native chemical transformations, is challenging but highly valuable, as they promise novel, sustainable solutions for chemistry and biotechnology. However, current machine learning methods are poorly suited to discovering unseen chemistry as they often frame function prediction as closed set classification or a retrieval task. Here, we present Fluxion, a generative deep learning framework that learns enzymatic catalysis by modeling dynamic electron flow trajectories across the enzyme’s catalytic residues. By combining both synthetic chemistry and biochemical datasets with protein language model representations, Fluxion generates multi-step electron-flow trajectories analogous to the arrow-pushing representations used to describe enzyme reaction mechanisms. Generation is conditioned on enzyme context, including the enzyme sequence, catalytic residues, substrates, and cofactors. We show that this conditioning allows Fluxion to learn enzyme-dependent regioselectivity across cytochrome P450 enzymes with different sequences shifting the predicted reaction sites for the same substrate. We then demonstrate that Fluxion’s embeddings are useful for downstream tasks, such as specificity prediction on two experimental datasets, with and without finetuning. Finally, we show that Fluxion has the potential to transfer synthetic chemical logic to biology; it can generate the observed non-native product from real-world non-native directed evolution screens. Our results establish a proof of concept that generative modeling through mechanistic representations of enzymes can shift enzyme function prediction beyond static database retrieval and closed set classification to function generation. This conceptual framework provides a stepping stone towards an in silico generative method to discover non-native biocatalysts.
W. J. Rieger, Sebastian Häussermann, Luca C. Herrmann et al.· bioRxiv· 0 citations
Enzymes present a sustainable alternative to traditional chemical industries, drug synthesis, and bioremediation applications. Because catalytic residues are the key amino acids that drive enzyme function, their accurate prediction facilitates enzyme function prediction. Sequence similarity-based approaches such as BLAST are fast but require previously annotated homologues. Machine-learning (ML) approaches aim to overcome this limitation; however, current gold-standard ML-based methods require high-quality 3D structures limiting their application to large datasets. To address these challenges, we developed Squidly, a sequence-only tool that leverages contrastive representation learning with a biology-informed, rationally designed pairing scheme to distinguish catalytic from non-catalytic residues using per-token Protein Language Model embeddings. Squidly surpasses state-of-the-art ML annotation methods in catalytic residue prediction while remaining sufficiently fast to enable wide-scale screening of databases. We ensemble Squidly with BLAST to provide an efficient tool that annotates catalytic residues with high precision and recall for both in- and out-of-distribution sequences.
W. J. Rieger, Mikael Bodén, Frances H. Arnold et al.· eLife· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.