This project presents an AI-guided simulated annealing framework for automated gene sequence design using two approaches, combining a fast native optimisation core with an adaptive and context aware evaluation model, demonstrating a flexible approach to large scale gene sequence optimisation.
Abstract
Designing effective gene and mRNA sequences is a difficult optimisation problem because the number of possible nucleotide combinations grows extremely quickly with sequence length. Traditional optimisation methods such as simulated annealing are well suited to exploring these large search spaces, but their performance depends heavily on the quality of the scoring function used to evaluate candidate sequences. Hand crafted scoring rules are often slow to compute and cannot easily adapt to different biological contexts or patient specific constraints. This project presents an AI-guided simulated annealing framework for automated gene sequence design using two approaches. The first replaces fixed rule-based scoring with an adaptive model evaluating candidates using biological reference data and patient-specific information. By adjusting biological trait importance based on age, disease background, and treatment goals, the scoring model dynamically changes sequence evaluation without modifying the optimization algorithm. The second approach employs Gradient Boosting Regression on CRISPR guide RNA sequences with extracted biological features including GC content, positional nucleotides, and sequence complexity metrics. This model learns from validated literature guides, providing interpretable, deterministic scoring while maintaining adaptability. The framework is designed to support long running and repeated simulated annealing searches with minimal human intervention. Sequence evaluation is decoupled from the optimisation engine so that scoring models and reference databases can be updated as new experimental or clinical data becomes available. This allows the same optimisation pipeline to be reused across different applications such as vaccine design, cancer related gene targets or personalised therapies. By combining a fast native optimisation core with an adaptive and context aware evaluation model, this work demonstrates a flexible approach to large scale gene sequence optimisation. The proposed system highlights how AI driven scoring can improve the practicality of heuristic search methods and move sequence design closer to personalised and data driven biomedical applications.
A language-model-guided framework that iteratively refined industry-optimized coding sequences of clinical-stage therapeutics through synonymous exploration of codon space establishes directed evolution as a practical strategy to improve biologic expression, a key manufacturing bottleneck, without altering protein sequ...
James Heuschkel, Laura Kingsley, Jon Reed et al.· bioRxiv· 0 citations
This work uses a general-purpose "post-training" algorithm grounded in statistical physics that employs quantitative experimental rankings to directly produce a sampler for diverse, high fitness sequences with fewer data points than competing methods.
Sebastian Ibarraran, Shriram Chennakesavalu, Frank Hu et al.· Journal of Chemical Informat...· 0 citations
Directed evolution is a method for engineering biological systems or components, such as proteins, wherein desired traits are optimised through iterative rounds of mutagenesis and selection of fit variants. The process of protein directed evolution can be envisaged as navigation over high-dimensional landscapes with nu...
S. Towers, Jessica James, Harrsion Steel et al.· bioRxiv· 0 citations
This PhD research focuses on the development and analysis of Monte Carlo Tree Search (MCTS) and related stochastic search algorithms for de novo RNA design at the tertiary structure level to explore how such methods can efficiently navigate the high-dimensional sequence space while integrating feedback from modern stru...
This work introduces seqme, a modular and highly extendable open-source Python library, containing model-agnostic metrics for evaluating computational methods for biological sequence design, and can be used to evaluate both one-shot generation and iterative optimization.
Rasmus Møller-Larsen, Adam Izdebski, Jan Olszewski et al.· Bioinformatics Advances· 2 citations
EvoMOBO is established as a modular framework for multi-objective protein engineering using experimental or mechanism-derived labels using simulation-derived mechanistic descriptors, with experiments reserved for final validation.
Kai Wen, Sirui Wang, Yixin Sun et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.