Skip to content
#protein folding Review Open access

Toward a Minimal Amino Acid Alphabet for Protein Design

Aug 2026 · bioRxiv · 0 citations · 47 references
Biology

TL;DR

The results show that globular proteins may have formed early in evolution and show that it is possible to design proteins with interesting properties for biotechnology and synthetic biology.

Abstract

Proteins are built from 20 canonical amino acids. It is interesting to explore whether proteins can be formed from significantly reduced amino acid alphabets. Our bioinformatics survey of UniProt (more than 250 M sequences) revealed that proteins composed of reduced amino acid alphabets (< 10) are extremely rare among existing proteins. Next, we used computational protein design to design proteins composed of all 1,013 possible alphabets of 2-10 early amino acids (Ala, Asp, Glu, Gly, Ile, Leu, Pro, Ser, Thr, and Val). The length of all proteins was 100 amino acid residues. Small amino acid alphabets preferred simple helices or helix bundles. Larger amino acid alphabets allowed for the design of more complex structures. A protein composed of 8 amino acid types (Ala, Asp, Gly, Leu, Val, Ser, Thr, and Pro) was successfully experimentally verified. It adopts the β-sheet-rich fibronectin type III domain architecture. Attempts to experimentally verify designs composed of 6 and 4 amino acid types were unsuccessful. We show by a computational experiment with an experimental validation that inverse folding models, namely ProteinMPNNsol, can stabilize a designed protein within the same eight-amino-acid alphabet. Our results show that globular proteins may have formed early in evolution. Furthermore, we show that it is possible to design proteins with interesting properties for biotechnology and synthetic biology.

Read PDF

Similar papers

Open access Jul 2026

Charged amino acid propensities and solubility: lysine is elevated at the termini of helices in E. coli proteins

An emerging result in the relationship between amino acid sequence and protein solubility is a preference, on average, for lysine over arginine in more soluble proteins. The termini of helices are known to be prone to partial unfolding, often employing N- and C-cap amino acids to maintain stability. Hypothesising that lysine/arginine differences in relation to solubility may be evident at helical termini, their propensities and predicted charge interactions in helices were examined. There is enrichment of lysine over arginine at helical termini in AlphaFold models of Escherichia coli proteins, more so (on average) for the most soluble proteins. Similar effects are seen for the sum of charged amino acids at helical termini. Regions other than helical termini also show correlation of lysine composition, and overall charged amino acid composition, with solubility. These results suggest that protein design protocols could improve solubility through targeting lysine enrichment in regions such as helical termini, in addition to the more conventional consideration of helix capping interactions.

Sifan Zhang, Jim Warwicker · 0 citations
Open access Jul 2026

Expanding all-α-helical protein space through rational computational design

De novo protein design is advancing rapidly1,2. This is being driven by AI to generate protein backbones, sequences, and structural models3–7. As a result, de novo designed proteins are becoming larger and more complex8–10, and increasingly explore new protein structures11,12. By contrast, natural proteins have evolved structural and functional complexity by modular combination of recurring protein domains13. Approximately 25% of these natural domains are mostly α-helical structures14. Here we show how these can be expanded using rational computational design. Following the domain classification scheme CATH15, we build complex all-α de novo proteins hierarchically using sequence-to-structure relationships for helix-helix interactions, systematic rules to connect helices, computational tools to design loops, and in silico evaluation. The pipeline starts with a target architecture of free-standing helices. These are connected into a topology by considering local arrangements of helical bundles using understood sequence-to-structure relationships for helix packing. Single-chain sequences are completed using template- and AI-based methods. Finally, AlphaFold models are assessed to give small numbers of designs for experimental validation. We test 31 designs for 14 different architectures and 25 topologies. 75% of these express as stable, monomeric, water-soluble proteins; and >30% yield X-ray crystal structures matching the designs to atomic accuracy and with new-to-nature structures. Finally, several of the scaffolds are functionalised through one-shot designs to deliver ion, small-molecule and protein binders.

K. I. Albanese, Joel J. Chubb, L. Gutierrez-Rus et al. · 0 citations
Review Aug 2026

Exome-matched protein as the ideal dietary amino acid pattern: A narrative review.

With few exceptions, all organisms on Earth use a common amino acid alphabet for synthesizing protein that consists of 20 canonical amino acids. Heterotrophic organisms have to obtain these amino acids from their food, especially the indispensable amino acids. The amount of and ratio between the indispensable amino acids is therefore a major determinant of protein quality. There exist several measures of protein quality for human nutrition and all of them rely on the empirical estimation of amino acid requirements. With the availability of whole genome sequencing data, however, a new method of deductively deriving amino acid requirements became available which uses the information encoded on the exome (the totality of protein-coding exons). By translating an organism's exome into the corresponding proteome in silico, the average encoded amino acid pattern can be computed and used ex hypothesi for defining an ideal amino acid pattern. Here I review the theoretical concepts behind exome-matched proteins and summarize the preclinical data that provide preliminary evidence for the hypothesis that such proteins indeed constitute an ideal protein source to maximize growth, reproduction and health. While clinical trials are yet to confirm this hypothesis in humans, there is a broad hypothetical range of applications of exome-matched proteins for human consumption. This review concludes that the concept of exome-matched proteins is interesting theoretically, promising for practical applications in animal and human nutrition and stimulating for further transdisciplinary research.

R. Klement · 0 citations
Aug 2026

HighMorph: De Novo Cyclic Peptide Sequence Design via Protein–Protein Interaction Recapitulation

Cyclic peptides have emerged as a compelling class of bioactive scaffolds, but de novo design of target-binding cyclic peptides from protein structures remains challenging. Here, we present HighMorph, an interaction-guided framework that combines protein–protein interaction information with artificial intelligence for rational cyclic peptide design. HighMorph integrates Monte Carlo tree search with a Transformer-based policy-value network to efficiently explore cyclic peptide sequence space, while incorporating explicit atomic-level hydrogen bond constraints extracted from reference protein–protein complexes to guide sequence optimization. The framework is systematically validated on two clinically relevant targets, programmed death-ligand 1 (PD-L1) and kallikrein-related peptidase 4 (KLK4). Notably, 33.3% and 40% of the generated candidates are active against PD-L1 and KLK4, respectively, with active cyclic peptides exhibiting micromolar binding affinities (approximately 10–6 M). These results validate our approach for cyclic peptide design. Additionally, interaction analysis provides insights for developing therapeutics targeting challenging protein interfaces.

Minhui Lan, Chengyun Zhang, Wentong Wang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.