Hierarchy of grammar rules for the language of transcriptional activation domains
Summary Transcriptional activation domains (ADs) of eukaryotic gene activators have remained enigmatic for decades as short, consensus-less, extremely variable amino acid sequences that lack a specific structure and interact fuzzily with an uncertain number of targets. Understanding AD sequence grammar is critical for solving the enigma. Using rational design of AD sequences and high-throughput in vivo experimentation combined with bioinformatic analysis and machine learning, we refined grammar rules for AD sequences, calculated the relative importance of each rule to define a rule hierarchy, and linked each rule to the biochemical features essential for biological function. The key dominating and common features—redundant representation in the sequence of aromatic residues and the obligatory net negative charge of the sequence are consistent with the novel idea that ADs function as acidic-hydrophobic surfactants, which is crucial for understanding eukaryotic gene regulation and the function of intrinsically disordered protein regions.