RegimeFormer: A Large Protein Model of Global Perturbation Regimes
Abstract
Protein language models organize sequence and structure at scale; here we add a global coordinate system for how proteins respond to mutation.1–12 We present RegimeFormer, a large protein perturbation model coupled to RegimeAtlas, which we constructed by harmonizing, representing and indexing 202,556,313 non-redundant protein sequences across the tree of life. This full sequence universe defines the atlas-wide protein-regime map. A diversity-preserving one-million-protein subset provides the high-resolution training and inference layer, with 995,995 proteins yielding 407,048,356 residue summaries and substitution-specific predictions available on demand. Across experimental DMS, molecular benchmarks, structural confidence and evolutionary constraint, RegimeFormer reveals reproducible protein-level perturbation regimes that organize residue fragility, adaptability and uncertainty. Explicit regime conditioning improves substitution-specific reconstruction and shows its largest relative advantage under unseen-protein, unseen-family and low-homology evaluation. The resulting molecular priors also improve downstream transcriptomic and drug-response modelling. Together, RegimeFormer and the 202.6-million-sequence RegimeAtlas establish a large-model framework for mapping, predicting and querying protein perturbation landscapes across global sequence space.