Summary Embeddings, numerical vectors learned by deep-learning models, are increasingly used to represent complex molecular biology data and support predictive tasks and generative design. There is a growing need for systematic approaches to interpret and explain the information encoded in high-dimensional embedding spaces. Here, we introduce EmmaEmb, a quantitative, model-agnostic framework for geometric correction, direct analysis, and comparison of embedding spaces. Our framework encompasses local and global analysis methods to quantify data distribution within an embedding space and enable comparisons of representations across spaces in relation to known biological features. Through experiments with seven embedding models across six molecular biology tasks, we demonstrate that our methods reveal insights from embedding spaces that align with downstream predictive tasks, uncover misclassification patterns, and contextualize differences in biological information captured by ProtT5, AlphaFold2, and ESM C. We provide an open-source Python library implementing all analysis methods and a guided diagnostic workflow.
Pia Francesca Rissom, Vít Škrhák, Paulo Yanez Sarmiento et al.· Patterns· 0 citations
This study reveals how the organization of the protein energy landscape shapes universal "allosteric grammar" and algorithmic detectability of regulatory binding sites and proposes that allosteric sites are encoded in persistent neutrally frustrated regions optimized for context‐dependent regulatory modulation.
Will Gatlin, Max Ludwick, L. Turano et al.· Protein Science· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.