Single-cell foundation models have transformed transcriptomic analysis, yet most rely on fixed gene identifiers that limit transfer across species and data types. Here we present LucaCell, a sequence-centric foundation model that represents genes through pre-trained mRNA sequence embeddings rather than static gene annotations. Gene expression is discretized into bins and modeled with a Transformer encoder, enabling sequence-informed cell representation without a fixed gene-ID vocabulary. Pre-training on 85 million human and mouse single cells, LucaCell is evaluated on human, mouse and lemur gene expression profiles, human chromatin accessibility data, unaligned reads from more than 50 prokaryotic taxa, and five influenza A virus genomes. LucaCell enables manual-mapping-free cross-species cell type annotation and an alignment-free microbial embedding framework that simultaneously distinguishes bacterial species identity and intra-species physiological states. It also improves gene expression reconstruction by incorporating donor-specific exonic SNP information into mRNA sequence embeddings, and predicts cellular viral load across influenza A virus strains while highlighting infection-like transcriptional states in mock-infected cells. These results show that sequence-informed gene representation can improve the generalization of single-cell foundation models across species, data types, and predictive tasks.
Yan Sun, Yong He, Min-Si Ren et al.· bioRxiv· 0 citations
The rapid evolution of Large Language Models (LLMs) has brought unprecedented capabilities across reasoning, coding, and multimodal tasks. However, as performance scales, their opaque ''black-box'' nature raises a critical challenge: How can we trace the origins of emergent intelligence, and more importantly, how can we leverage these internal mechanisms to guide model optimization? This tutorial provides a comprehensive, end-to-end view of LLM interpretability, transitioning from microscopic neural analysis to macroscopic application and deployment. It is systematically organized into five core sections: i) Unlocking the Black Box: We begin with the evolution of LLM interpretability and highlight recent breakthroughs from leading research teams. ii) Methodology: We present a rigorous overview of foundational theories (e.g., mathematical framework for transformer, biological mechanisms in LLMs) and essential methods (e.g., path patching, logit lens, and neuron description). iii) Anatomy of LLMs: Using advanced techniques to decode internal semantic features, neural circuits, and complex behaviors, we interpret how models perform reasoning, factual recall, and in-context learning. iv) Applications: We show how to transfer interpretability insights into actionable improvements across the LLM pipeline, including interpretability-guided data synthesis (data value scoring, corpus filtering, and activation-based data diagnosis). We also present Pinpoint Training and Steering for precise capability gains, and Pinpoint Quantization for extreme low-bit compression with minimal capability loss. v) Advanced Topics: We conclude by exploring how these interpretability paradigms scale and inspire the design of frontier architectures, agentic systems, and thinking models. In this tutorial, researchers and engineers will gain the theoretical frameworks and practical engineering toolkits needed to understand, steer, and efficiently deploy LLMs in real-world production environments.
Wei Zhang, Zhengfu He, Lu-Lu Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
This tutorial provides a comprehensive, end-to-end view of LLM interpretability, transitioning from microscopic neural analysis to macroscopic application and deployment, and explores how these interpretability paradigms scale and inspire the design of frontier architectures, agentic systems, and thinking models.
Wei Zhang, Zhengfu He, Lucia Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.