Single-cell foundation models have transformed transcriptomic analysis, yet most rely on fixed gene identifiers that limit transfer across species and data types. Here we present LucaCell, a sequence-centric foundation model that represents genes through pre-trained mRNA sequence embeddings rather than static gene annotations. Gene expression is discretized into bins and modeled with a Transformer encoder, enabling sequence-informed cell representation without a fixed gene-ID vocabulary. Pre-training on 85 million human and mouse single cells, LucaCell is evaluated on human, mouse and lemur gene expression profiles, human chromatin accessibility data, unaligned reads from more than 50 prokaryotic taxa, and five influenza A virus genomes. LucaCell enables manual-mapping-free cross-species cell type annotation and an alignment-free microbial embedding framework that simultaneously distinguishes bacterial species identity and intra-species physiological states. It also improves gene expression reconstruction by incorporating donor-specific exonic SNP information into mRNA sequence embeddings, and predicts cellular viral load across influenza A virus strains while highlighting infection-like transcriptional states in mock-infected cells. These results show that sequence-informed gene representation can improve the generalization of single-cell foundation models across species, data types, and predictive tasks.
Yan Sun, Yong He, Min-Si Ren et al.· bioRxiv· 0 citations
BACKGROUND
Cystic echinococcosis, caused by the tapeworm Echinococcus granulosus sensu stricto (ss), is a globally distributed, zoonotic disease that is recognised by WHO as a neglected tropical disease. Despite its clinical and economic importance, nuclear genomic variation in this parasite has not been systematically characterised across global populations. In this study, we aimed to characterise the genome-wide nuclear genetic diversity and population structure of E granulosus ss across globally distributed populations.
METHODS
We conducted a genomic study of 137 E granulosus ss samples from endemic regions across five continents, derived from previously collected parasite material from livestock, wildlife, and human infections. Using a chromosome-scale reference genome, we applied population genomic approaches to investigate genome-wide nuclear genetic diversity, population structure, and patterns of evolutionary constraint.
FINDINGS
We identified 1 071 085 nuclear single-nucleotide polymorphisms across 137 samples, with heterozygosity ranging from 46% to 93% per sample. Genome-wide analyses identified two major clades associated with geographical origin. Distinct regions of genetic differentiation were observed, particularly on chromosome 9. Conserved genes under purifying selection included those involved in glycan biosynthesis and core cellular functions, whereas variable genes were enriched in pathways such as ribosome biogenesis. Mitochondrial genotypes (G1 and G3) did not align with the nuclear genomic structure.
INTERPRETATION
To the best of our knowledge, this study provides the first broad atlas of nuclear genomic diversity in E granulosus ss, uncovering genetic diversity and population structure. The findings have important implications for molecular epidemiology, genomic surveillance, and translational development of diagnostics and vaccines. Incorporating genomic data into cystic echinococcosis control programmes could enhance WHO-aligned efforts to reduce the burden of this neglected tropical disease.
FUNDING
Australian Research Council and the Estonian Ministry of Education and Research.
Liina Anijalg, Pasi K. Korhonen, Neil D. Young et al.· The Lancet Microbe· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.