Automated GenePy gene-burden computation via a reproducible Nextflow workflow integrated with the Genomics England (GEL) Lifebit platform
Abstract
Abstract Interpretation of rare-disease genomes remains constrained by variant-centric analytical frameworks that insufficiently capture the cumulative impact of multiple variants within a gene. GenePy provides an individual-level, gene-based burden metric that integrates variant consequence, allele frequency, and zygosity into a unified quantitative score, enabling a transition from discrete variant annotation to aggregated gene-level interpretation. In the context of Genomics England (GEL), this formulation supports a panel-agnostic, genotype-to-phenotype diagnostic strategy for unresolved monogenic disorders by prioritizing genes with elevated mutational burden per individual. Here, we present a fully automated, containerized GenePy workflow deployed through Nextflow and integrated within the GEL Research Environment via the Lifebit CloudOS platform. This implementation provides scalable, secure, and governance-compliant computation of gene-level burden scores across population-scale cohorts. The workflow harmonizes variant annotation, quality control, and chunked data aggregation within modular, reproducible processes designed for high-throughput execution on cloud-native infrastructure. By enabling robust, portable, and auditable gene-level scoring across large rare-disease sequencing datasets, this framework enhances analytical resolution and supports downstream statistical prioritization, integrative phenotype matching, and hypothesis generation within genotype-to-phenotype diagnostic workflows.