Aug 2026· Bioinform.· Vol 42· 0 citations· 35 references
Computer ScienceMedicine
TL;DR
DeepGeSeq is introduced, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis that bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development of deep learning methods in genomics research.
Abstract
Abstract Motivation Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. Results By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq’s versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. Availability and implementation https://github.com/JiaqiLi1024/DeepGeSeq.
This work presents an empirical analysis of UQ in deep learning models, focusing on genomics applications, and shows that Bayesian Neural Networks are better at capturing uncertainty caused by strong class imbalance and out-of-distribution data in genomics, despite their computational disadvantages.
This work introduces seqme, a modular and highly extendable open-source Python library, containing model-agnostic metrics for evaluating computational methods for biological sequence design, and can be used to evaluate both one-shot generation and iterative optimization.
Rasmus Møller-Larsen, Adam Izdebski, Jan Olszewski et al.· Bioinformatics Advances· 2 citations
This work establishes a controlled, extensible platform for benchmarking GRN reconstruction and single-cell analysis methods and introduces a flexible GRN generation script, allowing users to design or perturb regulatory networks to test specific hypotheses.
This tutorial provides a comprehensive, end-to-end workflow from raw data to deployed models,icitly designed for environmental chemists with limited prior experience in ML modeling while also providing practical guidance for other users seeking to strengthen their modeling workflows.
Kai Zhang, Yu-Shu Cheng, Hai-Ping Ai et al.· ACS Environmental Au· 0 citations
Deep learning models can predict the functional impact of genetic variants. However, training and applying these sequence models to personalized genome sequences are severely constrained by the computational burden of storing and reading massive datasets.
We developed GenVarLoader, an accelerated dataloader that creates personalized genomic sequences and functional tracks on-the-fly to overcome these bottlenecks. GenVarLoader stores this personalized genomic data in formats specifically optimized for machine learning. Our tool reads variants up to 26,000 times faster than a pre-subset BCF, loads personalized data up to 130 times faster than tuned FASTA and pyBigWig pipelines, and reduces storage approximately 2,000-fold compared to existing alternatives.
By eliminating severe data loading bottlenecks, GenVarLoader ensures efficient model training and inference. It provides a scalable solution for integrating biobank-scale personalized genomic data with deep learning applications.
David Laub, Aaron Ho, Loukik Raina et al.· BMC Genomics· 0 citations
iDCF (Interpretable Deconvolution of Cell Fractions) is a novel framework that enforces biological topology onto deep neural networks, bridging the gap between computational inference and biological intuition.
Hongming Guo, Ting-Fang Wu, Wen-Zheng Wang et al.· PLoS Computational Biology· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.