Aug 2026· The Plant Journal· Vol 127· 0 citations· 76 references
Medicine
TL;DR
Here, BINN is extended for genomic prediction and selection in crops by integrating thousands of single‐nucleotide polymorphisms with multi‐omics measurements and prior biological knowledge and substantially reduces prediction error relative to conventional neural nets and correctly identifies the most important nonlinear pathway.
Abstract
SUMMARY Traditional genotype‐to‐phenotype models depend heavily on direct mappings that achieve only modest accuracy, forcing breeders to conduct large, costly field trials to maintain or marginally improve genetic gain. Models that incorporate intermediate molecular phenotypes can achieve higher predictive fit, but remain impractical since such data are unavailable at deployment or design time. Biology‐informed neural networks (BINNs) overcome this limitation by encoding pathway‐level inductive biases and leveraging multi‐omics data only during training, while using genotype data alone during inference. Here, we extend BINNs for genomic prediction and selection in crops by integrating thousands of single‐nucleotide polymorphisms with multi‐omics measurements and prior biological knowledge. By directly embedding omics‐derived priors, BINN outperforms conventional models in low‐data (n < p) regimes and enables sensitivity analyses that expose biologically meaningful traits. Applied to maize gene expression and multi‐environment field trial data, BINN improves rank correlation accuracy within and across most subpopulations under sparse data conditions and nonlinearly identifies genes that GWAS/transcriptome‐wide association studies may fail to uncover. With complete domain knowledge for a synthetic metabolomics benchmark, BINN substantially reduces prediction error relative to conventional neural nets and correctly identifies the most important nonlinear pathway. Importantly, both cases show that highly sensitive BINN latent variables correlate with the experimental quantities they represent, despite not being trained on them. This suggests that BINNs learn biologically relevant representations, nonlinear or linear, from genotype to phenotype. Together, BINNs establish a framework for improved genomic prediction accuracy and biological discovery that can guide genomic selection, candidate gene selection, pathway enrichment, and gene‐editing prioritization.
Genomic prediction of multiple phenotypes is crucial in modern plant breeding; however, existing methods struggle with negative transfer and lack interpretability, particularly across high‐dimensional small‐sample data and diverse species. To address this, we propose Mul‐PheG2P, a novel paradigm based on decoupled learning and predictive space fusion. It employs a two‐stage design: first training phenotype‐specific encoders using genetic data, then decoupling phenotype‐specific learning from cross‐phenotype aggregation via an interpretable prediction layer. Mul‐PheG2P outperforms existing methods across diverse crop datasets, including maize (Zea mays), wheat (Triticum aestivum), and tomato (Solanum lycopersicum). It provides a multi‐scale interpretability chain: at the macro level, it quantifies phenotypic contributions via attention‐based weighting; at the micro level, Integrated Gradients reveal the genetic basis of predictions. Notably, the model successfully identified the CCT (CONSTANS, CO‐like, and TOC) motif regulating photoperiodism and the SQUAMOSA (SQUAMOSA promoter binding protein) promoter for inflorescence development, confirming its ability to capture functional biological mechanisms. These results highlight the high performance and interpretability of Mul‐PheG2P, showcasing its value for low‐cost, large‐scale screening to advance precision breeding.
Jia-Hui Wang, Yong Zhang, Bo Li et al.· New Phytologist· 0 citations
Crop improvement increasingly depends on extracting useful breeding signals from data that span DNA sequence variation, gene regulation, molecular phenotypes, high-throughput field measurements and environmental exposure. Multi-omics can connect genotype to phenotype through intermediate biological layers, while artificial intelligence (AI) and machine-learning methods can model nonlinear, high-dimensional relationships that are difficult to represent with conventional approaches. Yet greater data volume and model complexity do not automatically translate into greater genetic gain. This critical narrative review evaluates how genomics, pangenomics, transcriptomics, epigenomics, proteomics, metabolomics, phenomics and environmental covariates are being integrated with statistical learning, machine learning and deep learning for crop improvement. Literature was selected from accessible scholarly databases and indexes through 13 June 2026, with emphasis on peer-reviewed studies that permit evaluation of predictive value, biological interpretation and breeding relevance. The evidence is strongest where additional modalities capture non-redundant information that is biologically proximal to the target trait or environment, as demonstrated in hybrid prediction, stress adaptation, grain-quality analysis and environment-aware genomic prediction. Conversely, classical genomic best linear unbiased prediction and related models remain competitive in many settings, particularly when sample size is modest, relationships among individuals dominate prediction, or nonlinear signal is weak. Reported AI advantages are sensitive to validation design, relatedness between training and test sets, tissue and developmental stage, environmental transfer, missing modalities and hyperparameter tuning. Pangenomes, single-cell regulatory maps and interpretable multimodal models broaden the biological search space, but evidence for routine breeding utility remains less mature than their mechanistic promise. The most defensible path forward is therefore not unrestricted model escalation, but decision-focused integration: biologically informed feature representation, prospective multi-environment validation, explicit uncertainty, robust missing-data handling and functional validation of discovered mechanisms. Multi-omics and AI are most likely to accelerate crop improvement when evaluated against breeding decisions and realised genetic gain rather than prediction accuracy alone.
A comprehensive overview of GS methodologies is provided, first covering the statistical foundations of linear mixed and Bayesian models, and then modern ML and DL approaches, to provide practical guidance for optimizing genomic evaluation strategies in the era of big data breeding.
Li-Fei Zhang, Mingzhu Zhang, Xinle Wang et al.· Journal of Animal Science an...· 1 citation
Genetic prediction of complex phenotypes typically relies on additive linear models, which scale well but cannot capture non-additive effects or deeply integrate molecular and clinical data. Domain-specific neural networks have driven advances in images, text, and other modalities, but genome-scale neural networks remain challenging because genotypes are sparse and high-dimensional, effective sample sizes are limited, and generic architectures lack interpretability. Here, we introduce the omnigenic neural network, a biologically structured architecture inspired by the omnigenic model of complex traits. The model learns hierarchical representations of biological processes, accommodates multimodal inputs, supports transfer learning, and enables multitask prediction. Models trained in the UK Biobank and evaluated in the All of Us cohort for ischemic heart disease, type 2 diabetes, and schizophrenia outperformed published PGS Catalog and PRS-CSx scores. A multitask model trained across 36 cardiovascular endpoints further outperformed corresponding single-phenotype models and baselines. The architecture provides systems-level interpretability by quantifying the contributions of biological processes, which were consistent with established disease mechanisms. It also captures non-linear interactions between variants. Analysis of these interactions using Integrated Hessians revealed patterns concordant with previously reported epistatic associations. Together, these findings establish the omnigenic neural network as a flexible framework for interpretable, multimodal, and multitask genomic prediction.
J. Upmeier zu Belzen, L. Arnoldt, N. Hollmann et al.· medRxiv· 0 citations
ABSTRACT Integrating proteomic and metabolomic data is essential for understanding complex diseases, yet current approaches that rely primarily on statistical associations often overlook the structured biochemical relationships between molecular entities and suffer from discriminative instability in small clinical cohorts. Here, we present ProMetNet, a biochemically constrained framework that incorporates pathway‐derived connectivity from the Reactome database into neural network architecture. By encoding protein–metabolite relationships based on reaction topology, ProMetNet models structured cross‐omics dependencies rather than relying solely on statistical correlations, reducing spurious associations while preserving global molecular context and improving robustness in data‐limited settings. Across four heterogeneous disease cohorts, including Alzheimer's disease, type 2 diabetes, COVID‐19, and glioblastoma, ProMetNet consistently outperforms evaluated multi‐omics integration methods, including MOGONET, P‐NET, PEARL, and MOINER, maintaining high discriminative performance under substantial data downsampling. In addition to classification accuracy, the framework prioritizes biologically plausible protein–metabolite dependencies that are not captured by conventional differential or correlation‐based analyses. Importantly, pathway‐level signals identified by ProMetNet demonstrate consistent discriminative performance in independent large‐scale population data from the UK Biobank (N = 47,507), supporting their robustness and generalizability. Together, these results establish ProMetNet as a biologically grounded and interpretable framework for multi‐omics integration, enabling robust identification of structured molecular dependencies across diseases.
Ming-Hui Zhao, Na Zhou, Ruo-Tong Liu et al.· Advancement of science· 0 citations
Biomarker discovery from high-dimensional transcriptomic data is frequently hindered by the “curse of dimensionality” and model selection bias. To address this, we propose the Grouping–Scoring–Modeling (G-S-M) framework, a knowledge-driven pipeline that anchors feature selection in established disease–gene associations. G-S-M operates within a 100-iteration Monte Carlo ensemble architecture utilizing internal cross-validation to ensure unbiased evaluation. We evaluated this framework on seven cancer datasets, where it demonstrated robust discrimination with an overall mean F1-score of 0.84 across all datasets. The framework achieved the strongest performance on Acute Myeloid Leukemia (mean F1 = 0.99, AUC-ROC = 1.00) and maintained competitive accuracy even on challenging cohorts, while producing biologically interpretable gene panels traceable to named disease associations. Permutation tests (10,000 iterations) confirmed statistically significant disease–gene enrichment (p < 0.0001) in five of seven datasets, and independent protein interaction network analyses demonstrated significant enrichment of the selected features. Released as an open-source software suite with interactive interfaces, G-S-M provides a reproducible computational framework for candidate biomarker discovery.
Malik Yousef, Jens Allmer, Yasin Inal et al.· Applied Sciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.