Overall, XAI should be considered a framework for model interpretation, feature prioritization, and hypothesis generation rather than a replacement for experimental validation in plant genomics and breeding, and the common misconception that feature importance implies biological causality is emphasized.
Abstract
Machine learning has become an important tool in plant genomic prediction for modeling complex genotype–phenotype relationships and improving breeding decisions. However, many high-performing models, particularly ensemble and deep learning approaches, remain difficult to interpret, limiting their biological applicability. This review summarizes major machine learning methods and explainable artificial intelligence (XAI) approaches used in plant genomics, including SHAP (SHapley Additive exPlanations), LIME (Local Interpretable Model-Agnostic Explanations), attention mechanisms, permutation importance, tree-based feature importance, and gradient-based attribution methods. XAI can help identify influential SNPs, genomic regions, candidate genes, regulatory elements, and omics features associated with complex traits. For example, SHAP analysis in an almond germplasm collection identified a genomic region associated with shelling fraction, illustrating how XAI can generate testable hypotheses for further validation. The review further discusses applications in trait prediction, breeding, functional genomics, and multi-omics integration. Importantly, we emphasize major limitations, including data bias, model instability, correlated genomic markers, limited model transferability, and the common misconception that feature importance implies biological causality. We recommend integrating XAI with linkage disequilibrium pruning, stability assessment, biological annotation, and experimental validation before prioritizing candidate genes. Overall, XAI should be considered a framework for model interpretation, feature prioritization, and hypothesis generation rather than a replacement for experimental validation in plant genomics and breeding.
Predicting complex phenotypes from plant-microbiome multi-omics data has advanced significantly with machine learning and deep learning. However, traditional black-box models remain limited by their focus on predictive performance rather than mechanistic biological understanding. To bridge this gap, explainable artific...
Lu-Pin Deng, Lei Liu, Yu Luo et al.· International Journal of Mol...· 0 citations
Crop improvement increasingly depends on extracting useful breeding signals from data that span DNA sequence variation, gene regulation, molecular phenotypes, high-throughput field measurements and environmental exposure. Multi-omics can connect genotype to phenotype through intermediate biological layers, while artifi...
Abstract Genomic prediction has become a central paradigm in biology, enabling quantitative inference of genetic contributions to complex traits across humans, animals, and plants. Although genomic research in human genetics and animal breeding shares a highly homologous methodological foundation, significant barriers...
Jin-Bu Wang, Li-Li Du, Zhi-Da Zhao et al.· Briefings in Bioinformatics· 0 citations
Understanding the complex relationship between genotype and phenotype is a critical goal in biological research, with significant implications for fields such as medicine, agriculture, and biosciences. The ability to predict phenotypes from genetic information can advance crop improvement, precision medicine, and disea...
A. Rahman, Manideep Kolla, Hai-Feng Wang et al.· IISE Annual Conference &...· 0 citations
In this study, a dataset derived from a natural rice population was used to develop two genomic prediction models, genomic best linear unbiased prediction (GBLUP) and a convolutional neural network (CNN), together with a gene-based crop modeling framework, which provided valuable insights into modeling genotype-by-envi...
Jinhan Zhang, Wei-Jie Tang, Hong-Wei Ma et al.· Theoretical and Applied Gene...· 0 citations
A novel tree-based feature engineering framework that uses Classification and Regression Trees (CART) to explicitly encode high-order interaction decision paths as dummy variables that significantly improves classification accuracy and model interpretability compared to using the original feature space alone is propose...
J. Byun, Dheeman Saha, Younghun Han et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.