The results suggest that compared with traditional GS models that rely on a genomic kinship matrix, ML-based approaches offer greater flexibility in feature reduction, and that GS integration is promising for enhancing breeding efficiency in grapevines.
Abstract
Although grapevine (Vitis spp.) is among the oldest and most economically significant fruit species globally, its genetic improvement faces major bottlenecks due to long juvenile periods and extended cycles for phenotypic evaluation. In this context, genomic selection (GS) has emerged as an effective alternative to traditional selection, offering a robust framework to optimize breeding programs by significantly reducing generation intervals while enhancing predictive accuracy (PA) in early generations and expected genetic gains (EGGs). Nevertheless, factors such as minor allele frequency (MAF) and population size can significantly affect predictive models, even to the point of making their use unfeasible in breeding programs. In this context, this study evaluated the effect of data dimensionality reduction on GS accuracy by selecting single-nucleotide polymorphisms (SNPs) based on MAF thresholds. The experimental design tested the predictive capacities of four machine learning (ML) algorithms (ElasticNet, K-Neighbors, Support Vector Machine Regression, and XGBoost) alongside the conventional Genomic Best Linear Unbiased Prediction (gBLUP) model. These were validated using three SNP datasets (11,115, 9,494, and 6,100 markers) filtered by MAF levels of 0.05, 0.1, and 0.2 across six genetic traits, and EGGs were compared between conventional breeding and GS via the breeders’ equation. The results revealed that the ML models exhibited remarkable stability, with no significant differences in PA across the different MAF-based SNP densities, except for berry length, which showed a substantial difference with XGBoost at an MAF of 0.2. Conversely, gBLUP demonstrated high sensitivity to dimensionality reduction, with its performance significantly impacted by MAF filtering across all the traits. These results suggest that compared with traditional GS models that rely on a genomic kinship matrix, ML-based approaches offer greater flexibility in feature reduction. Additionally, compared with chemical traits, morphological traits generally had greater predictive ability. Furthermore, every GS model provided estimated genetic gains superior to traditional breeding, with improvements ranging from an 8.90-fold increase in berry length to a 2.86-fold increase in total soluble solids, confirming that GS integration is promising for enhancing breeding efficiency in grapevines.
Proper analysis of phenotypic data is essential for reliable genomic prediction (GP) and sustained genetic gain in breeding programs. In this study, we used phenotypic and genotypic data generated from the winter wheat breeding program of Deutsche Saatveredelung AG (DSV), Lippstadt, Germany, comprising 1,941 genotypes and 6,335 SNP markers across three traits: grain yield, plant height, and heading date. We evaluated the impact of different two-stage analysis strategies: one using the environment (year × location combination) as the analysis unit (TS-S1) and the other using the breeding stage as the analysis unit (TS-S2), each with and without accounting for breeding-stage effects (-YesBS and -NoBS), on the computation of best linear unbiased estimates (BLUEs) and genome-wide prediction ability (PA), defined as the correlation between BLUEs and predicted values. The performance of these 4 different second stage models (TS-S1-NoBS, TS-S1-YesBS, TS-S2-NoBS, and TS-S2-YesBS) were evaluated using the extended genomic best linear prediction (EGBLUP) model under 5-fold cross validation (5-fold CV), leave one year out cross validation (LOY-CV) and leave one breeding stage out cross validation (LOBS-CV) scenarios. In the presence of the strong breeding stage effect ignoring the breeding stage in the model resulted in biased BLUEs and overestimated prediction abilities in the 5-fold CV, and underestimated prediction abilities in the LOY-CV and LOBS-CV. These unstable prediction abilities are driven by confounded environmental effects in the BLUEs due to omission of the breeding-stage effect. In contrast, models that accounted for the breeding stage produced more reliable BLUEs and more stable prediction results across all cross-validation scenarios, with TS-S1-YesBS performing best overall. Overall, our findings demonstrate that breeding stage effects mainly arise from differences in growing conditions and must be adequately considered. Therefore, including breeding stage in phenotypic models is critical to obtaining unbiased BLUEs and ensuring accurate genomic prediction and selection decisions.
Ravindra Reddy Gundala, Georg Witte, Jost Doernte et al.· Frontiers in Plant Science· 0 citations
ABSTRACT Genomic selection (GS) is a crop and livestock improvement method suited for predicting complex agronomical traits, while genomic prediction (GP) is the development of GS models prior to their practical use in breeding programs. One of the challenges in GP is accounting for how complex genomic interactions, such as epistasis, affect the resulting phenotype. Incorporating haplotypes and Machine Learning (ML) into GP models are two methods for accounting for local epistasis and non‐linear relationships. This study compared linear‐ and ML‐GP models for single nucleotide polymorphisms (SNPs) and haplotypes in Brassica napus . A publicly available dataset of 991 B. napus individuals, 4 286 896 SNPs, and the traits flowering time, oil content and oleic acid content was used for all GP models. The ML models improved trait prediction accuracy for all tested traits in both SNP‐ and haplotype‐based GP. While haplotypes did not significantly boost prediction accuracy over SNPs, they captured novel genetic variation and offered a broader diversity of variants for selection, suggesting a qualitative advantage for long‐term breeding goals. Here, we demonstrate how haplotypes and ML can improve GS in B. napus .
Tessa R. Macnish, H. Al-Mamun, Thomas Bergmann et al.· Plant Biotechnology Journal· 0 citations
This narrative review critically examines recent advances in genomic selection for rice and its integration with high-throughput genotyping, high-throughput phenotyping, machine learning, multi-environment prediction, and speed breeding.
Ha Duc Chu, T. Q. Nguyen, Loc Van Nguyen et al.· Genes· 1 citation
The results suggest that, in elite wheat germplasm characterized by long-range linkage disequilibrium and strong realized genomic relationships, medium-density targeted genotyping platforms can retain most of the predictability achieved by higher-density systems.
In this study, we assessed the impact of prior-induced regularization using six real datasets from wheat, rice, and potato, spanning 107–758 genotypes, 2–12 environments, 1–18 traits, and 2744–108,024 molecular markers. Two modeling scenarios were evaluated: (i) unimodal genomic prediction based solely on marker information and (ii) multimodal (multi-component) prediction integrating genomic, environmental, and genotype-by-environment (G × E) effects. Predictive performance was evaluated using Pearson’s correlation (COR) and normalized root mean squared error (NRMSE) under 10 repeated random 50% training–50% testing partitions, representing prediction of untested lines in tested environments. Bayesian genomic prediction (BGP) relies on prior distributions to regulate shrinkage and stabilize inference in high-dimensional settings. We evaluated whether predictive performance was driven primarily by the type of Bayesian prior or by the presence of effective prior-induced regularization. Across most datasets, regularized Bayesian models achieved higher predictive correlations and markedly lower NRMSE than the weakly regularized or unregularized baseline. Differences among regularized prior families were generally modest, whereas weakening or removing regularization frequently produced unstable estimates and inflated prediction error. Predictive results were obtained for both winter-wheat datasets as well as for the rice, potato, and DMario datasets. In multimodal analyses, models with coherent regularization across genomic, environmental, and genotype-by-environment components were generally more accurate and stable than configurations in which regularization was absent or weakened in key components. Rice_Kim_2020 was an informative exception in which the baseline remained competitive. These results show that the principal empirical contrast is the presence versus absence of effective prior-induced regularization, rather than a universal ranking of Bayesian prior families. Appropriate regularization should therefore be treated as a central model-design decision in genomic prediction.
Osval A. Montesinos-López, José Elías Peregrina-Chavarría, A. Montesinos-López et al.· Plants· 0 citations
The exponential increase in the number of genotyped animals, combined with the availability of high-density SNP chips has introduced computational challenges for routine genomic evaluations, particularly during the construction of the genomic relationship matrix. Although higher-density SNP panels can facilitate the identification of causal mutations, their use substantially increases computational requirements without a proportional gain in genomic prediction performance. To optimize computational efficiency while maintaining accuracy of genomic predictions, this study compared five SNP selection strategies (i.e., random sampling, random sampling with inclusion of informative SNPs, linkage disequilibrium (LD)-based pruning, a Shannon entropy–based machine learning approach, and FST-based prioritization) to develop reduced-density panels for Nellore cattle. Using high-density (HD) genotype data comprising 437,650 SNPs from 304,782 animals (after quality control) as reference, three reduced-density panels (25K, 45K, and 65K SNPs) panels were tested across five traits (i.e., Age at first calving, Stayability, Weaning weight, Yearling weight, Muscling) with diverse genetic architectures. Genomic estimated breeding values (GEBVs) derived from these reduced panels were compared to those obtained from the HD reference panel using Pearson’s correlations, under both genomic best linear unbiased prediction (GBLUP) and single-step GBLUP (ssGBLUP) methods. In the GBLUP model, prediction accuracy generally improved with increased marker density. Random selection with and without the informative SNPs consistently yielded the highest accuracies, whereas the FST-based approach showed the lowest agreement with the HD reference across all densities. In contrast, ssGBLUP demonstrated strong robustness to marker reduction, producing uniformly high correlations (≈1.00) across all SNP densities and selection strategies. These findings indicate that optimized low-density SNP panels maintain prediction accuracy comparable to HD panels, offering a cost-effective tool for large-scale genomic evaluations. Author Summary Genomic selection has transformed cattle breeding by allowing producers to identify animals with superior genetic potential using DNA information. However, modern genomic evaluations often rely on very large genetic datasets that require substantial computing power and increase genotyping costs, particularly in large breeding populations such as Nellore cattle in Brazil. In this study, we evaluated whether reduced-density marker panels could maintain the same level of prediction accuracy as high-density panels commonly used in genomic evaluations. We compared five different strategies for selecting informative genetic markers and tested panels containing different numbers of markers across economically important traits. We found that reduced-density panels, particularly those developed using random or linkage-based selection methods, produced genomic predictions highly similar to those obtained with high-density panels. In addition, prediction methods that combined genomic and pedigree information remained highly robust even with fewer markers. Our findings suggest that reduced-density panels can support accurate and cost-effective genomic evaluations, allowing breeding programs to evaluate more animals more frequently while reducing computational demands.
A. R. Ogunbawo, H. Mulim, J. Hidalgo et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.