SNPoptimizer: a scalable genetic-algorithm framework to derive minimal discriminatory SNP panels from large genotyping datasets
The ability to efficiently discriminate genotypes is a critical step in genomics-assisted breeding, population genomics, biodiversity studies, traceability along food chains, and germplasm management. However, identifying the minimal and most informative subset of SNPs capable of uniquely distinguishing a large set of individuals remains a computationally challenging task. Here, we present SNPoptimizer, a user-friendly Shiny application that uses a genetic algorithm–based framework to optimally select discriminatory SNPs from large-scale genotyping datasets. By leveraging the evolutionary principles of selection, mutation, and crossover, SNPoptimizer iteratively identifies compact SNP panels that maximize genotype resolution. The application supports HapMap-formatted and VCF genotype files and includes an optional second-round optimization for resolving putative duplicates. We benchmarked SNPoptimizer across three independent datasets, including a tomato diversity panel, 820 Cauliflower genotypes, and a soybean diversity panel comprising 30 million variants across 1,511 samples. Across the three datasets, panels of 17–22 SNPs yielded R-VDP values ranging from 0.8744 to 0.9973, with complete discrimination obtained in Dataset III, demonstrating robust performance across different datasets. Cross-tool comparisons revealed complementary trade-offs among discriminatory power, panel size, runtime, and run-to-run reliability. SNPoptimizer provides a flexible solution for researchers seeking to reduce genotyping costs while maintaining high discriminative power.