Systematic Benchmarking of DNA Sequence Encoding Strategies for Predicting Regulatory Effects of Non-Coding SNPs
Abstract
Non-coding single nucleotide polymorphisms (SNPs) are key modulators of gene regulation and have been implicated in diverse complex traits and diseases. With the growing demand for accurate functional interpretation of non-coding variants, the choice of encoding strategies becomes critical in downstream predictive modeling. Despite recent advances, a systematic evaluation of encoding approaches tailored for non-coding SNPs remains lacking. To address this gap, we present a comprehensive benchmark that evaluates six representative encoding strategies, including categorical, semantic, and functional embeddings, across three quantitative trait loci (QTL)-related prediction tasks. The study encompasses nine machine learning and deep learning models and incorporates experimental controls and repeated trials to ensure robustness and reproducibility. We assess each strategy along multiple dimensions, such as interpretability, representation abundance, and computational efficiency. Rather than ranking individual methods, our analysis emphasizes the interaction between encoding strategies, model types, and preprocessing protocols, and highlights their collective influence on predictive performance. This work establishes a standardized framework for evaluating non-coding SNP representations and offers guidance for selecting and optimizing prediction pipelines in regulatory genomics.