BINND is developed, a binding and interaction neural network to predict non-orthogonal DNA interactions, which could aid diagnostic, bioengineering, and DNA origami design, and supporting a shift toward exploiting the full sequence space.
Abstract
Common frameworks in molecular bioengineering and synthetic biology focus on orthogonality, viewing weak or non-specific interactions as problems to avoid. This constrains the usable sequence space, limits scalability, and neglects scenarios where synthetic systems must operate within natural backgrounds of high sequence diversity. Harnessing the full space is difficult because models are lacking that can accurately and quickly predict non-orthogonal interactions and be validated against ground truth data. Here we develop BINND — Binding and Interaction Neural Network for DNA — using DNA-DNA interactions as a testbed. BINND combines an ultra-high throughput platform measuring millions of interactions with a deep learning model attaining accuracies above 80%, generalizing across diverse sequences and running 50 times faster than current models. We demonstrate its value with a searchable DNA network of fictitious storybook characters. BINND enables accurate prediction for diagnostics, bioengineering, and DNA origami, supporting a shift toward exploiting the full sequence space. Weak or non-specific DNA-DNA interactions can confound system orthogonality in synthetic biology. Here the authors develop BINND, a binding and interaction neural network to predict non-orthogonal DNA interactions, which could aid diagnostic, bioengineering, and DNA origami design.
Abstract Accurate predictions of DNA-binding residues (DBRs) in protein sequences facilitate decoding molecular-level mechanisms underlying cellular functions that involve protein–DNA interactions. While dozens of these predictors have been released, they target either structured or intrinsically disordered regions (IDRs), and the latter were trained to predict less detailed DNA-binding IDRs rather than DBRs. Given this dichotomy, the structure-trained methods underperform on disordered proteins, and vice versa. Moreover, they suffer from high cross-prediction rates, incorrectly labeling many residues that interact with non-DNA ligands as DBRs. We address these issues by introducing DNAreader, the first predictor specifically designed to predict DBRs in the structured and disordered sequence regions. DNAreader relies on an innovative stacked transformer encoder network that combines batch training and contrastive learning, which substantially boosts predictive performance. Using two low-similarity test datasets, we demonstrate that DNAreader statistically outperforms existing tools, performs well for structured and disordered regions, and produces very few cross-predictions. We also developed the DNAreaderDBIDR module, which accurately predicts DNA-binding IDRs, providing flexibility to identify DBRs within IDRs or to predict entire disordered DNA-binding regions. We release DNAreader as a user-friendly web server at http://biomine.cs.vcu.edu/servers/DNAreader/, with the corresponding source code at https://github.com/jianzhang-xynu/DNAreader.
Genomic foundation models pretrained on DNA sequence have achieved strong performance across a range of tasks, but sequence-only representations cannot fully capture regulatory information reflected by additional DNA-centric modalities. Existing multimodal genomic models are often optimized for specific prediction tasks rather than for learning reusable embeddings shared across downstream analyses. However, directly fusing heterogeneous genomic modalities is challenging because sparse, peak-shaped regulatory signals and dense sequence representations have markedly different statistical structures, making naive multimodal alignment prone to degenerate near-zero solutions. We present a self- supervised DNA-centric multimodal foundation model that addresses this gap, integrating DNA sequence embeddings with local and global chromatin accessibility in a shared multimodal encoder to produce reusable window-level embeddings that support both masked reconstruction during pre-training and downstream prediction tasks. We diagnose this heterogeneous-modality alignment failure and show that global normalization substantially alleviates collapse, enabling effective joint learning across modalities. The resulting embeddings improve multiple downstream evaluations of regulatory function, including regulatory activity prediction, regulatory signal ranking and chromatin accessibility peak detection, achieving a 4.6-fold AUPRC improvement over the DNA-only baseline in peak detection, and further improving external validation on ClinVar, GTEx eQTL and PBMC caQTL datasets. The framework is extensible to additional regulatory modalities, providing a methodological basis for multimodal DNA foundation models.
Gene-therapy design depends on identifying regulatory sequences that drive the right level, timing, and cell-type specificity of expression. Regulatory DNA models offer a way to prioritize such sequences computationally before committing candidates to biological testing. Biological validation involves DNA synthesis, cloning, cell culture, sequencing, and functional screening, so training compute is part of the same constrained discovery pipeline rather than an isolated modeling expense. Reducing the compute required to reach a target pretraining quality could shift time and budget toward larger candidate screens, additional assays, more cell contexts, and broader follow-up validation. Given that Adam-style optimizers are widely used for training genomic sequence models, we study whether Muon can provide a more compute-efficient alternative for regulatory DNA pretraining. We provide an in-depth analysis by training Transformer models (26M–420M parameters) on ENCODE cis-regulatory sequences with Adam and Muon while holding architecture, data, and non-optimizer hyperparameters fixed and varying optimizer family, norm-control scheme, learning rate, and model width. In the largest-scale matched-target comparison, Muon reaches Adam-matched perplexity targets with a median FLOP reduction of 35.4% and a median wall-clock time reduction of 38.5%. The analysis further shows that optimizer rankings depend on norm control: independent weight decay pairs more favorably with Muon than Hyperball in this setting. These findings indicate that optimizer update structure and norm-control choices are practical levers for reducing the training resources required to reach matched perplexity targets in regulatory DNA pretraining.
Deep-Interact Studio is, to the authors' knowledge, the only such platform to combine fine-grained per-layer model customization with multi-model comparison and interpretability, offering a flexible and transparent alternative to fixed, single-purpose tools.
Dipayan Sarkar, K. Bardhan, Chiranjib Sarkar· bioRxiv· 0 citations