Aug 2026· Match-communications in Mathematical and in Computer Chemistry· 0 citations· 32 references
Abstract
Identifying transcription factor binding sites (TFBSs) is fundamental to understanding complex gene regulatory mechanisms and the functions of non-coding regions. Although existing methods have achieved substantial strides, capturing both local structural features and long-range spatial dependencies within DNA sequences remains a major challenge for improving prediction accuracy. In this study, we propose DNCLA, a deep learning model that synergizes multisize convolutional fusion, Bidirectional Long ShortTerm Memory (Bi-LSTM) networks, and a multi-head self-attention
mechanism. At the feature extraction level, DNCLA breaks through the limitations of traditional single-sequence encoding by fusing Nucleotide Chemical Properties (NCP) with Dinucleotide Physicochemical Properties (DPCP). NCP provides a refined characterization of chemical differences between bases based on ring structures, hydrogen bond sites, and functional group properties, while DPCP introduces parameters such as local structural stability and geometric flexibility of the DNA. Subsequently, the model extracts spatial evolution from these high-dimensional features through a multi-size convolutional module; captures long-range spatial dependencies using Bi-LSTM layers; and employs a multi-head self-attention mechanism to achieve adaptive weight distribution of global features, thereby enhancing the perception of key regulatory motifs. Results from training and testing the proposed model on 165 ChIPseq datasets demonstrate that DNCLA possesses robust generalization capabilities and high predictive performance in TFBSs identification. This suggests that the incorporation of physicochemical features better elucidates the essence of interactions between transcription factors and DNA.
Lysine crotonylation (Kcr) is an important post-translational modification (PTM) involved in diverse biological processes, including chromatin regulation, protein function modulation, and cellular signaling. Although mass spectrometry-based proteomics has substantially expanded the identification of Kcr sites, experimental screening remains labor-intensive, costly, and difficult to apply at proteome scale. Computational methods provide an efficient strategy for prioritizing candidate Kcr sites. However, most existing predictors mainly rely on sequence-derived representations and insufficiently exploit protein structural context. In this study, we propose BLOSSOM-Kcr, a structure-informed deep learning framework for Kcr site prediction. BLOSSOM-Kcr integrates BLOSUM62-based sequence substitution features with residue-level structural descriptors, including secondary structure, solvent accessibility, backbone geometry, and spatial neighborhood information. The fused residue-level representation is further processed by residual convolutional blocks, channel attention, bidirectional long short-term memory (BiLSTM) layers, and attention pooling to capture local motif patterns, informative feature dimensions, and contextual dependencies surrounding candidate lysine residues. Fivefold cross-validation was performed for model optimization and comparative analysis, while an independent test set was used for final evaluation against existing Kcr site predictors. On the independent test set, BLOSSOM-Kcr achieved an AUC of 0.9023, an MCC of 0.6479, and an F1-score of 0.8336, outperforming representative Kcr site predictors. These results suggest that BLOSSOM-Kcr provides an effective structure-aware framework for Kcr site prediction.
The ESM2-Kcr model not only enhances the understanding of protein regulation but also holds great potential in identifying disease biomarkers and facilitating drug development.
Kai Liu, Sheng-Li Zhang, Jingyi Ren· Journal of Computer-Aided Mo...· 0 citations
Deep3MVPF, a multiview deep learning framework for 3'UTR stability prediction and m6A site identification, integrates a multiscale convolutional neural network, a k-mer de Bruijn graph neural network, and a secondary-structure graph neural network to jointly model sequence, topological, and structural representations.
Jun-Yi Liu, Qi Zhang, Jiangning Song et al.· Journal of Chemical Informat...· 0 citations
PLM-ArgMe is presented that is based on a symmetry-sensitive Transformer framework using context-aware ESM-2 residue embeddings, which is mapped through a novel Bio-Symmetric Mirrored Sinusoidal Encoding strategy to address the biological symmetry hypothesis of arginine methylation.
Nitika Bhatt, Kartik Joshi, R. Rout et al.· Biochemical and Biophysical...· 0 citations
DNA binding proteins play essential roles in numerous biological mechanisms. The DBPs can be either single-stranded binding (SSBs) or double-stranded binding (DSBs) to a DNA molecule. The in-depth identification of SSBs and DSBs has been a hot topic in bioinformatics and is involved in the drug discovery process. Traditional experimental methods failed to characterize the types of DBPs because of high cost and time constraints. While computational prediction of novel SSBs and DSBs has made significant progress, there are still challenges remaining in enhancing overall prediction performance. Methods: Here, we develop a novel BAN-SDBPred (Bilinear Attention Network for Single and Double Stranded DNA-Binding Protein Prediction) method. BAN-SDBPred leverages the evolutionary features by protein language model-based Evolutionary Scale Modeling 2 (ESM2), ProtT5, and a histogram of oriented gradient-based residue pairwise energy content matrix (RECM-HOG)-transformed energy estimation features from sequence alone. Then, the adaptive neighborhood-based sampling (ANBS) algorithm was adopted to solve the imbalance issue. Compared to other deep learning models, the bilinear attention network (BAN) learns the local and global enriched features from the sequences. Extensive experimental results anticipate that BAN-SDBPred outperforms the existing predictors in terms of all performance measures, such as Acc, F1, MCC, etc., on the training and independent test data. Our designed model has significant advantages in discriminating SSBs and DSBs from DBPs with an improved Acc of 2%, Precision of 4.5%, F1 of 21%, MCC of 9%, and area under curve (AUC) of 20%, respectively. We expect this research will help to predict large-scale novel SSBs and DSBs in particular and other binding problems in general. All data and models are available at 10.5281/zenodo.18718092.
K. Arshad, Muhammad Arif, A. Worachartcheewan et al.· ACS Omega· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.