Skip to content

Author

Dongjun Yu

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

BAN-SDBPred: Improving Single-Stranded and Double-Stranded DNA-Binding Protein Prediction Using an Attention Network with Bilinear Convolution and Adaptive Sampling Strategy

DNA binding proteins play essential roles in numerous biological mechanisms. The DBPs can be either single-stranded binding (SSBs) or double-stranded binding (DSBs) to a DNA molecule. The in-depth identification of SSBs and DSBs has been a hot topic in bioinformatics and is involved in the drug discovery process. Traditional experimental methods failed to characterize the types of DBPs because of high cost and time constraints. While computational prediction of novel SSBs and DSBs has made significant progress, there are still challenges remaining in enhancing overall prediction performance. Methods: Here, we develop a novel BAN-SDBPred (Bilinear Attention Network for Single and Double Stranded DNA-Binding Protein Prediction) method. BAN-SDBPred leverages the evolutionary features by protein language model-based Evolutionary Scale Modeling 2 (ESM2), ProtT5, and a histogram of oriented gradient-based residue pairwise energy content matrix (RECM-HOG)-transformed energy estimation features from sequence alone. Then, the adaptive neighborhood-based sampling (ANBS) algorithm was adopted to solve the imbalance issue. Compared to other deep learning models, the bilinear attention network (BAN) learns the local and global enriched features from the sequences. Extensive experimental results anticipate that BAN-SDBPred outperforms the existing predictors in terms of all performance measures, such as Acc, F1, MCC, etc., on the training and independent test data. Our designed model has significant advantages in discriminating SSBs and DSBs from DBPs with an improved Acc of 2%, Precision of 4.5%, F1 of 21%, MCC of 9%, and area under curve (AUC) of 20%, respectively. We expect this research will help to predict large-scale novel SSBs and DSBs in particular and other binding problems in general. All data and models are available at 10.5281/zenodo.18718092.

K. Arshad, Muhammad Arif, A. Worachartcheewan et al. · 0 citations
Aug 2026

GAPEK: A General Framework for Multiparameter Enzyme Kinetic Prediction with Adaptive Learning.

The enzyme kinetic parameters, including the turnover number, Michaelis constant, and inhibition constant, are key metrics for assessing catalytic performance. Although deep learning models have recently incorporated multimodal information from enzymes and substrates to predict these parameters, several obstacles still persist. First, current data sets suffer from limited size, inconsistency, and a lack of unified standards. Second, most existing approaches prioritize cross-modal consistency but fail to sufficiently exploit the unique information residing in each individual modality. Meanwhile, although a limited number of studies have recognized that collaborative exploration of shared and specific information can enhance model performance, these methods remain difficult to directly apply to enzyme-substrate pairs, as enzyme-substrate relationships are inherently interactive rather than semantically equivalent counterparts. Third, the measured kinetic parameters are often unevenly distributed, which severely undermines the predictive accuracy of existing models when dealing with extreme value ranges. To resolve the above challenges, we first compile Kinetic-DB, a large-scale and consistently formatted data set from public resources. Building upon this data set, we develop GAPEK, a new framework for estimating enzyme kinetic parameters. In particular, an adaptive data augmentation module is devised to enrich the diversity of both enzyme and substrate sequences, thereby alleviating the adverse effects of data imbalance. Subsequently, we perform feature extraction using two pretrained models, ESM-2 for enzymes and Mole-BERT for substrates, to obtain multimodal embeddings. To decouple these complex interacting features, we introduce a tailored dual information exploration module to capture both modality-specific and cross-modal information, further refined by domain classification and distribution alignment loss functions. To explicitly handle the imbalanced data distribution, our base model, GAPEK, incorporates an adaptive density-weighted loss function. Building on this, we propose GAPEK+, which integrates the Squared Error Relevance Area (SERA) function to reconfigure the learning objective. By prioritizing high-relevance regions, GAPEK+ effectively calibrates the model's sensitivity to rare but critical extreme values, substantially mitigating the prediction bias inherent in heavy-tailed regression tasks. Experimental results demonstrate that both GAPEK and GAPEK+ achieve superior performance over existing state-of-the-art approaches, particularly across extreme parameter ranges, highlighting their potential as valuable tools applicable to enzyme engineering, synthetic biology, and drug discovery.

Cheng-Hao Zhu, Weiping Ding, Wei Zhang et al. · 0 citations
Aug 2026

Deep3MVPF: Multiview Deep Framework for the Prediction of Stability and m6A in mRNA 3'UTR.

Deep3MVPF, a multiview deep learning framework for 3'UTR stability prediction and m6A site identification, integrates a multiscale convolutional neural network, a k-mer de Bruijn graph neural network, and a secondary-structure graph neural network to jointly model sequence, topological, and structural representations.

Jun-Yi Liu, Qi Zhang, Jiangning Song et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.