Aug 2026· Engineering, Technology & Applied Science Research· 0 citations· 26 references
TL;DR
This work presents a deep learning approach for multilingual vulnerability detection that emphasizes semantic transfer rather than architectural complexity and suggests that stabilizing semantic representations during transfer is key to improving generalization while maintaining practical efficiency under moderate computational constraints.
Abstract
Detecting vulnerabilities in source code remains a major challenge as modern software systems increasingly span multiple programming languages. This work presents a deep learning approach for multilingual vulnerability detection that emphasizes semantic transfer rather than architectural complexity. The proposed method combines contextualized embeddings from CodeBERT/GraphCodeBERT with a BiLSTM and attention mechanism to capture code semantics and adopts a cross-lingual transfer setting where models trained on one language (e.g., Python) are evaluated on another (e.g., Java). To improve robustness under data imbalance, SMOTE and stratified cross-validation are incorporated into the training process. Experiments on Juliet and CodeXGLUE show that the model achieves an F1-score of about 0.87 and a ROC-AUC of 0.85 in intra-language settings, while guided fine-tuning improves cross-language F1-score by an average of 0.18 and ROC-AUC by approximately 0.13 compared with the direct transfer baseline. These results suggest that stabilizing semantic representations during transfer is key to improving generalization while maintaining practical efficiency under moderate computational constraints.
LWVIC4Code is proposed, a non-contrastive representation learning approach specifically designed for Type-IV clone detection that achieves competitive or superior performance without negative samples, benefits from layer-wise supervision, and generalizes effectively from Python to other languages, particularly Java and...
Luciano Marchezan, Kevin Delcourt, Eugene Syriani et al.· 0 citations
The results show that the ensemble techniques are a practical approach to boost the precision of LLMs in the detection of vulnerabilities and suggest that ensemble methods offer great potential in the advancement of software security analysis.
H. Al-Ofeishat, Azhar Hussain, M. Faheem et al.· Engineering, Technology &...· 0 citations
This research proposes the first end-to-end multi-task architecture that jointly detects Python syntax errors and localizes their exact line position by adaptively fusing CodeBERT’s semantic embeddings with handcrafted syntactic features via a lightweight gating mechanism, without relying on Abstract Syntax Trees.
Rozali Ilham, Bahtiar Imran, Hasan Basri et al.· 0 citations
This paper introduces a programming language-agnostic, transferable, adversarial attack that exploits this CLM vulnerability and demonstrates that this attack, even when computed using smaller code embedding models, is highly effective and transferable to larger, closed-source embedding models.
Kaicheng Wang, Liyan Huang, Jesse Thomason et al.· 0 citations
The case of synthetic code detection (i.e., written by a human or generated by AI) becomes more important in the field of education and security. Current code detection tools that depend on probabilistic and classification-based approaches have limitations in detecting various programming styles. To address this challe...