Skip to content
Open access

Multilingual Source Code Vulnerability Detection Using Deep Learning: A Semantic Representation and Transfer Learning Approach

Aug 2026 · Engineering, Technology & Applied Science Research · 0 citations · 26 references

TL;DR

This work presents a deep learning approach for multilingual vulnerability detection that emphasizes semantic transfer rather than architectural complexity and suggests that stabilizing semantic representations during transfer is key to improving generalization while maintaining practical efficiency under moderate computational constraints.

Abstract

Detecting vulnerabilities in source code remains a major challenge as modern software systems increasingly span multiple programming languages. This work presents a deep learning approach for multilingual vulnerability detection that emphasizes semantic transfer rather than architectural complexity. The proposed method combines contextualized embeddings from CodeBERT/GraphCodeBERT with a BiLSTM and attention mechanism to capture code semantics and adopts a cross-lingual transfer setting where models trained on one language (e.g., Python) are evaluated on another (e.g., Java). To improve robustness under data imbalance, SMOTE and stratified cross-validation are incorporated into the training process. Experiments on Juliet and CodeXGLUE show that the model achieves an F1-score of about 0.87 and a ROC-AUC of 0.85 in intra-language settings, while guided fine-tuning improves cross-language F1-score by an average of 0.18 and ROC-AUC by approximately 0.13 compared with the direct transfer baseline. These results suggest that stabilizing semantic representations during transfer is key to improving generalization while maintaining practical efficiency under moderate computational constraints.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning

LWVIC4Code is proposed, a non-contrastive representation learning approach specifically designed for Type-IV clone detection that achieves competitive or superior performance without negative samples, benefits from layer-wise supervision, and generalizes effectively from Python to other languages, particularly Java and...

Luciano Marchezan, Kevin Delcourt, Eugene Syriani et al. · 0 citations
Open access Aug 2026

Enhancing Vulnerability Detection Precision through Ensemble Learning with Large Language Models

The results show that the ensemble techniques are a practical approach to boost the precision of LLMs in the detection of vulnerabilities and suggest that ensemble methods offer great potential in the advancement of software security analysis.

H. Al-Ofeishat, Azhar Hussain, M. Faheem et al. · 0 citations

CodeBERT-SENet: Adaptive Syntax-Semantic Fusion via Gated Attention for Python Bug Detection and Localization

This research proposes the first end-to-end multi-task architecture that jointly detects Python syntax errors and localizes their exact line position by adaptively fusing CodeBERT’s semantic embeddings with handcrafted syntactic features via a lightweight gating mechanism, without relying on Abstract Syntax Trees.

Rozali Ilham, Bahtiar Imran, Hasan Basri et al. · 0 citations
#small language model Preprint Aug 2026

Vulnerable Code Search: Transferable Attack for Code Language Models

This paper introduces a programming language-agnostic, transferable, adversarial attack that exploits this CLM vulnerability and demonstrates that this attack, even when computed using smaller code embedding models, is highly effective and transferable to larger, closed-source embedding models.

Kaicheng Wang, Liyan Huang, Jesse Thomason et al. · 0 citations
Conference Sep 2026

Comparative Analysis of Rewriting-Based Large Language Model-Generated Synthetic Code Detection

The case of synthetic code detection (i.e., written by a human or generated by AI) becomes more important in the field of education and security. Current code detection tools that depend on probabilistic and classification-based approaches have limitations in detecting various programming styles. To address this challe...

Maulana Arya Alambana, Arif Nurwidyantoro · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.