Aug 2026· Engineering, Technology & Applied Science Research· 0 citations· 17 references
TL;DR
The results show that the ensemble techniques are a practical approach to boost the precision of LLMs in the detection of vulnerabilities and suggest that ensemble methods offer great potential in the advancement of software security analysis.
Abstract
This study investigates the use of ensemble learning with Large Language Models (LLMs) to improve the accuracy of software vulnerability prediction, following a structured experimental approach to assess whether combining multiple models can enhance performance. Three baseline models, CodeBERT, GraphCodeBERT, and CodeT5, were trained and assessed on the Devign dataset, which provides a large collection of labeled source code snippets. Their outputs were then integrated using three ensemble techniques: Majority Voting, Weighted Voting, and Stacking. Precision, recall, and F1-score metrics were used to gauge performance. Ensemble approaches outperformed all standalone models. In particular, Majority Voting increased precision from 0.601 (CodeBERT) to 0.690, representing a 14.81% improvement. Keeping in view the detection accuracy, this study focused on reducing the false positives. The results show that the ensemble techniques are a practical approach to boost the precision of LLMs in the detection of vulnerabilities. Ensemble learning can address the challenges faced by standalone models by reducing false positives and improving the overall trade-off between accuracy and reliability. The study suggests that ensemble methods offer great potential in the advancement of software security analysis.
The results indicate that traditional ML models, especially random forest and extra trees, are still very effective for metric-based defect prediction, while DL and multi-modal approaches need to be fed with richer software artifacts to reach their full potential.
Amro Mohammad Abed Alfattah Abdin, Mohanad Alayedi, Ahmad M. Jaradat· Journal of Supercomputing· 0 citations
This work presents a deep learning approach for multilingual vulnerability detection that emphasizes semantic transfer rather than architectural complexity and suggests that stabilizing semantic representations during transfer is key to improving generalization while maintaining practical efficiency under moderate comp...
Tuan Nguyen Kim, Nin Ho Le Viet, Chieu Ta Quang· Engineering, Technology &...· 0 citations
This study benchmarks machine learning (ML) and deep learning models for intrusion detection using the publicly available FWAF dataset, and emphasizes the value of balancing performance with interpretability, empowering Security Operations Centers (SOC) to validate automated decisions and foster trustworthy AI-driven w...
Fatih Ünlü, Y. Sönmez, Murat Dener· El-Cezeri Fen ve Mühendislik...· 0 citations
The machine-learning model proposed in the paper has improved performance (efficiency, accuracy, and false-negative rate) for identifying vulnerabilities compared to conventional methods, and shows much greater flexibility and reliability when analyzing large code bases and different categories of vulnerabilities.
Findings indicate that ensembling approaches can be statistically significant and effective on larger datasets, where the best-performing ensemble improved performance by 37% over its individual LLMs on the commercial large-scale code.
M. Chochlov, Gul Aftab Ahmed, J. Patten et al.· ACM Transactions on Software...· 1 citation
This article proposes a learning-based malware detection approach including two complementary parts, the development of binary classifiers, on an enriched dataset of related files, with an extended feature set to achieve high accuracy.
Rasoul Rezvani-Jalal, Morteza Zakeri, S. Parsa et al.· International Journal of Inf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.