A comparative analysis of several promising automated techniques to identify security-related bug reports using benchmark datasets indicates that SetFit achieves the best overall performance, and cross-project experiments demonstrate that transfer learning can improve performance for projects with limited data, but may degrade results for projects with strong project-specific characteristics.
Abstract
Timely identification of security-related bug reports is essential to minimize the window of vulnerabilities in software systems. Manually screening incoming bug reports to identify security-related issues is time-consuming, error-prone, and non-scalable for large-scale software systems. Thus, a variety of automatic techniques, including traditional machine learning (ML) techniques and large language models, have been proposed to facilitate this task. However, the literature remains fragmented. Most studies introduce or optimize a particular technique and evaluate it against a limited set of baselines, often under different experimental setups. As a result, it is difficult to compare their results and draw reliable conclusions about the effectiveness of existing approaches, leaving researchers and practitioners without clear guidance on which techniques are most suitable for the task. To address this gap, we conducted a comparative analysis of several promising automated techniques to identify security-related bug reports using benchmark datasets. We evaluated Logistic Regression, Support Vector Machines, Random Forest, OpenAI's GPT-5.2, BERT-base, RoBERTa, and SetFit (a state-of-the-art few-shot learning framework). Our results indicate that SetFit achieves the best overall performance, achieving an F1-score of 0.80 and outperforming other techniques on three of the four datasets. RoBERTa performs competitively and approaches SetFit in some projects, while traditional ML techniques, particularly Logistic Regression, remain a strong baseline in certain contexts. In contrast, GPT-5.2 performs poorly in both zero-shot and few-shot settings. In addition, cross-project experiments demonstrate that transfer learning can improve performance for projects with limited data, but may degrade results for projects with strong project-specific characteristics.
A class-level automated multi-objective search-based system to identify and rank potentially buggy classes from bug reports, with the main objective is to maximize similarity while minimizing the number of suggested faulty files.
W. Ahmad, Mehtab Kiran Suddle, Maryam Bashir· 0 citations
Experimental results show that AST-based structural features substantially improve recall compared with the TF-IDF baseline, while the combined TF-IDF and AST representation maintains this improved performance.
Vani Pasupula, M. N. V. Manikanth, Nagaraju Vassey· International Journal of Cre...· 0 citations
The findings support the claim that BF is a more structured and automation-friendly framework than CWE, and exploration reveals specific gaps in BF, including under-specified guidance on attributes.
Mohammad Nazmul Hoque, Shaswata Mitra, Subash Neupane et al.· 0 citations
JavaScript powers approximately 98.8% of all websites, making vulnerabilities in its code a significant security risk, yet existing detection approaches such as Static Application Security Testing (SAST) tools often fail to identify many real-world vulnerabilities when applied to isolated code snippets. This paper pres...
Manit Kaushik, Ishir Bhardwaj, Pranav Gupta et al.· 0 citations
Overall, developer-reported SST challenges extend beyond vulnerability detection and include workflow, configuration, interpretation, and remediation concerns and can help researchers, practitioners, educators, and tool providers improve SST usability, documentation, result interpretation, and remediation support.
Md Erfan, A. Ryan, Md. Rayhanur Rahman· 0 citations
This article proposes a learning-based malware detection approach including two complementary parts, the development of binary classifiers, on an enriched dataset of related files, with an extended feature set to achieve high accuracy.
Rasoul Rezvani-Jalal, Morteza Zakeri, S. Parsa et al.· International Journal of Inf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.