Skip to content

A comparative analysis of automated techniques for security bug report identification

Jul 2026 · arXiv.org · Vol abs/2607.27893 · 0 citations · 28 references
Computer Science

TL;DR

A comparative analysis of several promising automated techniques to identify security-related bug reports using benchmark datasets indicates that SetFit achieves the best overall performance, and cross-project experiments demonstrate that transfer learning can improve performance for projects with limited data, but may degrade results for projects with strong project-specific characteristics.

Abstract

Timely identification of security-related bug reports is essential to minimize the window of vulnerabilities in software systems. Manually screening incoming bug reports to identify security-related issues is time-consuming, error-prone, and non-scalable for large-scale software systems. Thus, a variety of automatic techniques, including traditional machine learning (ML) techniques and large language models, have been proposed to facilitate this task. However, the literature remains fragmented. Most studies introduce or optimize a particular technique and evaluate it against a limited set of baselines, often under different experimental setups. As a result, it is difficult to compare their results and draw reliable conclusions about the effectiveness of existing approaches, leaving researchers and practitioners without clear guidance on which techniques are most suitable for the task. To address this gap, we conducted a comparative analysis of several promising automated techniques to identify security-related bug reports using benchmark datasets. We evaluated Logistic Regression, Support Vector Machines, Random Forest, OpenAI's GPT-5.2, BERT-base, RoBERTa, and SetFit (a state-of-the-art few-shot learning framework). Our results indicate that SetFit achieves the best overall performance, achieving an F1-score of 0.80 and outperforming other techniques on three of the four datasets. RoBERTa performs competitively and approaches SetFit in some projects, while traditional ML techniques, particularly Logistic Regression, remain a strong baseline in certain contexts. In contrast, GPT-5.2 performs poorly in both zero-shot and few-shot settings. In addition, cross-project experiments demonstrate that transfer learning can improve performance for projects with limited data, but may degrade results for projects with strong project-specific characteristics.

View source

Similar papers

Preprint Aug 2026

Bug Localization from Bug Reports: A Multi-Objective Approach

A class-level automated multi-objective search-based system to identify and rank potentially buggy classes from bug reports, with the main objective is to maximize similarity while minimizing the number of suggested faulty files.

W. Ahmad, Mehtab Kiran Suddle, Maryam Bashir · 0 citations
Open access Aug 2026

Static Code Analysis Framework for Automated Security Vulnerability Detection

Experimental results show that AST-based structural features substantially improve recall compared with the TF-IDF baseline, while the combined TF-IDF and AST representation maintains this improved performance.

Vani Pasupula, M. N. V. Manikanth, Nagaraju Vassey · 0 citations
#artificial intelligence Preprint Sep 2026

Exploring Automated Vulnerability Identification in JavaScript Code Using Large Language Models

JavaScript powers approximately 98.8% of all websites, making vulnerabilities in its code a significant security risk, yet existing detection approaches such as Static Application Security Testing (SAST) tools often fail to identify many real-world vulnerabilities when applied to isolated code snippets. This paper pres...

Manit Kaushik, Ishir Bhardwaj, Pranav Gupta et al. · 0 citations
#software testing Preprint Sep 2026

Investigating Developer-Reported Software Security Testing Challenges

Overall, developer-reported SST challenges extend beyond vulnerability detection and include workflow, configuration, interpretation, and remediation concerns and can help researchers, practitioners, educators, and tool providers improve SST usability, documentation, result interpretation, and remediation support.

Md Erfan, A. Ryan, Md. Rayhanur Rahman · 0 citations
Open access Aug 2026

Enhancing malware detection reliability in non-executable files using confidence score prediction

This article proposes a learning-based malware detection approach including two complementary parts, the development of binary classifiers, on an enriched dataset of related files, with an extended feature set to achieve high accuracy.

Rasoul Rezvani-Jalal, Morteza Zakeri, S. Parsa et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.