Skip to content
Open access

Malicious Repository Detection Using Transformer-Based Models: A Metadata-Driven Approach

2026 · IEEE Access · Vol 14, pp. 146250-146269 · 0 citations · 34 references

TL;DR

A metadata-driven detection framework that fine-tunes DistilBERT on repository metadata text, specifically the description, README, and topics, and formulates detection as a binary classification task between malware and benign repositories, suggesting that the framework is promising for threshold-based screening scenarios.

Abstract

GitHub and similar platforms host millions of software repositories, a fraction of which openly distribute malicious source code. Existing work on malicious repository detection has primarily relied on handcrafted features, structural analysis, or multimodal code embeddings, leaving the potential of transformer-based natural-language models largely unexplored. We present a metadata-driven detection framework that fine-tunes DistilBERT on repository metadata text, specifically the description, README, and topics, and formulates detection as a binary classification task between malware and benign repositories. Our dataset combines malware repositories sourced from hacker-community crawls with benign repositories from the CrossSim benchmark, yielding 1,478 BERT-ready samples after leakage-aware preprocessing. Training under a 10-fold GroupKFold protocol stratified by repository identifier to reduce same-repository leakage across training and validation splits yields a mean accuracy of 98.98 % and a mean F1-score of 99.22 %, outperforming all Repo2Vec baseline variants by up to 9.22 percentage points in F1-score. A bagging ensemble of three DistilBERT models further improves the F1-score to 99.75 %. Calibration analysis via temperature scaling indicates well-calibrated probabilistic outputs (ECE = 0.0041 for the cross-validation model and ECE = 0.0074 for the ensemble after scaling), suggesting that the framework is promising for threshold-based screening scenarios. Interpretability and robustness analyses further show that the model’s decisions are driven by distributed contextual signals rather than a handful of overt security terms: malware recall remains at 99.18 % even after masking all 22 such keywords, providing direct evidence against a keyword-shortcut explanation of the model’s performance. To the best of our knowledge, no prior work has applied a BERT-family transformer model to supervised repository-level malware-versus-benign binary classification using metadata alone. The main limitation is that the benign benchmark is restricted to Java-language repositories, whereas the malware repositories span multiple languages.

Read PDF

Similar papers

Preprint Aug 2026

MalTotal: Cost-Effective and Language-Agnostic Malicious Code Poisoning Detection for Millions of Repositories

MalTotal leverages LLM-assisted semantic reasoning to identify sensitive APIs, perform hybrid semantic slicing, and reconstruct malicious behavior contexts while reducing analysis overhead, demonstrating the effectiveness, scalability, and cost-efficiency of MalTotal in mitigating large-scale code poisoning attacks.

Jian Zhao, Shenao Wang, Qingyang Wu et al. · 1 citation · ⚡1
Open access

Recovering vulnerability detection on realistically-distributed code

This thesis rebuilds Real-Vul through a controlled Code Property Graph pipeline, measure and remove a 37% content leak intrinsic to whole-codebase sampling, and train a relational graph neural network with a disciplined class-imbalance recipe: focal loss, class-aware undersampling, and fine-tuning of a GraphCodeBERT no...

Brian Kade Betterton · 0 citations
Open access Aug 2026

Multilingual Source Code Vulnerability Detection Using Deep Learning: A Semantic Representation and Transfer Learning Approach

This work presents a deep learning approach for multilingual vulnerability detection that emphasizes semantic transfer rather than architectural complexity and suggests that stabilizing semantic representations during transfer is key to improving generalization while maintaining practical efficiency under moderate comp...

Tuan Nguyen Kim, Nin Ho Le Viet, Chieu Ta Quang · 0 citations
Open access Sep 2026

Metadata Compressibility and Evaluation Bias in Malicious Package Detection for NPM and PyPI

Malicious packages in open source-software supply chains are a growing security concern, yet machine learning detectors built on registry metadata are difficult to interpret and are typically evaluated under protocols susceptible to data leakage. We construct a dataset of 3330 package versions from NPM and PyPI in whic...

Hanan Moufid, M. El Ghazouani, Moulay Ahmed el Kiram · 0 citations
Open access Aug 2026

AuthProtect: An Incremental Learning Framework for Android Malware Detection via Permission-Exploitation Mapping

Android's widespread adoption and open ecosystem make it a primary target for malware, a challenge exacerbated by internet fragmentation resulting in non-stationary data distributions across regions. This work presents AuthProtect, a scalable malware detection framework based on incremental learning and a novel permiss...

Maksim Iavich, Razvan Bocu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.