A metadata-driven detection framework that fine-tunes DistilBERT on repository metadata text, specifically the description, README, and topics, and formulates detection as a binary classification task between malware and benign repositories, suggesting that the framework is promising for threshold-based screening scenarios.
Abstract
GitHub and similar platforms host millions of software repositories, a fraction of which openly distribute malicious source code. Existing work on malicious repository detection has primarily relied on handcrafted features, structural analysis, or multimodal code embeddings, leaving the potential of transformer-based natural-language models largely unexplored. We present a metadata-driven detection framework that fine-tunes DistilBERT on repository metadata text, specifically the description, README, and topics, and formulates detection as a binary classification task between malware and benign repositories. Our dataset combines malware repositories sourced from hacker-community crawls with benign repositories from the CrossSim benchmark, yielding 1,478 BERT-ready samples after leakage-aware preprocessing. Training under a 10-fold GroupKFold protocol stratified by repository identifier to reduce same-repository leakage across training and validation splits yields a mean accuracy of 98.98 % and a mean F1-score of 99.22 %, outperforming all Repo2Vec baseline variants by up to 9.22 percentage points in F1-score. A bagging ensemble of three DistilBERT models further improves the F1-score to 99.75 %. Calibration analysis via temperature scaling indicates well-calibrated probabilistic outputs (ECE = 0.0041 for the cross-validation model and ECE = 0.0074 for the ensemble after scaling), suggesting that the framework is promising for threshold-based screening scenarios. Interpretability and robustness analyses further show that the model’s decisions are driven by distributed contextual signals rather than a handful of overt security terms: malware recall remains at 99.18 % even after masking all 22 such keywords, providing direct evidence against a keyword-shortcut explanation of the model’s performance. To the best of our knowledge, no prior work has applied a BERT-family transformer model to supervised repository-level malware-versus-benign binary classification using metadata alone. The main limitation is that the benign benchmark is restricted to Java-language repositories, whereas the malware repositories span multiple languages.
This thesis rebuilds Real-Vul through a controlled Code Property Graph pipeline, measure and remove a 37% content leak intrinsic to whole-codebase sampling, and train a relational graph neural network with a disciplined class-imbalance recipe: focal loss, class-aware undersampling, and fine-tuning of a GraphCodeBERT no...
This work presents a deep learning approach for multilingual vulnerability detection that emphasizes semantic transfer rather than architectural complexity and suggests that stabilizing semantic representations during transfer is key to improving generalization while maintaining practical efficiency under moderate comp...
Tuan Nguyen Kim, Nin Ho Le Viet, Chieu Ta Quang· Engineering, Technology &...· 0 citations
Malicious packages in open source-software supply chains are a growing security concern, yet machine learning detectors built on registry metadata are difficult to interpret and are typically evaluated under protocols susceptible to data leakage. We construct a dataset of 3330 package versions from NPM and PyPI in whic...
Hanan Moufid, M. El Ghazouani, Moulay Ahmed el Kiram· Journal of Cybersecurity and...· 0 citations
Android's widespread adoption and open ecosystem make it a primary target for malware, a challenge exacerbated by internet fragmentation resulting in non-stationary data distributions across regions. This work presents AuthProtect, a scalable malware detection framework based on incremental learning and a novel permiss...
Maksim Iavich, Razvan Bocu· International Journal of Eng...· 0 citations