Comparative evaluation of active learning, random sampling, and deep learning for smart contract vulnerability detection
Abstract
Smart contracts are widely used to automate blockchain-based transactions and decentralised services, but vulnerabilities in their code can expose users and systems to financial loss, service disruption, and security compromise. Detecting vulnerable smart contracts remains challenging because reliable labelling often requires expert analysis, which makes fully labelled datasets costly to produce. This study investigates whether uncertainty-based active learning can reduce the labelling requirements for smart contract vulnerability detection using machine-learning classifiers. The BCCC-VulSCs-2023 dataset, containing 36,670 smart contract samples with 73 raw feature columns and a binary vulnerability label, was used for model development and evaluation. The dataset was preprocessed by removing constant and irrelevant columns, handling missing values, class balancing via SMOTE, an 80/20 train-test split, SelectKBest feature selection (top 14 features), and numerical feature normalisation. Random Forest, XGBoost, and CatBoost classifiers were evaluated under three conditions: uncertainty-based active learning, a random-sampling baseline with an identical labelling schedule, and a fully supervised baseline trained on the complete training set. LSTM, DNN, and GRU models were implemented as deep learning baselines. Under a fixed labelling budget (1,000 initial samples plus 4,000 samples queried per iteration over 10 iterations, totaling 41,000 of 43,062 available training labels), Random Forest and CatBoost with active learning achieved the strongest performance (73.7% and 73.9% test accuracy, respectively), outperforming XGBoost with active learning (69.1% accuracy) and all three deep learning baselines (66.7%–68.2% accuracy). However, under this budget, active learning showed no consistent advantage over the random-sampling baseline: test accuracy differed by less than 0.3 percentage points between active learning and random sampling across all three classifiers, and both approached the accuracy of the fully supervised baseline trained on the complete label set. These findings indicate that ensemble-based classifiers (Random Forest, CatBoost) are better suited to this structured smart contract feature representation than the evaluated deep learning baselines. However, at the labelling budget evaluated in this study, uncertainty-based active learning did not provide a measurable advantage over random sampling, suggesting that its benefit, if present, may only emerge at smaller labelling budgets than those tested here.