Skip to content
Open access

Machine Learning-Based Electricity Theft Detection in Smart Grids using Class Imbalance Handling and Shapley Additive Explanations (Shap) Explainability

Sep 2026 · Ethiopian International Journal of Engineering and Technology · 0 citations

Abstract

Among the various forms of non-technical losses plaguing power distribution networks, electricity theft stands out as particularly damaging draining utility revenues and undermining grid stability, with the worst effects felt in developing countries where oversight infrastructure is weak. Smart meter data opens a practical path to automated detection, one that scales far better than manual field inspections. Four supervised classifiers were evaluated in this work: Logistic Regression, Random Forest, Support Vector Machine, and XGBoost. All four were tested on the publicly available State Grid Corporation of China (SGCC) dataset, which covers consumption records for 42,372 customers across roughly 1,035 days. Customers missing more than 50% of their daily readings were excluded, leaving 31,188 usable records. Of these, 7.7% were labelled as theft cases — a ratio that closely mirrors real-world distribution network conditions. From each customer's consumption history, nine statistical features were extracted: daily mean, daily standard deviation, daily maximum, daily minimum, load factor, weekday-to-weekend consumption ratio, peak load, load variance, and coefficient of variation. The class imbalance was addressed through three distinct strategies — leaving it uncorrected, applying cost-sensitive class weighting, and generating synthetic minority samples through oversampling. Baseline models hit accuracy figures above 92%, yet their recall stayed below 6%, meaning nearly all theft cases went undetected. Once cost-sensitive weighting was introduced, recall climbed to 58.3% for the Support Vector Machine — a tenfold improvement over baseline — and 51.0% for Logistic Regression, with AUC-ROC values holding steady around 72%. Among tree-based models, XGBoost with class weighting delivered the best recall-F1 balance at 42.3% and 25.0% respectively. SHapley Additive exPlanations (SHAP) analysis on the XGBoost model pointed to peak load, daily maximum consumption, and load variance as the three features with the greatest influence on theft predictions, giving the results a degree of physical interpretability that raw metrics alone cannot provide. Taken together, the results make clear that ignoring class imbalance renders theft detection systems ineffective in practice, and that SHAP explanations are a meaningful bridge between statistical performance and the operational confidence utilities need to act on model outputs. The framework presented here provides a transparent, reproducible starting point for future work in this area. Keywords: Electricity Theft Detection, Smart Grid, Machine Learning, Class Imbalance, SHAP Explainability

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.