Limits of Short-Term Seismicity Features in Disaster Management: Machine Learning Based Magnitude Classification Along the NAF
Abstract
This study evaluated whether daily earthquake counts and daily mean magnitudes recorded during the 30 days preceding reference earthquakes along the North Anatolian Fault (NAF) contain discriminatory information for separating events of 4.0 ≤ M < 5.0 from those of M ≥ 5.0. AFAD catalogue records from 1990 to February 2026 were used to identify 448 reference earthquakes. For each event, earthquakes occurring within a 50 km radius during the preceding 30 complete local-calendar days were retrieved, while the target day was excluded. A magnitude-of-completeness threshold of M c = 2.8 was applied before feature engineering. The final dataset contained 60 predictors, comprising 30 daily earthquake-count variables and 30 daily mean-magnitude variables. To reduce leakage arising from overlapping retrospective windows and repeated catalogue events, dependency-controlled groups were used in both the model-development and hold-out partitions. Five dependency-controlled cross-validation folds were applied to the 361-observation model-development set, and SMOTE was restricted to the training portion of each fold. Logistic Regression, Support Vector Machine, Multi-Layer Perceptron, Gradient Boosting, and Random Forest were compared using mean cross-validated ROC-AUC. Logistic Regression achieved the highest mean cross-validated ROC-AUC of 0.597. On the 87-observation hold-out set, however, its ROC-AUC was 0.486, recall was 0.167, and F1-score was 0.125. Excluding completely zero-valued retrospective windows did not improve discrimination. SHAP analysis showed that both daily count and mean-magnitude variables contributed to model outputs, but their effects were distributed across the 30-day window and did not form a stable pattern. The findings indicate that these short-term daily seismicity features provide limited and unstable information for distinguishing the two magnitude classes. The SHAP results describe the behaviour of a weakly discriminative model and should not be interpreted as evidence of causal precursors or reliable earthquake-prediction signals.