Reliable auto insurance fraud detection using boosting and deep learning models through comprehensive predictive performance, calibration, statistical significance, and economic impact
Insurance fraud detection poses a key challenge due to substantial class imbalance, heterogeneous claim types, and the continual evolution of fraudulent practices. Despite several Machine Learning (ML) approaches having been developed, comparative assessments that simultaneously address predictive performance, calibration stability, cost-effectiveness, and comprehensive statistical analyses remain scarce. To bridge the gap, this study presents a robust evaluation framework for identifying auto insurance fraud that incorporates boosting-based classifiers, advanced deep tabular architectures, and adaptive resampling methods. Six classification models -namely Categorical Boosting (CatBoost), Light Gradient Boosting Machine (LightGBM), eXtreme Gradient Boosting (XGBoost), Attentive Interpretable Tabular Learning Architecture (TabNet), Feature Tokenizer Transformer (FT-Transformer), and Multi-Layer Perceptron Residual Network (MLP-ResNet)- were tested under three data balancing methods (Original, SMOTE, and ADASYN). Contrary to previous studies that focus exclusively on classification performance, the proposed framework includes sensitivity analysis, probabilistic calibration scoring, effect-size evaluation, expected-cost analysis, Pareto frontier optimization, and statistical significance tests. The comparative evaluation reveals that no single model systematically outperforms all others. CatBoost with ADASYN provides the most balanced predictive performance, while TabNet presents superior probabilistic prediction quality and the lowest expected cost. These findings show that the optimal model depends on the target objective, whether predictive accuracy or cost-sensitive fraud detection, a result further validated by the Friedman and Nemenyi statistical tests. In addition, the Pareto frontier analysis identifies CatBoost and TabNet as supplementary optimal trade-offs between predictive performance and operational cost. Statistical analysis confirms significant differences among the evaluated configurations, attesting to the robustness of the findings. Overall, the proposed framework highlights the complementary strengths of boosting-based approaches and TabNet, offering a reliable and cost-effective method for insurance fraud detection, depending on the targeted operational objective.