A machine learning-based decision support system for multiclass defect detection, utilizing exclusively heterogeneous legacy process data to minimize manual inspection time in foundries, and a Naive Bayes Stacking meta-classifier effectively neutralizes single-algorithm inductive biases.
Abstract
This paper presents a machine learning-based decision support system for multiclass defect detection, utilizing exclusively heterogeneous legacy process data to minimize manual inspection time in foundries. Validated on 51,377 products across 193 defect categories, the methodology resolves structural data inconsistencies through k-nearest neighbor (kNN) imputation and piecewise winsorization. A multi-stage feature selection cascade, incorporating variance thresholding, correlation filtering, and Random Forest Feature Importance (RFFI), reduces the feature space from 139 to 78 process-critical variables. Following Synthetic Minority Over-sampling Technique (SMOTE)-based class balancing, five classifiers were benchmarked via 10-fold cross-validation and optimized using the macro-averaged F3-score to mathematically penalize undetected defects. Light Gradient Boosting Machine (LightGBM) and Random Forest (RF) provided superior predictive baselines. To enforce strict zero-defect constraints, an asymmetric risk function shifted decision boundaries, enabling the risk-calibrated LightGBM model to reduce manual inspection volume by 9.72% with zero defect escapes. For resolving conflicting predictions, multi-algorithm decision fusion was implemented. By statistically evaluating the joint probabilities of the base models’ post-calibration outputs, a Naive Bayes Stacking meta-classifier effectively neutralizes single-algorithm inductive biases. Ultimately, synthesizing these F3-optimized, risk-calibrated base models via meta-learning successfully isolated true defect-free components, maximizing the final inspection time reduction to 14.82% while strictly maintaining zero defect escapes.
Software defect prediction (SDP) is essential for improving software quality since it finds error-prone modules early in the development lifecycle. Current methods produce inflated and erroneous performance metrics because of data leaks, inadequate class imbalance management, and reliance on antiquated classifiers. By...
B. V. Chowdary, Sendhil Kumar B. B, D. L. Sri et al.· International Conference on...· 0 citations
This study compared no correction, random oversampling, random undersampling, SMOTE, ADASYN, and class-weighted learning across logistic regression, decision tree, random forest, support vector machine, and neural network classifiers to find accuracy alone is unsuitable for selecting defect predictors.
L. Akpan· International Journal of App...· 0 citations
In smart city software systems, where interconnected services demand high reliability, Software Defect Prediction (SDP) plays a vital role and reducing maintenance costs by identifying defect-prone modules early in the Software Development Life Cycle (SDLC). Cross-Project Defect Prediction (CPDP) enables defect data fr...
Emediong Bassey Obot, Victor Anaga, Sadiq Thomas et al.· E3S Web of Conferences· 0 citations
Purpose: This study developed a Decision Support System (DSS) for fraud prediction using a Genetic Support Vector Machine (GSVM), a hybrid model that employs a Genetic Algorithm (GA) to optimize Support Vector Machine (SVM) hyperparameters (C and γ) on financial ratios extracted from MachameRatios®.
Research Methodolog...
C. Egbunike, C. Onyali, K. Okafor· International Journal of Fin...· 0 citations
Financial fraud detection is a critical challenge in modern banking systems, where fraudulent
transactions represent less than 0.13% of total transactions, creating severe class imbalance.
Traditional rule-based systems struggle to adapt to evolving fraud patterns, necessitating
machine learning approaches that can...
Richard Chafukira Phiri· IIARD INTERNATIONAL JOURNAL...· 0 citations
This paper develops and evaluates an early-warning credit risk classifier using a three-class formulation that distinguishes performing loans (L), delinquent accounts (DP), and non-performing loans (NPL). Early-warning modeling is challenging because deterioration events are relatively infrequent, yielding class imbala...