A machine learning-based decision support system for multiclass defect detection, utilizing exclusively heterogeneous legacy process data to minimize manual inspection time in foundries, and a Naive Bayes Stacking meta-classifier effectively neutralizes single-algorithm inductive biases.
Abstract
This paper presents a machine learning-based decision support system for multiclass defect detection, utilizing exclusively heterogeneous legacy process data to minimize manual inspection time in foundries. Validated on 51,377 products across 193 defect categories, the methodology resolves structural data inconsistencies through k-nearest neighbor (kNN) imputation and piecewise winsorization. A multi-stage feature selection cascade, incorporating variance thresholding, correlation filtering, and Random Forest Feature Importance (RFFI), reduces the feature space from 139 to 78 process-critical variables. Following Synthetic Minority Over-sampling Technique (SMOTE)-based class balancing, five classifiers were benchmarked via 10-fold cross-validation and optimized using the macro-averaged F3-score to mathematically penalize undetected defects. Light Gradient Boosting Machine (LightGBM) and Random Forest (RF) provided superior predictive baselines. To enforce strict zero-defect constraints, an asymmetric risk function shifted decision boundaries, enabling the risk-calibrated LightGBM model to reduce manual inspection volume by 9.72% with zero defect escapes. For resolving conflicting predictions, multi-algorithm decision fusion was implemented. By statistically evaluating the joint probabilities of the base models’ post-calibration outputs, a Naive Bayes Stacking meta-classifier effectively neutralizes single-algorithm inductive biases. Ultimately, synthesizing these F3-optimized, risk-calibrated base models via meta-learning successfully isolated true defect-free components, maximizing the final inspection time reduction to 14.82% while strictly maintaining zero defect escapes.
Software defect prediction (SDP) is essential for improving software quality since it finds error-prone modules early in the development lifecycle. Current methods produce inflated and erroneous performance metrics because of data leaks, inadequate class imbalance management, and reliance on antiquated classifiers. By...
B. V. Chowdary, Sendhil Kumar B. B, D. L. Sri et al.· International Conference on...· 0 citations
This study compared no correction, random oversampling, random undersampling, SMOTE, ADASYN, and class-weighted learning across logistic regression, decision tree, random forest, support vector machine, and neural network classifiers to find accuracy alone is unsuitable for selecting defect predictors.
L. Akpan· International Journal of App...· 0 citations
This study proposes an integrated CPDP framework that combines Particle Swarm Optimization with Domain Knowledge (PSO+DK) for feature selection and Adaptive Synthetic Sampling (ADASYN) for class imbalance handling that enhanced the discriminative power of the models.
Emediong Bassey Obot, Victor Anaga, Sadiq Thomas et al.· E3S Web of Conferences· 0 citations
Purpose: This study developed a Decision Support System (DSS) for fraud prediction using a Genetic Support Vector Machine (GSVM), a hybrid model that employs a Genetic Algorithm (GA) to optimize Support Vector Machine (SVM) hyperparameters (C and γ) on financial ratios extracted from MachameRatios®.
Research Methodolog...
C. Egbunike, C. Onyali, K. Okafor· International Journal of Fin...· 0 citations
Empirical evidence is provided that for fraud detection with well-engineered features in low-dimensional spaces, ensemble methods yield optimal results and the undersampling strategy proves effective for handling class imbalance while maintaining computational efficiency.
Richard Chafukira Phiri· IIARD INTERNATIONAL JOURNAL...· 0 citations
This paper develops and evaluates an early-warning credit risk classifier using a three-class formulation that distinguishes performing loans (L), delinquent accounts (DP), and non-performing loans (NPL). Early-warning modeling is challenging because deterioration events are relatively infrequent, yielding class imbala...