Imbalance-aware breast cancer classification using SMOTE, ADASYN, and cost-sensitive ensemble learning on the Wisconsin diagnostic dataset
Annotatsiya
One of the major and commonly overlooked sources of bias in the Wisconsin Diagnostic Breast Cancer (WDBC) data is due to class imbalance. The dataset is moderately imbalanced (357 benign cases vs. 212 malignant cases; ratio is approximately 1.68:1), but due to the impact that a false negative would have in the clinic, imbalance aware learning is crucial. This study compares four imbalance-handling methods—baseline (no method), Synthetic Minority Over-sampling Technique (SMOTE), Adaptive Synthetic Sampling (ADASYN), and cost-sensitive class-weighted learning—and five machine learning classifiers: logistic regression, k-nearest neighbors (kNN), support vector machine (RBF kernel), random forest (RF), and gradient boosting (GB), with soft voting, stacking, and bagging ensembles. The stratified 10-fold cross validation was used and resampling was done only within the train folds with a leakage-free design. Ten complementary performance metrics were used to evaluate model performance: accuracy, precision, recall, specificity, F1-score, Matthews correlation coefficient (MCC), ROC-AUC, PR-AUC, balanced accuracy and Brier score. To bolster the analysis, three additional experiments were performed: (i) controlled imbalance-ratio experiments (IR = 1.68, 5, 10, and 20), (ii) re-implementation of the best published baseline of the WDBC random forest on the same partitions, and (iii) Holm–Bonferroni-corrected significance testing with 1000-bootstrap confidence intervals and clinical decision-curve analysis. For the native imbalance setting, soft voting ensemble with SMOTE-balanced classifiers (F1 = 0.9880, MCC = 0.9812) performed best, outperforming the re-implemented RF baseline (F1 = 0.9407, MCC = 0.9079). The most powerful malignancy indicators were the worst area, perimeter and concave points features, which were identified by both SHAP and permutation analysis. The proposed framework is a repeatable, understandable, and multi-IR-validated benchmark pipeline for the biomedical tabular classification task when the data is imbalanced.
Hali tarjima qilinmagan