Asosiy kontentga oʻtish
Maqola

Imbalance-aware breast cancer classification using SMOTE, ADASYN, and cost-sensitive ensemble learning on the Wisconsin diagnostic dataset

Ahmed Kateb Jumaah Al-NussairiMathematics Department, College of Basic Education, University of Misan, Misan, 62001, IraqMohammed Kadhim RahmaCommunications Engineering Department, Al Mustaqbal University, Hillah, IraqMaytham T. QasimCollege of Health and Medical Technology, Al-Ayen Iraqi University, AUIQ, Thi-QarKhayala MammadovaMedical and Biological Physics Department, Azerbaijan Medical University, Baku, AzerbaijanAhmed Shakir Al‐HitiDept. of Electrical Engineering, Faculty of Engineering, University of Anbar, Ramadi 31001, IraqFarrukh BakhritdinovDepartment of Exact Sciences, Kimyo International University in Tashkent, UzbekistanБаходир Бахтиёр ўғли РахимовTashkent State Medical University, Tashkent, UzbekistanMohammad KhisheImam Khomeini Naval Science University of Nowshahr, Nowshahr, Iran
2026en
ABI

Annotatsiya

One of the major and commonly overlooked sources of bias in the Wisconsin Diagnostic Breast Cancer (WDBC) data is due to class imbalance. The dataset is moderately imbalanced (357 benign cases vs. 212 malignant cases; ratio is approximately 1.68:1), but due to the impact that a false negative would have in the clinic, imbalance aware learning is crucial. This study compares four imbalance-handling methods—baseline (no method), Synthetic Minority Over-sampling Technique (SMOTE), Adaptive Synthetic Sampling (ADASYN), and cost-sensitive class-weighted learning—and five machine learning classifiers: logistic regression, k-nearest neighbors (kNN), support vector machine (RBF kernel), random forest (RF), and gradient boosting (GB), with soft voting, stacking, and bagging ensembles. The stratified 10-fold cross validation was used and resampling was done only within the train folds with a leakage-free design. Ten complementary performance metrics were used to evaluate model performance: accuracy, precision, recall, specificity, F1-score, Matthews correlation coefficient (MCC), ROC-AUC, PR-AUC, balanced accuracy and Brier score. To bolster the analysis, three additional experiments were performed: (i) controlled imbalance-ratio experiments (IR = 1.68, 5, 10, and 20), (ii) re-implementation of the best published baseline of the WDBC random forest on the same partitions, and (iii) Holm–Bonferroni-corrected significance testing with 1000-bootstrap confidence intervals and clinical decision-curve analysis. For the native imbalance setting, soft voting ensemble with SMOTE-balanced classifiers (F1 = 0.9880, MCC = 0.9812) performed best, outperforming the re-implemented RF baseline (F1 = 0.9407, MCC = 0.9079). The most powerful malignancy indicators were the worst area, perimeter and concave points features, which were identified by both SHAP and permutation analysis. The proposed framework is a repeatable, understandable, and multi-IR-validated benchmark pipeline for the biomedical tabular classification task when the data is imbalanced.

Hali tarjima qilinmagan

Identifikatorlar

Iqtiboslar va manbalar