An Integrated Statistical and Machine Learning Approach for Breast Cancer Classification Using Tumor Morphological Features.
Breast cancer is one of the most common cancers and a leading cause of death among women worldwide. Early and accurate classification of tumors as benign or malignant is essential for improving patient outcomes. In recent years, statistical and machine learning methods have been widely used to improve diagnostic accuracy; however, combining these approaches in a single framework remains limited.
A retrospective analysis was conducted using the Breast Cancer Wisconsin Diagnostic Dataset (569 tumor samples). Eight morphological features were analyzed. The dataset was split into training (70%) and testing (30%) sets using stratified random sampling. Logistic regression, random forest, and support vector machine (SVM) models were developed. Hyperparameter tuning was performed using five-fold cross-validation. Model performance was evaluated using accuracy, sensitivity, specificity, precision, F1-score, Matthews correlation coefficient (MCC), and area under the curve (AUC). Logistic regression assumptions and model diagnostics were also assessed.
Radius, texture, smoothness, and concavity were significant predictors of malignancy. Five-fold cross-validation indicated stable model performance. On the test set, logistic regression achieved the highest accuracy (95.3%) and AUC (0.983), followed by SVM (94.1%, AUC = 0.980) and random forest (92.9%, AUC = 0.979). Additional performance metrics, including precision, F1-score, and MCC, also demonstrated strong classification performance. Diagnostic analyses confirmed acceptable model fit and no major violations of logistic regression assumptions. No significant differences in AUC were observed among the models (DeLong test, p > 0.05).
The study demonstrates that an integrated statistical and machine learning approach provides a robust and accurate method for breast cancer classification. Logistic regression showed slightly better performance, whereas machine learning models also achieved comparable results. This approach has strong potential to support early detection and clinical decision-making.
A retrospective analysis was conducted using the Breast Cancer Wisconsin Diagnostic Dataset (569 tumor samples). Eight morphological features were analyzed. The dataset was split into training (70%) and testing (30%) sets using stratified random sampling. Logistic regression, random forest, and support vector machine (SVM) models were developed. Hyperparameter tuning was performed using five-fold cross-validation. Model performance was evaluated using accuracy, sensitivity, specificity, precision, F1-score, Matthews correlation coefficient (MCC), and area under the curve (AUC). Logistic regression assumptions and model diagnostics were also assessed.
Radius, texture, smoothness, and concavity were significant predictors of malignancy. Five-fold cross-validation indicated stable model performance. On the test set, logistic regression achieved the highest accuracy (95.3%) and AUC (0.983), followed by SVM (94.1%, AUC = 0.980) and random forest (92.9%, AUC = 0.979). Additional performance metrics, including precision, F1-score, and MCC, also demonstrated strong classification performance. Diagnostic analyses confirmed acceptable model fit and no major violations of logistic regression assumptions. No significant differences in AUC were observed among the models (DeLong test, p > 0.05).
The study demonstrates that an integrated statistical and machine learning approach provides a robust and accurate method for breast cancer classification. Logistic regression showed slightly better performance, whereas machine learning models also achieved comparable results. This approach has strong potential to support early detection and clinical decision-making.
Authors
Woudneh Woudneh, Shiferaw Shiferaw, Gebrie Gebrie, Nigussie Nigussie, Cherie Cherie, Gebeyehu Gebeyehu
View on Pubmed