A machine learning model for distinguishing gastric cancer from intestinal metaplasia in patients with psychological symptoms.
To develop and internally evaluate an interpretable machine learning model for distinguishing current gastric cancer (GC) from intestinal metaplasia (IM) among patients with psychological symptoms.
This retrospective, single-center, cross-sectional study included 302 patients with psychological symptoms, comprising 185 with IM and 117 with GC. Patients were randomly divided into a training cohort (n = 212) and an independently retained validation cohort (n = 90). Candidate features were evaluated using LASSO, Boruta, and recursive feature elimination. Seven machine learning algorithms were compared using nested cross-validation in the training cohort. The final model was assessed in the validation cohort using discrimination, calibration, decision curve analysis, and Shapley Additive Explanations (SHAP).
Seven features were retained: age, albumin, sex, psychological symptom type, total bilirubin, smoking status, and monocyte count. XGBoost achieved the highest mean outer-fold AUC (0.839 ± 0.029) and was selected as the final model. In the validation cohort, XGBoost achieved an AUC of 0.793 (95% CI 0.670-0.895), accuracy of 0.800, sensitivity of 0.657, specificity of 0.891, and F1 score of 0.719. The Brier score was 0.167. Decision curve analysis suggested potential clinical net benefit. SHAP analysis identified age and psychological symptom type as the leading contributors to classification.
The interpretable XGBoost model showed moderate discrimination for distinguishing current GC from IM and may provide supplementary information for clinical assessment. It should not replace endoscopic or histopathological diagnosis, and external multicenter validation is required before broader clinical implementation.
This retrospective, single-center, cross-sectional study included 302 patients with psychological symptoms, comprising 185 with IM and 117 with GC. Patients were randomly divided into a training cohort (n = 212) and an independently retained validation cohort (n = 90). Candidate features were evaluated using LASSO, Boruta, and recursive feature elimination. Seven machine learning algorithms were compared using nested cross-validation in the training cohort. The final model was assessed in the validation cohort using discrimination, calibration, decision curve analysis, and Shapley Additive Explanations (SHAP).
Seven features were retained: age, albumin, sex, psychological symptom type, total bilirubin, smoking status, and monocyte count. XGBoost achieved the highest mean outer-fold AUC (0.839 ± 0.029) and was selected as the final model. In the validation cohort, XGBoost achieved an AUC of 0.793 (95% CI 0.670-0.895), accuracy of 0.800, sensitivity of 0.657, specificity of 0.891, and F1 score of 0.719. The Brier score was 0.167. Decision curve analysis suggested potential clinical net benefit. SHAP analysis identified age and psychological symptom type as the leading contributors to classification.
The interpretable XGBoost model showed moderate discrimination for distinguishing current GC from IM and may provide supplementary information for clinical assessment. It should not replace endoscopic or histopathological diagnosis, and external multicenter validation is required before broader clinical implementation.