Development and Validation of an Interpretable Machine Learning-Based Clinical Prediction Model for Short-Term Mortality in Intracerebral Hemorrhage With Thrombocytopenia: A Multicenter Study.
Intracerebral hemorrhage (ICH) with thrombocytopenia is associated with poor outcomes, but early risk prediction tools for this subgroup are limited. We aimed to develop and externally validate an interpretable machine learning model for predicting 28-day all-cause mortality after ICU admission.
Internal data were derived from MIMIC-III/IV, eICU, and NWICU, and external validation data from a single tertiary academic teaching hospital in China. The internal cohort was randomly split 7:3 into internal training set and internal test set. Feature selection, hyperparameter optimization, and training of five machine learning models were performed in the internal training set. Performance was evaluated in the internal test set and external validation cohort. SHAP analysis was used for interpretation, and a web-based tool was developed.
Among 1190 included patients, 859 were in the internal cohort and 331 in the external validation cohort. Fifteen predictors were retained. LightGBM showed the best performance, with AUROCs of 0.840 and 0.764 in the internal test and external validation cohorts, respectively. Important predictors included GCS, diastolic blood pressure, glucose, and platelet count.
The developed LightGBM model showed good internal discrimination, acceptable external discrimination, and interpretable feature contributions, supporting its potential as a complementary early risk stratification aid for patients with ICH and thrombocytopenia. Prospective multicenter validation is required before clinical implementation.
Internal data were derived from MIMIC-III/IV, eICU, and NWICU, and external validation data from a single tertiary academic teaching hospital in China. The internal cohort was randomly split 7:3 into internal training set and internal test set. Feature selection, hyperparameter optimization, and training of five machine learning models were performed in the internal training set. Performance was evaluated in the internal test set and external validation cohort. SHAP analysis was used for interpretation, and a web-based tool was developed.
Among 1190 included patients, 859 were in the internal cohort and 331 in the external validation cohort. Fifteen predictors were retained. LightGBM showed the best performance, with AUROCs of 0.840 and 0.764 in the internal test and external validation cohorts, respectively. Important predictors included GCS, diastolic blood pressure, glucose, and platelet count.
The developed LightGBM model showed good internal discrimination, acceptable external discrimination, and interpretable feature contributions, supporting its potential as a complementary early risk stratification aid for patients with ICH and thrombocytopenia. Prospective multicenter validation is required before clinical implementation.