Development and validation of an interpretable machine learning model for predicting hypertension risk in patients with psoriasis.
Hypertension is a common yet frequently underdiagnosed comorbidity in psoriasis patients. Early identification and blood pressure control are critical to improving outcomes. Although machine learning (ML) is widely used in disease prediction, a model for hypertension risk within the psoriasis population remains unavailable. This study aims to develop and validate such a model in patients with psoriasis.
In this retrospective study, 2,957 psoriasis patients from a single tertiary center were used for model development and internal validation, and 567 psoriasis participants from the National Health and Nutrition Examination Survey (NHANES) served as the external validation cohort. After missForest imputation and consensus feature selection, the Synthetic Minority Oversampling Technique (SMOTE) was applied to the training set only. Nine machine learning algorithms were trained and evaluated for discrimination, calibration, and clinical utility. Shapley Additive Explanations (SHAP) were used for model interpretation.
Nine nonredundant predictors were retained. The SMOTE-enhanced logistic regression model showed the most balanced and generalizable performance, with area under the receiver operating characteristic curve values of 0.850, 0.816, and 0.789 in the training, internal validation, and external validation cohorts, respectively, together with acceptable calibration and favorable net clinical benefit. SHAP identified age, dyslipidemia, and type 2 diabetes mellitus as the leading contributors. The final model was deployed as a publicly accessible web application.
This interpretable and externally validated machine learning model provides a practical tool for hypertension risk stratification in psoriasis patients and may support earlier identification and individualized preventive management in clinical practice.
In this retrospective study, 2,957 psoriasis patients from a single tertiary center were used for model development and internal validation, and 567 psoriasis participants from the National Health and Nutrition Examination Survey (NHANES) served as the external validation cohort. After missForest imputation and consensus feature selection, the Synthetic Minority Oversampling Technique (SMOTE) was applied to the training set only. Nine machine learning algorithms were trained and evaluated for discrimination, calibration, and clinical utility. Shapley Additive Explanations (SHAP) were used for model interpretation.
Nine nonredundant predictors were retained. The SMOTE-enhanced logistic regression model showed the most balanced and generalizable performance, with area under the receiver operating characteristic curve values of 0.850, 0.816, and 0.789 in the training, internal validation, and external validation cohorts, respectively, together with acceptable calibration and favorable net clinical benefit. SHAP identified age, dyslipidemia, and type 2 diabetes mellitus as the leading contributors. The final model was deployed as a publicly accessible web application.
This interpretable and externally validated machine learning model provides a practical tool for hypertension risk stratification in psoriasis patients and may support earlier identification and individualized preventive management in clinical practice.