AI-Enabled Modeling for Alzheimer's Disease Risk Prediction and Validation.
To investigate the multimodal clinical influencing factors of Alzheimer's disease (AD) onset, and to establish and test a risk prediction tool derived from these determinants and additional clinical measures, thus facilitating early intervention and risk classification in individuals at high risk for AD.
A retrospective cohort of 502 high-risk individuals for AD (exhibiting cognitive decline or family history) who visited our hospital was included. A total of 502 participants were randomly split into a training cohort (n = 350) and a validation cohort (n = 152) in a 7:3 proportion. Demographic characteristics, clinical indicators, biomarkers, and genetic markers were collected. In the training set, univariate analysis and least absolute shrinkage and selection operator (LASSO) regression were first applied for variable screening, followed by multivariate logistic regression to pinpoint independent influencing factors. Random forest (RF), XGBoost, and deep learning models were constructed using Python, with performance evaluated using area under the curve (AUC). The optimal model was selected, and feature importance was analyzed.
Between the training and validation sets, no statistically significant baseline characteristic differences were found (p > 0.05). Multivariate logistic regression identified the apolipoprotein E epsilon 4 allele (APOE ε4) genotype, cerebrospinal fluid (CSF) p-tau181/amyloid-beta 42 (Aβ42) ratio, and diabetes as independent risk factors for AD (p < 0.05), while serum folate levels, Mini-Mental State Examination (MMSE) scores, and Montreal Cognitive Assessment (MoCA) scores served as independent protective factors (p < 0.05). In the validation set, the RF model achieved the highest AUC (0.879), followed by XGBoost (0.869) and deep learning (0.844), with the CSF p-tau181/Aβ42 ratio identified as the most predictive feature.
The RF model, based on integrated multimodal clinical influencing factors and clinical indicators, demonstrates potential for AD risk stratification in high-risk populations when evaluated on a validation cohort.
A retrospective cohort of 502 high-risk individuals for AD (exhibiting cognitive decline or family history) who visited our hospital was included. A total of 502 participants were randomly split into a training cohort (n = 350) and a validation cohort (n = 152) in a 7:3 proportion. Demographic characteristics, clinical indicators, biomarkers, and genetic markers were collected. In the training set, univariate analysis and least absolute shrinkage and selection operator (LASSO) regression were first applied for variable screening, followed by multivariate logistic regression to pinpoint independent influencing factors. Random forest (RF), XGBoost, and deep learning models were constructed using Python, with performance evaluated using area under the curve (AUC). The optimal model was selected, and feature importance was analyzed.
Between the training and validation sets, no statistically significant baseline characteristic differences were found (p > 0.05). Multivariate logistic regression identified the apolipoprotein E epsilon 4 allele (APOE ε4) genotype, cerebrospinal fluid (CSF) p-tau181/amyloid-beta 42 (Aβ42) ratio, and diabetes as independent risk factors for AD (p < 0.05), while serum folate levels, Mini-Mental State Examination (MMSE) scores, and Montreal Cognitive Assessment (MoCA) scores served as independent protective factors (p < 0.05). In the validation set, the RF model achieved the highest AUC (0.879), followed by XGBoost (0.869) and deep learning (0.844), with the CSF p-tau181/Aβ42 ratio identified as the most predictive feature.
The RF model, based on integrated multimodal clinical influencing factors and clinical indicators, demonstrates potential for AD risk stratification in high-risk populations when evaluated on a validation cohort.