Comparative performance of AI models and clinicians in evidence-based cardiovascular disease management for people living with HIV: Comparative Study.
Widespread antiretroviral therapy has greatly extended the life expectancy of people living with HIV (PLWH), making cardiovascular disease (CVD) one of their primary comorbidities. Nevertheless, significant cross-specialty knowledge gaps persist in routine clinical practice. Siloed disciplinary expertise results in low clinical adherence to guideline-recommended risk management interventions, highlighting an urgent demand for integrated, evidence-based tools that break down interdisciplinary barriers. Large language models (LLMs) have demonstrated robust medical knowledge retrieval and reasoning capacity in recent years, yet no systematic evaluation has determined whether these models can bridge such knowledge gaps and facilitate multidisciplinary collaborative CVD management for PLWH.
This study compared the performance of four mainstream AI models (Deepseek-V3, Deepseek-R1, ChatGPT-4o, ChatGPT-o4-mini) and 12 human clinicians (8 infectious disease specialists and 4 cardiologists) in addressing guideline-based CVD management tasks for PLWH.
Based on four authoritative domestic and international guidelines on HIV and CVD care, a structured 25-question assessment battery was developed via two rounds of Delphi expert consultation, with standard reference answers and an evaluation framework finalized through expert consensus. Responses of the four LLMs were generated with standardized prompts, while 12 clinicians answered identical questions in one-on-one structured interviews, with all verbal replies transcribed verbatim. Six multidisciplinary experts independently rated all responses across four dimensions: accuracy, completeness, readability and reliability, using a 4-point ordinal scale ranging from 1 (poor) to 4 (excellent). Cumulative link mixed models (CLMMs) were applied to analyze intergroup differences.
All AI models achieved statistically significantly higher scores than clinicians across all evaluation dimensions (p < 0.01). The AI group had mean scores of 3.44-3.68 (median = 4, CV: 0.145-0.178). Restricted by individual factors including specialty background, knowledge reserve, clinical experience, clinicians obtained lower mean scores of 1.78-2.05 (median = 2, CV: 0.428-0.473) with markedly greater score dispersion. Among all AI models, Deepseek-R1 delivered the optimal performance and showed statistically significant advantages over ChatGPT-4o, ChatGPT-o4-mini and Deepseek-V3 (all p < 0.01). Specialty-stratified CLMM analysis revealed no significant overall score difference between cardiologists and infectious disease specialists (OR = 0.92, 95% CI: 0.84-1.01, p = 0.094). Dimension-specific CLMMs combined with Wilcoxon rank-sum tests confirmed that cardiologists only earned significantly higher scores in the accuracy dimension (OR = 0.81, 95% CI: 0.67-0.97, p = 0.0261). Domain-specific performance divergence was observed: cardiologists outperformed infectious disease specialists in CVD risk assessment (2.26 vs 1.83), whereas infectious disease specialists achieved higher scores on drug adverse effect evaluation (2.23 vs 1.65).
This structured Q&A study on CVD management for PLWH found that LLMs outperformed human clinicians on all assessment metrics, with Deepseek-R1 attaining a distinctly superior composite score. The findings support the promising potential of Deepseek-R1 as a cross-disciplinary decision-support tool: it integrates multi-domain complex clinical knowledge, which may help address cross-specialty knowledge barriers, could improve the completeness and precision of clinical information output, and may enhance communication and decision-making efficiency for patients with complicated multimorbidity. To maximize clinical benefits, AI systems should be integrated into multidisciplinary care workflows alongside targeted clinical training to optimize the management of complex comorbidities among PLWH.
This study compared the performance of four mainstream AI models (Deepseek-V3, Deepseek-R1, ChatGPT-4o, ChatGPT-o4-mini) and 12 human clinicians (8 infectious disease specialists and 4 cardiologists) in addressing guideline-based CVD management tasks for PLWH.
Based on four authoritative domestic and international guidelines on HIV and CVD care, a structured 25-question assessment battery was developed via two rounds of Delphi expert consultation, with standard reference answers and an evaluation framework finalized through expert consensus. Responses of the four LLMs were generated with standardized prompts, while 12 clinicians answered identical questions in one-on-one structured interviews, with all verbal replies transcribed verbatim. Six multidisciplinary experts independently rated all responses across four dimensions: accuracy, completeness, readability and reliability, using a 4-point ordinal scale ranging from 1 (poor) to 4 (excellent). Cumulative link mixed models (CLMMs) were applied to analyze intergroup differences.
All AI models achieved statistically significantly higher scores than clinicians across all evaluation dimensions (p < 0.01). The AI group had mean scores of 3.44-3.68 (median = 4, CV: 0.145-0.178). Restricted by individual factors including specialty background, knowledge reserve, clinical experience, clinicians obtained lower mean scores of 1.78-2.05 (median = 2, CV: 0.428-0.473) with markedly greater score dispersion. Among all AI models, Deepseek-R1 delivered the optimal performance and showed statistically significant advantages over ChatGPT-4o, ChatGPT-o4-mini and Deepseek-V3 (all p < 0.01). Specialty-stratified CLMM analysis revealed no significant overall score difference between cardiologists and infectious disease specialists (OR = 0.92, 95% CI: 0.84-1.01, p = 0.094). Dimension-specific CLMMs combined with Wilcoxon rank-sum tests confirmed that cardiologists only earned significantly higher scores in the accuracy dimension (OR = 0.81, 95% CI: 0.67-0.97, p = 0.0261). Domain-specific performance divergence was observed: cardiologists outperformed infectious disease specialists in CVD risk assessment (2.26 vs 1.83), whereas infectious disease specialists achieved higher scores on drug adverse effect evaluation (2.23 vs 1.65).
This structured Q&A study on CVD management for PLWH found that LLMs outperformed human clinicians on all assessment metrics, with Deepseek-R1 attaining a distinctly superior composite score. The findings support the promising potential of Deepseek-R1 as a cross-disciplinary decision-support tool: it integrates multi-domain complex clinical knowledge, which may help address cross-specialty knowledge barriers, could improve the completeness and precision of clinical information output, and may enhance communication and decision-making efficiency for patients with complicated multimorbidity. To maximize clinical benefits, AI systems should be integrated into multidisciplinary care workflows alongside targeted clinical training to optimize the management of complex comorbidities among PLWH.