Evaluating large language models using the Type 2 Diabetes Health Education guideline: a comparative analysis of ChatGPT-4.1, Claude-4.0, DeepSeek-V3, and ERNIE Bot 4.5 Turbo.
To systematically evaluate the quality and readability of health information generated by four large language models (LLMs) in response to inquiries regarding type 2 diabetes mellitus (T2DM), using an authoritative Chinese clinical guideline as the reference standard.
A total of 124 standardized questions were extracted from the Chinese Type 2 Diabetes Popular Science Guidelines. Six endocrinologists and diabetes specialists conducted independent, blind evaluations using the CLEAR tool (Completeness, Lack of false Information, Evidence, Appropriateness, Relevance) and PEMAT-P (Patient Education Materials Assessment Tool for Printable materials). Response characteristics were also recorded. Between-model differences were tested using the Kruskal-Wallis H test with Bonferroni pairwise comparisons.
All four models achieved total CLEAR scores within the "very good" range (19-25), with no significant differences seen between models (χ2 = 1.985, p = 0.576). No significant differences were observed in the dimensions of Lack of false information (χ2 = 7.644, p = 0.054), Evidence (χ2 = 2.309, p = 0.511), and Relevance (χ2 = 7.516, p = 0.057). However, significant differences emerged in Completeness (χ2 = 47.661, p < 0.001) and Appropriateness (χ2 = 88.360, p < 0.001). Claude-4.0 received the lowest score in Completeness (median 4.00, IQR 3.00-5.00) but achieved the highest ranking in Appropriateness (median 4.00, IQR 4.00-5.00). On the PEMAT-P, understandability differed significantly across models (χ2 = 159.120, p < 0.001), yet all models surpassed the 70% threshold, with ChatGPT-4.1 highest (median 91.91%, IQR 91.91-100.00%). However, despite significant differences among the various models (χ2 = 354.023, p < 0.001), only ERNIE Bot 4.5 Turbo (median 75.00%, IQR75.00-75.00%) surpassed the 70% threshold, with no single model demonstrating consistent superiority across all dimensions.
Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education. Future development should prioritize stronger step-by-step behavioral guidance and differentiated, scenario-specific model deployment to enhance their value in patient-facing diabetes self-management support.
A total of 124 standardized questions were extracted from the Chinese Type 2 Diabetes Popular Science Guidelines. Six endocrinologists and diabetes specialists conducted independent, blind evaluations using the CLEAR tool (Completeness, Lack of false Information, Evidence, Appropriateness, Relevance) and PEMAT-P (Patient Education Materials Assessment Tool for Printable materials). Response characteristics were also recorded. Between-model differences were tested using the Kruskal-Wallis H test with Bonferroni pairwise comparisons.
All four models achieved total CLEAR scores within the "very good" range (19-25), with no significant differences seen between models (χ2 = 1.985, p = 0.576). No significant differences were observed in the dimensions of Lack of false information (χ2 = 7.644, p = 0.054), Evidence (χ2 = 2.309, p = 0.511), and Relevance (χ2 = 7.516, p = 0.057). However, significant differences emerged in Completeness (χ2 = 47.661, p < 0.001) and Appropriateness (χ2 = 88.360, p < 0.001). Claude-4.0 received the lowest score in Completeness (median 4.00, IQR 3.00-5.00) but achieved the highest ranking in Appropriateness (median 4.00, IQR 4.00-5.00). On the PEMAT-P, understandability differed significantly across models (χ2 = 159.120, p < 0.001), yet all models surpassed the 70% threshold, with ChatGPT-4.1 highest (median 91.91%, IQR 91.91-100.00%). However, despite significant differences among the various models (χ2 = 354.023, p < 0.001), only ERNIE Bot 4.5 Turbo (median 75.00%, IQR75.00-75.00%) surpassed the 70% threshold, with no single model demonstrating consistent superiority across all dimensions.
Although the four LLMs generally provide accurate and pertinent information regarding type 2 diabetes, enduring limits in actionability and inconsistencies among models in content completeness and understandability restrict their effective use in diabetic patient education. Future development should prioritize stronger step-by-step behavioral guidance and differentiated, scenario-specific model deployment to enhance their value in patient-facing diabetes self-management support.
Authors
Huang Huang, Bai Bai, Luo Luo, Zhang Zhang, Hu Hu, Hu Hu, Ruan Ruan, Zeng Zeng, Jia Jia, Deng Deng, Xiong Xiong, Zhou Zhou, Qi Qi, Lyu Lyu, Li Li
View on Pubmed