Diagnostic accuracy and clinical performance of deep learning models for grading diabetic retinopathy: a systematic review and meta-analysis.
Diabetic retinopathy (DR) is a leading cause of preventable visual impairment worldwide, and its precise severity grading is critical for optimizing clinical management. Conventional frameworks, notably the International Clinical Diabetic Retinopathy (ICDR) scale, are often hindered by substantial inter-observer variability and high dependency on specialist expertise. While deep learning (DL) has recently emerged as a transformative approach for automated stratification, a comprehensive synthesis of evidence regarding its diagnostic performance and clinical application remains lacking.
This systematic review and meta-analysis aimed to comprehensively assess the diagnostic accuracy of fundus image-based deep learning models in the grading of diabetic retinopathy.
PubMed, Embase, Web of Science, and the Cochrane Library were systematically searched for relevant studies published up to October 28, 2025. Diagnostic accuracy studies utilizing DL algorithms alongside ICDR criteria for diabetic retinopathy grading were included. Literature screening and data extraction were performed independently by two researchers, and the risk of bias was assessed using the QUADAS-2 tool.
A total of 41 studies were included, encompassing various DL architectures and multiple public and private fundus image datasets. In the five-class classification task based on ICDR criteria, the pooled sensitivities of DL-based models varied significantly across severity levels: 95.19% (95% CI: 93.00%-97.00%) for no DR (stage 0), 72.06% (95% CI: 62.06%-81.09%) for mild NPDR (stage 1), 84.33% (95% CI: 78.90%-89.10%) for moderate NPDR (stage 2), 75.84% (95% CI: 68.42%-82.57%) for severe NPDR (stage 3), and 78.82% (95% CI: 71.76%-85.13%) for PDR (stage 4). In the simplified four-class classification task, sensitivities markedly improved across all grades: 96.85% (95% CI: 90.18%-99.93%) for stage 0, 92.94% (95% CI: 79.50%-99.72%) for stage 1, 92.75% (95% CI: 79.31%-99.61%) for stage 2, and 88.19% (95% CI: 68.99%-98.93%) for stage 3.
DL exhibits high sensitivity and substantial potential for DR grading, particularly in screening for no DR and vision-threatening DR. Nevertheless, precisely differentiating between adjacent non-proliferative stages remains a clinical challenge. The observed heterogeneity underscores the imperative for methodological standardization, rigorous external validation, and multimodal data integration. Future research should prioritize enhancing clinical utility and generalizability to facilitate their translation into real-world clinical practice.
https://www.crd.york.ac.uk/PROSPERO/, identifier CRD420261338867.
This systematic review and meta-analysis aimed to comprehensively assess the diagnostic accuracy of fundus image-based deep learning models in the grading of diabetic retinopathy.
PubMed, Embase, Web of Science, and the Cochrane Library were systematically searched for relevant studies published up to October 28, 2025. Diagnostic accuracy studies utilizing DL algorithms alongside ICDR criteria for diabetic retinopathy grading were included. Literature screening and data extraction were performed independently by two researchers, and the risk of bias was assessed using the QUADAS-2 tool.
A total of 41 studies were included, encompassing various DL architectures and multiple public and private fundus image datasets. In the five-class classification task based on ICDR criteria, the pooled sensitivities of DL-based models varied significantly across severity levels: 95.19% (95% CI: 93.00%-97.00%) for no DR (stage 0), 72.06% (95% CI: 62.06%-81.09%) for mild NPDR (stage 1), 84.33% (95% CI: 78.90%-89.10%) for moderate NPDR (stage 2), 75.84% (95% CI: 68.42%-82.57%) for severe NPDR (stage 3), and 78.82% (95% CI: 71.76%-85.13%) for PDR (stage 4). In the simplified four-class classification task, sensitivities markedly improved across all grades: 96.85% (95% CI: 90.18%-99.93%) for stage 0, 92.94% (95% CI: 79.50%-99.72%) for stage 1, 92.75% (95% CI: 79.31%-99.61%) for stage 2, and 88.19% (95% CI: 68.99%-98.93%) for stage 3.
DL exhibits high sensitivity and substantial potential for DR grading, particularly in screening for no DR and vision-threatening DR. Nevertheless, precisely differentiating between adjacent non-proliferative stages remains a clinical challenge. The observed heterogeneity underscores the imperative for methodological standardization, rigorous external validation, and multimodal data integration. Future research should prioritize enhancing clinical utility and generalizability to facilitate their translation into real-world clinical practice.
https://www.crd.york.ac.uk/PROSPERO/, identifier CRD420261338867.