Correct but Incomplete: Limitations of AI-Assisted Decision Support in Rectal Cancer.
BackgroundArtificial intelligence (AI) increasingly supports clinical decision making. ChatGPT-5 (OpenAI, San Francisco, CA) and OpenEvidence (OpenEvidence Inc, Miami, FL) are regularly used by physicians, yet the limits of their reliability in decision-making remain poorly defined. This study evaluated both platforms against the National Comprehensive Cancer Network® (NCCN) Guidelines for Rectal Cancer to define where AI tools perform well and where they fall short.MethodsThe NCCN Guidelines for Rectal Cancer (Version 2.2025) were reviewed. Three questions were generated per decision-making page and classified into workup/diagnosis, treatment, and surveillance domains, yielding 138 clinical scenarios. Both platforms were queried. Responses were scored independently by two physicians on a 5-point Likert scale (5 = completely correct; 1 = absolutely incorrect). Two performance thresholds were defined: Correctness (≥3) and Accuracy (≥4). Proportions were compared with Fisher's exact test and score distributions with the Mann-Whitney U test.ResultsBoth platforms demonstrated high overall guideline concordance. ChatGPT-5 achieved Correctness in 136 (98.6%) and Accuracy in 116 (84.1%), compared to 128 (92.8%) and 112 (81.2%) for OpenEvidence, respectively. Overall Correctness favored ChatGPT-5 (P = 0.035), while Accuracy showed no significant difference (P = 0.634). Mean Likert scores were 4.57 vs 4.37 (P = 0.089). Both platforms achieved >90% Correctness across all domains.ConclusionBoth platforms demonstrate reliable identification of the broad direction of care but exhibit important limitations in completeness, nuance, and the handling of preference-sensitive decisions. These findings define the appropriate role of AI-assisted decision support as an adjunct to, rather than a substitute for, multidisciplinary review in rectal cancer management.
Authors
Meyer Meyer, Bresler Bresler, Makaryan Makaryan, Gheissari Gheissari, Mauch Mauch, Fujita Fujita
View on Pubmed