Bridging the data latency gap: automated extraction of genomic biomarkers from unstructured clinical documents to support real-world oncology data.
Real-world oncology data are essential for clinical research and precision cancer care. However, genomic biomarkers are often embedded in scanned, unstructured clinical documents requiring manual abstraction before becoming available in cancer registries, delaying real-world evidence generation. This study evaluated and compared three open-source optical character recognition (OCR) approaches, Tesseract, EasyOCR, and a hybrid implementation, to determine which best enables automated extraction of Oncotype DX recurrence scores and improves the timeliness and quality of real-world oncology data.
We evaluated the feasibility of automated genomic data extraction using 675 Oncotype DX reports from a Midwestern U.S. health system. EasyOCR, Tesseract, and a hybrid OCR approach were used to extract recurrence scores from scanned reports. OCR-derived values were compared with manually abstracted scores and local cancer registry data. Performance was assessed using agreement, precision, recall, F1 score, and processing time. Multivariable logistic regression was performed to identify factors associated with discordance between registry-reported and manually abstracted scores.
The hybrid OCR approach demonstrated the highest performance, achieving 97% agreement with manual abstraction, precision of 0.997, recall of 0.972, and an F1 score of 0.984. Registry abstraction demonstrated comparable performance but required greater manual effort. Automated extraction substantially reduced processing time while maintaining high accuracy. Logistic regression showed registry discordance was largely independent of patient and tumor characteristics, with unknown progesterone receptor (PR) status as the only significant predictor.
Automated extraction of genomic biomarkers represents a scalable approach to reducing delays in cancer data availability. Earlier capture of genomic information may support cancer registry modernization and improve real-world evidence generation in precision oncology.
We evaluated the feasibility of automated genomic data extraction using 675 Oncotype DX reports from a Midwestern U.S. health system. EasyOCR, Tesseract, and a hybrid OCR approach were used to extract recurrence scores from scanned reports. OCR-derived values were compared with manually abstracted scores and local cancer registry data. Performance was assessed using agreement, precision, recall, F1 score, and processing time. Multivariable logistic regression was performed to identify factors associated with discordance between registry-reported and manually abstracted scores.
The hybrid OCR approach demonstrated the highest performance, achieving 97% agreement with manual abstraction, precision of 0.997, recall of 0.972, and an F1 score of 0.984. Registry abstraction demonstrated comparable performance but required greater manual effort. Automated extraction substantially reduced processing time while maintaining high accuracy. Logistic regression showed registry discordance was largely independent of patient and tumor characteristics, with unknown progesterone receptor (PR) status as the only significant predictor.
Automated extraction of genomic biomarkers represents a scalable approach to reducing delays in cancer data availability. Earlier capture of genomic information may support cancer registry modernization and improve real-world evidence generation in precision oncology.