OmniPathoVQA: Benchmarking pathology vision-language models with Encyclopedia-scale knowledge.
Pathology vision-language models (VLMs) are promising for building the clinical decision support systems. However, a key barrier to real-world clinical deployment lies in the lack of rigorous and clinically meaningful model evaluation. Existing pathology visual question answering (VQA) benchmarks exhibit the weak alignment of vision-language information, insufficient support for knowledge-intensive reasoning, and a narrow disease spectrum, thereby undermining the ability to provide clinical-level evaluation of model performance. In this work, we introduce OmniPathoVQA, a new pathology VQA benchmark offering a broad disease coverage, spanning all human anatomical systems, and thousands of disease entities. To enhance well-organized visual-language alignment, we leverage pathology educational materials and design a fine-grained extraction pipeline for linking pathology images with the correct knowledge. To examine the knowledge depth for pathological reasoning, the hard-version questions are designed based on microanatomic features underlying similar diseases. A comprehensive evaluation of eighteen VLMs reveals that closed-source VLMs attain higher scores compared to open-source counterparts, yet their performance drops dramatically in difficult questions. We find that strong general reasoning ability is a crucial catalyst for the successful fine-tuning and pathology feature interpretation. Our study establishes a rigorous standard for assessing high-level reasoning and guiding the rapid development of clinical pathology VLMs.
Authors
Chen Chen, Wei Wei, Rui Rui, Yuan Yuan, Zhang Zhang, Du Du, He He, Wang Wang, Liu Liu, Zhou Zhou, Chen Chen
View on Pubmed