2607.19261v1 Jul 21, 2026 cs.CV

PathAgentBench: 전체 슬라이드 병리 이미지에 대한 증거 탐색 비전-언어 모델 성능 평가

PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

Tianyi Zhang
Tianyi Zhang
Citations: 38
h-index: 2
Zeyu Liu
Zeyu Liu
Citations: 17
h-index: 3
Qiaochu Xue
Qiaochu Xue
Citations: 5
h-index: 1
Yufeng Wu
Yufeng Wu
Citations: 0
h-index: 0
Yueming Jin
Yueming Jin
Citations: 26
h-index: 2
Dankai Liao
Dankai Liao
Citations: 9
h-index: 2
Xinyue Zhang
Xinyue Zhang
Citations: 0
h-index: 0
Dachun Zhao
Dachun Zhao
Citations: 1
h-index: 1
Linghan Cai
Linghan Cai
Citations: 53
h-index: 1

전체 슬라이드 이미지(WSI) 진단은 진단적으로 관련된 영역을 식별하고, 다양한 배율에서 이를 검토하며, 다중 척도 정보를 통합하는 것을 요구합니다. 그러나 대부분의 기존 병리 벤치마크는 모델을 사전 크롭된 패치 또는 사전 추출된 슬라이드 특징에 대해 평가하여, 기가픽셀 WSI로부터 직접적으로 증거를 획득하는 모델의 능력은 충분히 검증되지 않았습니다. 본 논문에서는 증거 탐색 비전-언어 모델(VLM)을 평가하기 위한 벤치마크인 PathAgentBench를 소개합니다. 이 벤치마크는 다음 네 가지 상호 보완적인 기능을 평가합니다: 증거 해석을 위한 이미지-텍스트 매칭, 증거 검증을 위한 텍스트-이미지 검색, 증거 획득을 위한 진단 영역 위치 지정, 그리고 증거 통합을 위한 다중 척도 추론. 이 벤치마크는 다양한 배율에서 중첩된 영역을 연결하고, 각 척도별 결과 및 슬라이드 수준 진단을 제공하는 진단 트리 구조로 구성되어 있습니다. 본 벤치마크에는 10명의 전문의가 주석을 단 1,822개의 TCGA WSI와 17,135개의 진단 경로가 포함되어 있으며, 추가적으로 상세한 주석이 달린 190개의 유방암 WSI 세트를 사용하여 완전 자동 슬라이드 탐색 성능을 평가했습니다. 본 논문에서는 20개의 범용 모델, 의료 특화 모델 및 병리학 전문 모델을 평가합니다. 선도적인 오픈 소스 모델은 다중 척도 추론에서 93% 이상의 정확도를 달성했으며, 양방향 매칭 작업에서도 50% 이상의 정확도를 보였습니다. 반면, 진단 영역 위치 지정은 여전히 어려운 과제로 남아 있으며, 텍스트 기반의 최적 평균 IoU(Intersection over Union) 점수는 0.09 미만으로, 단순한 중심 기반 휴리스틱 방법보다 성능이 낮습니다. 완전 자동 탐색 과정에서, 낮은 배율에서는 0.522의 높은 hit rate를 보였지만, 중간 배율에서는 0.185로 감소하고, 높은 배율에서는 0.020으로 더욱 감소합니다. 이러한 결과는 선별된 증거에 대한 추론과 WSI로부터 직접적으로 증거를 획득하는 것 사이의 상당한 격차가 있음을 보여줍니다. PathAgentBench는 증거 탐색 병리학 모델을 측정하고 개선하기 위한 통합 프레임워크를 제공합니다.

Original Abstract

Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence. However, most existing pathology benchmarks evaluate models on pre-cropped patches or pre-extracted slide features, leaving their ability to acquire evidence directly from gigapixel WSIs largely untested. We introduce PathAgentBench, a benchmark for evaluating evidence-seeking vision-language models (VLMs) across four complementary capabilities: image-to-text matching for evidence interpretation, text-to-image retrieval for evidence verification, diagnostic-region localization for evidence acquisition, and multi-scale reasoning for evidence integration. The benchmark is organized as a diagnostic tree that links nested regions across magnifications with scale-specific findings and path-level diagnoses. It contains 1,822 TCGA WSIs and 17,135 diagnostic paths annotated by ten board-certified pathologists. An additional private cohort of 190 breast cancer WSIs with detailed annotations is used to evaluate autonomous whole-slide exploration. We evaluate 20 general-purpose, medical, and pathology-specialized models. Leading open-weight models achieve over 93% accuracy in multi-scale reasoning and over 50% accuracy in both cross-modal matching tasks. In contrast, diagnostic-region localization remains challenging: the best text-guided mean intersection-over-union is below 0.09, underperforming a simple center-based heuristic. During autonomous exploration, the unconditional hit rate decreases from 0.522 at low magnification to 0.185 at intermediate magnification and 0.020 at high magnification. These results reveal a pronounced gap between reasoning over curated evidence and acquiring that evidence directly from WSIs. PathAgentBench provides a unified framework for measuring and improving evidence-seeking pathology models.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!