다중 스케일 추론을 통한 병리학적 시각 언어 모델 성능 향상
Enhancing Pathological VLMs with Cross-scale Reasoning
병리학 이미지는 본질적으로 다중 스케일을 가지며, 정확한 진단을 위해서는 병리학자들은 낮은 배율에서 전체 조직 구조에 대한 정보와 높은 배율에서의 세포 형태학 정보를 통합해야 합니다. 기존의 시각-언어 모델(VLM)을 위한 병리학 데이터셋은 다양한 스케일을 포함하는 경우가 많지만, 명시적인 다중 스케일 추론 목표를 가지고 있는 경우는 드뭅니다. 이러한 제한점은 VLM이 필수적인 다중 스케일 표현을 학습하고 증거 기반의 추론 능력을 갖추는 것을 방해합니다. 이 문제를 해결하기 위해, 우리는 병리학 해석을 다중 배율 추론으로 정의하는 최초의 다중 스케일 훈련 및 평가 패러다임을 제안합니다. 그러나 이러한 작업을 구축하는 과정에서 중요한 과제가 드러났습니다. 즉, 멀티 이미지 기반 시각 질의 응답(VQA)은 텍스트만으로 해결 가능한 문제에 취약하며, 모델이 시각적 증거가 아닌 배율에 의존적인 요소들을 사용하여 답을 추측할 수 있다는 것입니다. 이를 해결하기 위해, 우리는 적대적인 텍스트 기반 필터링과 제약 조건 기반 질문 설계 방법을 결합한 누출 방지 데이터 정제 파이프라인을 제안합니다. 이 파이프라인을 통해, 우리는 다양한 배율 레벨의 2,537개의 병리학 이미지에 기반한 4,685개의 객관식 질문으로 구성된 고품질 벤치마크인 Scale-VQA를 구축했습니다. 마지막으로, 우리는 다중 스케일 VQA 작업에서 성능을 최적화하기 위해 강화 학습으로 훈련된 모델인 ScaleReasoner-R1을 제시합니다. ScaleReasoner-R1은 우리의 다중 스케일 추론 벤치마크에서 최고 수준의 성능을 달성했으며, 기존의 단일 스케일 벤치마크에서도 뛰어난 성능을 보입니다. 이러한 결과는 제한적인 다중 스케일 학습만으로도 병리학적 이해를 크게 향상시킬 수 있음을 시사합니다. 코드 및 데모는 오픈 소스로 공개될 예정입니다.
Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnification to cellular morphology at higher magnification for accurate diagnosis. While existing pathological datasets for vision-language model (VLM) include various scales, they often lack an explicit cross-scale reasoning objective. This limitation prevents VLMs from capturing essential cross-scale representations and learning evidence-based reasoning. To bridge this gap, we introduce the first cross-scale training and evaluation paradigm that formulates pathology interpretation as multi-magnification reasoning. However, creating such a task reveals a critical challenge: multi-image visual question answering (VQA) is prone to text-only shortcuts, which allow models to guess answers using magnification-dependent artifacts rather than visual evidence. To address this, we propose a leakage-aware curation pipeline that combines adversarial text-only screening with constraint-guided question design. Using this pipeline, we construct Scale-VQA, a high-quality benchmark with 4,685 multiple-choice questions grounded in 2,537 pathology images across multiple magnification levels. Finally, we present ScaleReasoner-R1, a model trained via reinforcement learning to optimize performance on the cross-scale VQA task. ScaleReasoner-R1 achieves state-of-the-art performance on our cross-scale reasoning benchmark and generalizes to SOTA performance on established single-scale benchmarks. Findings suggest that even the limited cross-scale supervision can significantly improve pathological understanding. The code and demos will be open-sourced.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.