2608.03508v1 Aug 04, 2026 cs.CV

다중 해상도 세포부터 기가픽셀 전체 슬라이드 이미지까지: 계산 병리학을 위한 기반 모델

From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

Sajid Javed
Sajid Javed
Citations: 403
h-index: 10
B. Alawode
B. Alawode
Citations: 222
h-index: 9
Moshira Abdalla
Moshira Abdalla
Citations: 54
h-index: 2
Dwarikanath Mahapatra
Dwarikanath Mahapatra
Citations: 92
h-index: 5
Muhammad Muzammal Naseer
Muhammad Muzammal Naseer
Citations: 0
h-index: 0

비전 트랜스포머(ViT) 및 그 계층적 변형은 계산 병리학(CPath) 분야에서 뛰어난 성능을 보여왔습니다. 그러나 대부분의 모델은 단일 해상도의 전체 슬라이드 이미지(WSI)로 사전 학습되어, 임의의 해상도에 대한 일반화 능력이 제한됩니다. 기가픽셀 WSI는 세포 형태, 조직 구조 및 전반적인 맥락과 같이 다양한 크기에서 진단 패턴을 내포하고 있으며, 이는 숙련된 병리학자들이 WSI를 검사하는 방식과 일치합니다. 본 연구에서는 세포부터 조직, 그리고 전체 슬라이드 수준까지 다중 해상도 정보를 계층적으로 통합하는 모델인 Multi-Resolution Pyramid Transformer (MRPT)를 소개합니다. MRPT는 생물학적으로 의미 있는 연속적 교차 해상도 어텐션(CCRA) 메커니즘을 사용하여 크기에 독립적인 상호 작용을 포착하고, 해상도 간의 임베딩을 정렬하여 다중 해상도의 의미론적 일관성을 강화함으로써 강력하고 일반화 가능한 WSI 표현을 제공합니다. MRPT는 6억 2천4백만 개의 패치, 240만 개의 영역 및 3만 6천 개의 WSI를 사용하여 다중 해상도 방식으로 사전 학습되었으며, 이를 통해 풍부한 거친 것부터 세밀한 것에 이르기까지의 조직 병리학적 특징을 학습합니다. 34개의 다양한 데이터셋에 대한 광범위한 실험 결과, MRPT는 최근의 기반 모델 및 다중 모드 대규모 언어 모델(MLLM)보다 암 아형 분류, 조직 현상 분석 및 WSI 이해를 위한 시각 질의 응답(VQA)에서 우수한 성능을 보였습니다.

Original Abstract

Vision Transformers (ViTs) and their hierarchical variants have achieved strong performance in Computational Pathology (CPath). However, most are pre-trained on single-resolution Whole Slide Images (WSIs), limiting their generalization across arbitrary resolutions. Gigapixel WSIs inherently contain diagnostic patterns at multiple scales, including cellular morphologies, tissue architectures, and global context, mirroring how expert pathologists examine WSIs. We introduce Multi-Resolution Pyramid Transformer (MRPT), a model that hierarchically aggregates multi-resolution information from cellular to tissue and WSI levels. MRPT employs a biologically meaningful Consecutive Cross-Resolution Attention (CCRA) mechanism to capture scale-independent interactions and enforces multi-resolution semantic consistency by aligning embeddings across resolutions, yielding robust and generalizable WSI representations. Pre-trained in a multi-resolution self-supervised manner on 624M patches, 2.4M regions, and 36K WSIs, MRPT learns rich coarse-to-fine histopathology features. Extensive experiments on 34 diverse datasets show that MRPT surpasses recent foundation models and Multimodal Large Language Models (MLLMs) in cancer subtype classification, tissue phenotyping, and Visual Question Answering (VQA) for WSI understanding.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!