단순 계산을 넘어: 병리학 기반 모델의 분포적 강건성 마진
Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models
병리학 기반 모델은 임상 적용에 점점 가까워지고 있지만, 여전히 센터 간의 체계적인 비생물학적 변동에 취약합니다. 조직 처리, 염색 및 스캔 과정에서의 차이는 모델의 표현 방식에 강하게 반영되어, 단기 학습을 유발하고 코호트 및 기관 간 일반화 능력을 약화시킵니다. Robustness Index (RI)는 지역 표현 방식의 기하학적 구조가 생물학적 요인에 의해 지배되는지 아니면 비생물학적 변동에 의해 지배되는지를 정량적으로 평가하지만, RI의 계산 기반 방식은 거리 정보를 고려하지 않습니다. 저희 연구에서는 거리 가중치를 추가해도 큰 변화가 없으며, 그 이유는 RI 자체가 고정된 이웃 구조를 사용하여 샘플 수준의 이질성을 숨기고 모델 의존적인 샘플 집합만을 평가하기 때문입니다. 따라서 저희는 Cross-confounder Robustness Margin (CRoMa)라는 샘플 단위로 분석되는 새로운 측정 방법을 제안합니다. CRoMa는 거리 정보를 활용하여 교차 요인(cross-confounder) 생물학적 매칭 항목과 동일 요인(same-confounder) 생물학적 주의 대상 항목 간의 거리를 직접 비교합니다. CRoMa는 강건성을 단일 통합 점수가 아닌 코호트 전체의 마진 분포로 재정의합니다. 저희는 세 개의 벤치마크에서 20개의 타일 수준 인코더와 네 번째 데이터셋에서 4개의 슬라이드 수준 인코더의 고정된 표현 방식을 평가했습니다. 중앙값 CRoMa를 기준으로 한 순위는 데이터셋 전반에 걸쳐 대체로 일관성을 보였지만, 기본 분포는 모델 내부의 상당한 이질성을 드러냈습니다. 모든 타일 인코더에서 요인 지배적인 하위 꼬리가 남아 있었으며, 그 빈도와 심각성은 모델마다 크게 달랐습니다. 이러한 다양한 강건성 프로필은 모델 선택을 일반적 강건성과 낮은 꼬리 강건성 간의 파레토 최적화 문제로 제시합니다. 또한 높은 CRoMa 값은 지도 학습 적응 후 발생하는 단축 경로(shortcut)에 의한 성능 저하를 줄이는 것과 관련이 있었습니다. CRoMa는 표현 방식의 기하학적 구조를 분포적 강건성 지표로 변환하여, 다운스트림에서 발생할 수 있는 단축 경로 취약성을 예측함으로써, 강건성 평가 및 모델 선택을 위한 체계적인 기반을 제공합니다.
Pathology foundation models are approaching clinical deployment, yet remain vulnerable to systematic non-biological variation across centres. Differences in tissue preparation, staining and scanning are strongly encoded in their representations, enabling shortcut learning and weakening generalisation across cohorts and institutions. The Robustness Index (RI) quantifies whether local representation geometry is dominated by biology or by non-biological variation, but its count-based formulation discards distance information. We show that adding distance weights changes little because the deeper limitation lies in RI's pooled, fixed-neighbourhood design, which obscures sample-level heterogeneity and effectively evaluates only a model-dependent subset of samples. We introduce the Cross-confounder Robustness Margin (CRoMa), a sample-resolved measure that directly compares distances to cross-confounder biological matches and same-confounder biological distractors. CRoMa recasts robustness as a cohort-wide margin distribution rather than a single pooled score. We evaluated frozen representations from 20 tile-level encoders across three benchmarks and 4 slide-level encoders on a fourth. Rankings by median CRoMa were broadly consistent across datasets, while the underlying distributions revealed substantial within-model heterogeneity. Every tile encoder retained a confounder-dominated lower tail, whose prevalence and severity varied markedly across models. These distinct robustness profiles frame model selection as a Pareto trade-off between typical and lower-tail robustness. Higher CRoMa was also associated with smaller shortcut-induced performance drops after supervised adaptation. By turning representation geometry into a distributional robustness readout that anticipates downstream shortcut susceptibility, CRoMa provides a principled basis for robustness assessment and model selection.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.