자기 지도 학습이 임상 감독보다 의료 기반 모델의 표현 수렴에 더 큰 영향을 미친다
Self-supervision drives representational convergence in medical foundation models more than clinical supervision
서로 다른 연구 그룹에서 개발된 의료 영상 인코더는 점점 더 상호 교환 가능하다고 여겨지는데, 이는 규모와 임상 감독이 이러한 인코더들의 표현을 공통 구조로 집중시킨다는 가정 때문입니다. 그러나 이러한 수렴 현상이 실제로 존재하는지, 무엇이 이를 유발하는지, 그리고 이것이 임상적으로 활용 가능한지에 대한 검증은 이루어지지 않았으며, 이러한 주장을 뒷받침하는 유사성 측정 방법은 취약합니다. 본 연구에서는 18개의 이미지 인코더와 7개의 텍스트 인코더를 대상으로 통제된 실험을 수행했습니다. 모든 인코더는 공개 가중치를 사용하며 로컬 환경에서 실행되며, 7백만 개에서 270억 개의 파라미터를 가지며, 650,982개의 흉부 X선 이미지를 포함한 5가지 영상 모달리티를 포괄합니다. 원인과 결과를 명확히 하기 위해, 데이터, 아키텍처 및 규모를 고정한 상태에서 학습 목표만 변경하여 인코더를 훈련하고, 합성 모델을 사용하여 동일한 효과를 재현했습니다. 수렴 현상은 미미하지만 무작위 수준 이상이며, 이는 임상 감독이 아닌 자기 지도 학습 목표에 의해 유발됩니다. 자기 지도 학습으로 훈련된 인코더 간의 정렬도가 가장 높았으며 (흉부 X선 이미지에서 40.4%), 레이블 기반 감독 (21.1%) 및 이미지-텍스트 방식 (3.3%)은 훨씬 낮았습니다. 또한, 수렴 정도는 모델 규모 (Spearman 상관 계수 0.302, p=0.223) 또는 성능과 함께 증가하지 않았습니다. 이러한 수렴 현상은 동일 모달리티 내에서만 나타나며, 임상 용어와 일치하지 않으며, 방사선 전문의가 판단하는 유사성을 정확하게 반영하지 못합니다. 그러나 선형 분류기는 인코더 간에 그리고 5개의 독립적인 병원 데이터셋으로 전이될 수 있으며, 원래 인코더에서의 성능의 약 85%를 유지합니다. 따라서 의료 영상에서 나타나는 표현 수렴은 모델 규모나 임상 감독으로부터 유래하는 것이 아니라 사전 학습 목표에 의해 결정됩니다. 이에 따라, 상호 운용성은 해당 목표를 통해 설계되어야 하며, 공유된 기하학적 구조가 가장 취약한 부분 (예: 환자 하위 그룹 간의 비교 및 임상 판단과의 비교)에서 검증되어야 합니다.
Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervision concentrate their representations onto a shared structure. Whether this convergence is real, what produces it, and whether it is clinically usable are untested, and the similarity measures behind such claims are fragile. We present a controlled dissection across 18 image and 7 text encoders, all open-weight and run locally, spanning 7M to 27B parameters and five imaging modalities, including 650,982 chest radiographs from six datasets. To isolate cause, we train encoders that vary only the objective under fixed data, architecture, and scale, and reproduce the effect in a synthetic model. Convergence is modest but above a random floor, driven by the self-supervised objective, not clinical supervision: matched self-supervised encoders aligned most (40.4% on chest radiography), with label-supervised (21.1%) and image-text (3.3%) far lower, and did not grow with size (Spearman 0.302, p=0.223) or capability. It is within-modality, does not reach clinical language, and does not reproduce how radiologists judge case similarity. Yet a linear classifier transfers across encoders and to five held-out hospitals, retaining about 85% of within-encoder performance. Convergence in medical imaging is therefore set by the pretraining objective, not inherited from scale or clinical supervision. Interoperability is accordingly something to design for through that objective, and to validate where the shared geometry is weakest, across patient subgroups and against clinical judgment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.