비전 기반 모델의 전이 한계 이해
Understanding the Transfer Limits of Vision Foundation Models
기반 모델은 대규모 사전 학습을 통해 광범위한 지식을 습득하고, 다양한 언어 작업에서 일반화 능력을 보여줍니다. 반면, 비전 기반 모델(VFMs)은 상당한 계산 자원이 투입되었음에도 불구하고, 하위 작업에 따라 성능 향상이 균일하지 않은 경우가 많습니다. 우리는 이러한 제한이 사전 학습 목표와 하위 비전 및 이미지 처리 작업의 요구 사항 간의 불일치에서 비롯된다고 가정합니다. 마스크 이미지 복원 또는 대비 학습과 같은 사전 학습 전략은 일반적인 시각 패턴이나 전역 의미 구조 복구와 같은 작업을 위한 표현을 형성하며, 이는 세분화, 분류 또는 이미지 합성 등 하위 응용 프로그램의 작업별 요구 사항과 일치하지 않을 수 있습니다. 이러한 가설을 실제 임상 영역에서 구체적으로 검증하기 위해, 우리는 두 가지 VFM 모델, 즉 복원 중심의 MAE 기반 모델(ProFound)과 대비 학습 기반 모델(ProViCNet)을 사용하여 전립선 다중 매개변수 MRI 작업 5가지에 대해 평가했습니다. 이를 통해 사전 학습과 하위 작업 간의 일치성이 전이 성능, 즉 사전 학습에서 미세 조정으로의 성능에 미치는 영향을 조사했습니다. 우리의 결과는 사전 학습과 하위 작업 간의 더 나은 일치성이, 미세 조정 전후의 동일한 특징 간의 최대 평균 불일치(MMD)와 같은 간단한 발산 지표로 측정되며, 더 큰 성능 향상과 더 빠른 수렴 속도와 상관관계가 있음을 보여줍니다. 이는 하위 응용 프로그램에 대한 적용 가능성을 염두에 두고 사전 학습 목표를 설계하고 분석하는 것의 중요성을 강조합니다.
Foundation models leverage large-scale pretraining to capture extensive knowledge, demonstrating generalization in a wide range of language tasks. By comparison, vision foundation models (VFMs) often exhibit uneven improvements across downstream tasks, despite substantial computational investment. We postulate that this limitation arises from a mismatch between pretraining objectives and the demands of downstream vision-and-imaging tasks. Pretraining strategies like masked image reconstruction or contrastive learning shape representations for tasks such as recovery of generic visual patterns or global semantic structures, which may not align with the task-specific requirements of downstream applications including segmentation, classification, or image synthesis. To investigate this in a concrete real-world clinical area, we assess two VFMs, a reconstruction-focused MAE-based model (ProFound) and a contrastive-learning-based model (ProViCNet), on five prostate multiparametric MR imaging tasks, examining how such task alignment influences transfer performance, i.e., from pretraining to fine-tuning. Our findings indicate that better alignment between pretraining and downstream tasks, measured by simple divergence metrics such as maximum-mean-discrepancy (MMD) between the same features before and after fine-tuning, correlates with greater performance improvements and faster convergence, emphasizing the importance of designing and analyzing pretraining objectives with downstream applicability in mind.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.