CT 기반 모델의 해부학적 맥락을 고려한 적응 방법
Anatomy Contextualized Adaption of CT Foundation Models
CT 영상과 언어 결합 모델은 다양한 후속 작업에서 뛰어난 성능을 보였지만, 일반적으로 전체 볼륨 표현으로 훈련되어 미세한 해부학적 정보를 희석시키는 경향이 있습니다. 미세 수준의 영상-언어 사전 학습은 해부학 단위의 시각적 특징과 해당 해부학에 특화된 텍스트를 연결하여 이러한 문제를 해결하지만, 이 과정에서 전체 볼륨 모델이 제공하는 전반적인 맥락을 손실합니다. 또한 기존의 미세 수준 접근 방식은 처음부터 훈련하므로 계산 비용이 많이 듭니다. 본 연구에서는 Anatomy Contextualized Adaptation (ACA)이라는 가벼운 프레임워크를 제안하며, 이는 동결된 CT 기반 모델의 표현 방식을 해부학 단위의 영상-언어 정렬에 적응시키는 동시에 전반적인 맥락을 향상시킵니다. ACA는 TotalSegmentator를 사용하여 CT 볼륨을 해부학 단위의 임베딩으로 분해하고, 트랜스포머를 통해 해부학 간의 관계를 파악하여 이러한 임베딩을 개선합니다. 또한, 방사선 보고서에서 추출된 각 해부학 및 전체 스캔 수준의 텍스트와 연결합니다. Merlin과 CT-RATE 데이터셋에 대한 평가 결과, ACA는 동결된 기본 모델과 기존의 미세 수준 방법보다 일관되게 우수한 성능을 보이며, 임베딩이 캐시되면 1시간 미만의 짧은 훈련 시간으로 달성됩니다. ACA의 해부학 간 트랜스포머가 학습한 어텐션 가중치는 해부학 간의 가능한 맥락 연결을 나타냅니다. 종합적으로 볼 때, 이러한 결과는 ACA가 CT 기반 모델을 해부학적 지식을 기반으로 한 영상-언어 정렬에 적응시키는 데 효과적인 방법이며, 동시에 전반적인 해부학적 맥락을 유지하고 향상시킬 수 있음을 시사합니다.
CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual features with anatomy-specific text, but in doing so discards the global context that whole-volume models provide. Furthermore, existing fine-grained approaches train from scratch, making them computationally expensive. We introduce Anatomy Contextualized Adaptation (ACA), a lightweight framework that adapts frozen CT foundation model representations for anatomy-level vision-language alignment while enhancing global contextualization. ACA uses TotalSegmentator to decompose CT volumes into anatomy-level embeddings, which are refined via a transformer that captures cross-anatomy relationships, and aligned to both per-anatomy and scan-level text extracted from radiology reports. Evaluated on Merlin and CT-RATE, ACA consistently outperforms both the frozen foundation model baselines and existing fine-grained methods in zero-shot finding classification, while requiring less than one hour of training once embeddings are cached. The attention weights learned by ACA's inter-anatomy transformer additionally indicate plausible cross-anatomy context routing. Altogether, these results support ACA as a lightweight approach for adapting CT foundation models to anatomically grounded vision-language alignment while preserving and enhancing global anatomical context.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.