동결된 3D CT 영상 인코더 내의 희소 개념 채널
Sparse Concept Channels in Frozen 3D CT Vision Encoders
대규모 시각-언어 모델이 3차원 의료 영상 해석 분야에서 점점 더 중요한 역할을 하고 있지만, 내부 유닛 중 어떤 부분이 임상적 발견을 인코딩하는지, 그리고 그 정보가 표현 어디에 존재하는지는 잘 알려져 있지 않습니다. 본 연구에서는 먼저 3차원 흉부 시각-언어 모델(Pillar-0)의 동결된 시각 임베딩을 분석하여 이 문제를 조사했습니다. (i) 각 방사선학적 발견은 전체 특징 분류 성능과 일치하고, 제로샷 텍스트 프롬프팅보다 훨씬 뛰어난 약 10개의 희소한 시각 인코더 채널에 의해 인코딩됨을 보여주었습니다. (ii) 특정 발견과 관련된 채널을 비활성화하면 해당 발견의 점수는 급격히 감소하는 반면, 관련 없는 레이블은 안정적으로 유지됩니다. (iii) 동일한 희소한 탐색 방법이 구조적으로 다른 3차원 복부 VLM(Merlin)에서도 재현되는 것을 확인했는데, 이는 동결된 의료 인코더의 일반적인 특성을 시사합니다. 저희는 훈련 없이 개념 채널을 탐색하는 CCP (Concept Channel Probe) 방법을 개발했으며, 이 방법은 코퍼스로부터 생성된 보고서 템플릿과 함께 기존 CT-CHAT 모델보다 임상 효능 및 자연어 생성 지표(F1: 0.549 vs. 0.184; BLEU: 0.483 vs. 0.373)에서 훨씬 우수한 성능을 보였으며, 동시에 지연 시간은 22배 더 짧습니다. 본 연구 결과는 동결된 의료 인코더가 발견 사항을 어떻게 표현하는지에 대한 명확하고 재현 가능한 특징을 제공하며, 이는 다양한 모델에 직접적으로 적용될 수 있습니다.
Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely know <i>which</i> internal units encode clinical findings or <i>where</i> that information lives in the representation. We first study this on a 3D chest vision-language model (Pillar-0) by probing its frozen vision embeddings. We show that (i) each radiological finding is encoded by a <i>sparse</i> set of ~10 vision-encoder channels that match full-feature classification performance and far exceed a zero-shot text prompting; (ii) turning off the channels tied to one finding, that finding's score collapses while unrelated labels stay stable; and (iii) the same sparse probe <i>replicates</i> on an architecturally unrelated 3D abdominal VLM (Merlin) suggesting a general property of frozen medical encoders. Our training-free concept channel probe (CCP) method, paired with a corpus-derived report template, outperforms published CT-CHAT on clinical efficacy and NLG metrics (F1 0.549 vs. 0.184; BLEU 0.483 vs. 0.373) at 22x lower latency. Our results provide a clear, reproducible characterization of how frozen medical encoders represent findings, demonstrating direct applicability across models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.