2608.05960v1 Aug 06, 2026 cs.CV

크고 밝거나 보이지 않는가: 3차원 CT 기반 모델의 기능 평가 지표

Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models

D. Rueckert
D. Rueckert
Citations: 1,828
h-index: 13
Philip Müller
Philip Müller
Citations: 11
h-index: 3
Maulik Chevli
Maulik Chevli
Citations: 51
h-index: 3
Johannes Brandt
Johannes Brandt
Citations: 167
h-index: 3
R. Braren
R. Braren
Citations: 9,093
h-index: 37

일상적인 CT 영상 판독은 전체 스캔 범위를 포괄하며, 우연히 발견되는 소견들을 포함합니다. 3차원 CT 기반 모델은 해부학적 구조와 병변에 대한 일반화된 표현을 제공함으로써 이러한 과정을 지원할 수 있습니다. 본 연구에서는 10개의 동결된 CT 인코더 모델을 세 그룹의 흉부 CT 스캔 데이터셋(내부 임상 데이터 포함)에 대해 $k$-최근접 이웃 방법, 제로샷 프롬프팅 및 선형 탐색법을 사용하여 진단 능력을 평가했습니다. 그 결과, 특정 평가 환경에 따라 순위가 크게 변동하며, 전반적으로 우수한 성능을 보이는 모델은 없었습니다. 일반적으로 미세한 이미지 토큰화와 시각-언어 정렬을 결합한 모델이 가장 좋은 성능을 보였지만, 경량의 지도 학습 인코더도 매우 경쟁력 있는 성능을 보여주었으며, 이는 명시적인 레이블이 크기(scale)를 효과적으로 대체할 수 있음을 나타냅니다. 중요한 점은, 모델 구조보다는 실제 성능에 영향을 미치는 주요 요인이 물리적 제약 조건이라는 것을 확인했습니다. 즉, 병변의 검출 가능성은 주변 조직과의 대비 정도와 공간적 범위에 따라 달라집니다. 통제된 장기 내 비교 실험을 통해 광범위하거나 높은 대비를 가진 이상 소견(예: 의료 기구 및 액체)은 안정적으로 식별될 수 있음을 확인했습니다. 반면, 작은 크기의 낮은 대비를 가진 병변은 평가된 모든 인코더에서 여전히 어려운 과제로 남아 있습니다. 이는 전역 풀링 기반 임베딩의 고유한 한계 때문이며, 작은 크기 및 낮은 대비 구조를 정확하게 표현하려면 영역 또는 병변 수준에서의 사전 학습이 필요할 것으로 판단됩니다.

Original Abstract

Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume. 3D CT foundation models could assist this process by providing generalizable representations of anatomy and pathology. To evaluate their diagnostic breadth, we benchmark ten frozen CT encoders across three cohorts of thoracic CT scans, including an unseen internal clinical dataset, using $k$-nearest neighbors, zero-shot prompting, and linear probing. We find no universal state-of-the-art, with rankings fluctuating significantly depending on the evaluation context. While models combining fine-grained image tokenization with vision-language alignment generally perform best, a lightweight supervised encoder remains highly competitive, demonstrating that explicit labels can effectively substitute for scale. Crucially, rather than model architecture, we observe that the primary determinant of performance is a physical bottleneck: a finding's detectability scales with its contrast against surrounding tissue and its spatial extent. Through controlled within-organ comparisons, we empirically demonstrate that widespread or high-contrast abnormalities, such as devices and effusions, are reliably recovered. Conversely, small, low-contrast focal lesions remain a persistent challenge across all evaluated encoders. We attribute this to the inherent limitations of globally pooled embeddings, suggesting that accurately representing small, low-contrast structures will require region- or lesion-level pretraining.

0 Citations
0 Influential
18.5 Altmetric
92.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!