OrganLens: CT 기반 모델을 위한 장기별 표현 학습
OrganLens: Organ-Specific Representation Learning for CT Foundation Models
CT 검사는 여러 장기를 포함하지만, 많은 생의학적 질문은 특정 장기의 이상, 예후 또는 경과적인 변화에 관한 것입니다. 이러한 질문에는 동일한 CT 데이터 내에서 각 장기에 대한 별도의 표현이 필요합니다. 기존의 CT 기반 모델은 일반적으로 단일 볼륨 수준의 표현을 생성하는 반면, 최근의 해부학적 구조 인지 방법은 사전에 분리된 장기 볼륨을 인코딩하거나 이미지를 명시적으로 장기 토큰 그룹으로 분리합니다. 전자는 임상적으로 중요한 주변 맥락을 제거할 수 있으며, 후자는 특징이 형성되기 전에 선택된 장기에 대한 공유 인코더를 활용하지 않습니다. 우리는 자기 지도 학습을 통해 장기별 표현 학습을 위한 OrganLens를 소개합니다. 장기 식별 정보는 공유 CT 인코더에 조건을 부여하며, 장기별 증류 및 해부학적 마스크 기반 감독은 해부학적 가중치를 적용한 풀링을 통해 장기별 표현을 형성하도록 특징을 조정합니다. 추론 단계에서, OrganLens는 외부 분할 마스크 없이 11개의 장기별 표현을 생성합니다. 우리는 다양한 촬영 환경 및 후속 평가를 위해 CT-RATE, RAD-ChestCT, INSPECT, 그리고 NLST 데이터셋에서 OrganLens를 평가했습니다. CT 사전 학습 모델인 DINOv2와 비교했을 때, 심장 표현은 CT-RATE의 심비대증 AUROC 값을 0.910에서 0.953으로 향상시켰으며, 폐 표현은 NLST의 폐암 사망률에 대한 Harrell C-index를 14.2% 개선했습니다. 전역 표현은 INSPECT 데이터셋에서 텍스트-이미지 및 이미지-텍스트 검색 모두 Recall@10을 각각 33.09% 및 32.04%로 달성했습니다. 장기 관련 작업에서, 해부학적으로 일치하는 표현은 보다 강력한 작업 관련 신호를 제공하며, 전역 표현은 광범위한 유용성을 유지합니다. OrganLens는 공유 인코더를 사용한 확장 가능한 장기별 CT 표현 학습 접근 방식을 제공합니다. 더 넓은 의미에서, 이는 의료 연구 커뮤니티에 다양한 코호트 및 임상 지표에서의 장기별 질병 연구를 위한 재사용 가능한 프레임워크를 제공합니다.
A CT examination captures multiple organs, but many biomedical questions concern abnormalities, prognosis, or longitudinal change in a specific organ. These questions require a separate representation for each organ within the same CT volume. Existing CT foundation models commonly produce a single volume-level representation, while recent anatomy-aware methods either encode pre-separated organ volumes or explicitly disentangle images into organ token groups. The former may remove clinically relevant surrounding context, while the latter does not condition a shared encoder on a selected organ before its features are formed. We introduce OrganLens for organ-specific representation learning through self-supervision. An organ identity conditions a shared CT encoder, while organ-specific distillation and anatomy-mask supervision shape features for anatomy-weighted pooling into organ-specific representations. At inference, the shared model produces 11 organ-specific representations without external segmentation masks. We evaluate OrganLens on CT-RATE, RAD-ChestCT, INSPECT, and NLST across diverse acquisitions and downstream evaluations. Relative to CT-pretrained DINOv2, heart representations raise CT-RATE cardiomegaly AUROC from 0.910 to 0.953, while lung representations improve the Harrell C-index for NLST lung-cancer mortality by 14.2\%. The global representation reaches INSPECT Recall@10 of 33.09\% and 32.04\% for text-to-image and image-to-text retrieval, respectively. Across organ-related tasks, anatomically matched representations provide stronger task-relevant signal, while the global representation retains broad utility. OrganLens offers a scalable approach to organ-specific CT representation learning with a shared encoder. More broadly, it provides the medical research community with a reusable framework for studying organ-specific disease across cohorts and clinical endpoints.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.