다중 기관 전립선 병변 분할을 위한 계층적 잠재 레이블 모델링 기반 딥 EM
Deep EM with Hierarchical Latent Label Modelling for Multi-Site Prostate Lesion Segmentation
전립선 병변 분할에서 레이블의 다양성은 주요 과제입니다. 다중 기관 데이터 세트에서, 어노테이션은 종종 기관별 특정 윤곽선 프로토콜을 반영하여, 분할 네트워크가 특정 기관의 스타일에 과적합되어 추론 시 새로운 기관에 대한 일반화 성능이 저하되는 문제를 야기합니다. 본 연구에서는 관찰된 각 어노테이션을 근본적인 '깨끗한' 병변 마스크의 노이즈가 있는 관찰값으로 간주하고, 계층적 기댓값 최대화 (HierEM) 프레임워크를 제안합니다. 이 프레임워크는 다음과 같은 과정을 반복합니다. (1) 각 픽셀에 대한 잠재 마스크의 사후 분포를 추론하고, (2) 이 사후 분포를 소프트 타겟으로 사용하여 CNN을 훈련하고, 계층적 사전 분포를 통해 기관별 민감도와 특이성을 추정합니다. 이 계층적 사전 분포는 레이블 품질을 전역 평균과 기관 및 사례 수준의 편차로 분해하여, 기관별 편향을 줄이기 위해 기관 편차에 의해만 기여되는 likelihood 항을 페널티로 적용합니다. 세 개의 코호트 데이터 세트에 대한 실험 결과, 제안된 계층적 EM 프레임워크가 최첨단 방법보다 교차 기관 일반화 성능을 향상시키는 것을 확인했습니다. 풀링된 데이터 세트에 대한 평가는 기관별 평균 DSC 값이 29.50%에서 39.69% 사이이며, 기관 하나를 제외한 일반화 성능은 27.91%에서 32.67% 사이로 나타났습니다. 이는 비교 방법보다 통계적으로 유의미한 개선을 보여줍니다 (p<0.039). 또한, 제안된 방법은 기관별 잠재 레이블 품질을 해석 가능한 형태로 제공합니다 (민감도 alpha는 31.5%에서 47.3% 사이이며, 특이도 beta는 약 0.99). 이는 교차 기관 어노테이션의 다양성에 대한 사후 분석을 지원합니다. 이러한 결과는 기관 의존적인 어노테이션을 명시적으로 모델링함으로써 교차 기관 일반화를 향상시킬 수 있음을 시사합니다.
Label variability is a major challenge for prostate lesion segmentation. In multi-site datasets, annotations often reflect centre-specific contouring protocols, causing segmentation networks to overfit to local styles and generalise poorly to unseen sites in inference. We treat each observed annotation as a noisy observation of an underlying latent 'clean' lesion mask, and propose a hierarchical expectation-maximisation (HierEM) framework that alternates between: (1) inferring a voxel-wise posterior distribution over the latent mask, and (2) training a CNN using this posterior as a soft target and estimate site-specific sensitivity and specificity under a hierarchical prior. This hierarchical prior decomposes label-quality into a global mean with site- and case-level deviations, reducing site-specific bias by penalising the likelihood term contributed only by site deviations. Experiments on three cohorts demonstrate that the proposed hierarchical EM framework enhances cross-site generalisation compared to state-of-the-art methods. For pooled-dataset evaluation, the per-site mean DSC ranges from 29.50% to 39.69%; for leave-one-site-out generalisation, it ranges from 27.91% to 32.67%, yielding statistically significant improvements over comparison methods (p<0.039). The method also produces interpretable per-site latent label-quality estimates (sensitivity alpha ranges from 31.5% to 47.3% at specificity beta approximates 0.99), supporting post-hoc analyses of cross-site annotation variability. These results indicate that explicitly modelling site-dependent annotation can improve cross-site generalisation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.