공간 의미론에서 시간적 맥락으로: 시선 추적 데이터를 활용한 약하게 지도된 의료 영상 분할
From Spatial Semantics to Temporal Context: Leveraging Gaze Trajectory for Weakly Supervised Medical Image Segmentation
의료 영상 분할은 노동 집약적이고 시간이 많이 소요되는 픽셀 단위 어노테이션에 크게 의존합니다. 시선 추적 기술은 비용 효율적인 솔루션을 제공하며, 임상 워크플로우에 자연스럽게 통합될 수 있습니다. 시선 추적 장치로 기록된 시선 데이터는 고정점을 통해 의료 전문가의 주의 영역을 나타내며, 궤적을 통해 시간적인 맥락과 점진적인 시각 인지 과정을 전달합니다. 그러나 시간적인 궤적 모델링은 여전히 어려운 과제이며, 탐색적인 고정점으로 인한 시선 데이터 노이즈는 분할 성능을 크게 제한합니다. 이러한 한계를 극복하기 위해, 우리는 시선 데이터를 활용하여 공간 의미론 모델링부터 시간적 맥락까지 학습하는 Trajectory-guided Uncertainty-aware Network (TrailNet)를 제안합니다. 구체적으로, 제안된 궤적 기반의 시공간 인코더는 시간적 맥락을 모델링하고, 이미지의 공간 의미론과 상호 작용하여 대상 인식 능력을 강화합니다. 또한, 다중 스케일 불확실성 디코더는 범주 간 상호 배타적인 제약을 활용하여 결정적인 예측을 수행하고, 노이즈로 인해 발생하는 지도 학습의 불확실성을 완화합니다. 시선 데이터 없이 추론할 수 있도록, 우리는 교사-학생 네트워크를 통해 특징 수준의 지식을 전달하는 사이클 증류 전략을 추가적으로 도입했습니다. 두 개의 공개 데이터셋에 대한 실험 결과는 TrailNet이 최첨단 방법보다 우수한 성능을 보이며, 각각 81.25% 및 81.85%의 Dice 점수를 달성했음을 보여줍니다.
Medical image segmentation heavily depends on labor-intensive and time-consuming pixel-level annotations. Eye tracking offers a cost-effective solution that can be naturally integrated into clinical workflows. Recorded by eye trackers, gaze conveys the spatial regions of clinicians' attention through fixations and the temporal context of clinicians' progressive visual perception from trajectories. Nevertheless, effective modeling of temporal trajectories remains challenging, and noise in gaze caused by exploratory fixations greatly limits segmentation performance. To overcome these limitations, we propose the Trajectory-guided Uncertainty-aware Network (TrailNet), which exploits gaze-supervised medical image segmentation from spatial semantics modeling to temporal context by jointly leveraging fixations and trajectories. Specifically, the proposed trajectory-guided spatio-temporal encoder models temporal context and establishes complementary interactions with image spatial semantics to strengthen target perception. Furthermore, the multi-scale uncertainty decoder leverages category mutual-exclusivity constraints to produce deterministic predictions and mitigate supervision uncertainty induced by noise. To enable gaze-free inference, we further introduce a cycle distillation strategy that transfers feature-level knowledge via teacher-student networks. Experimental results on two public datasets demonstrate that TrailNet outperforms state-of-the-art methods, achieving Dice scores of 81.25% and 81.85%, respectively.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.