EchoPilot: 스케일 공간 의미 프롬프트 및 신뢰도 기반 메모리를 활용한 학습 불필요 초음파 비디오 분할
EchoPilot: Training-Free Ultrasound Video Segmentation via Scale-Space Semantic Prompting and Reliability-Gated Memory
초음파 비디오 분할은 임상적으로 매우 유용하지만, speckle 노이즈, 약한 경계선, 그리고 빠른 해부학적 변형 때문에 어려운 기술입니다. 최근의 프롬프트 기반 기초 모델들은 포인트 가이드 분할을 가능하게 하지만, 초음파 영상에 직접 적용하는 것은 여전히 신뢰성이 낮습니다. 단일 포인트는 스케일 모호성을 해결하기 위한 충분한 공간 정보를 제공하지 못하며, 탐욕적인 메모리 업데이트는 초기 오류를 심각한 시간적 드리프트로 증폭시킵니다. 우리는 EchoPilot이라는 학습이 필요 없는 초음파 비디오 분할 프레임워크를 제안합니다. EchoPilot은 단 하나의 포인트 클릭과 해부학적 범주 이름만으로, 희소한 첫 번째 프레임 상호 작용을 통해 작동합니다. EchoPilot은 의미 기반 위치 설정을 위한 정적인 의료 비전-언어 모델(VLM), 밀집된 기하학적 특징 추출을 위한 비전 기초 모델(VFM), 그리고 마스크 예측 및 전파를 위한 프롬프트 가능한 비디오 분할기를 통합합니다. 초기화의 모호성을 해결하기 위해, 우리는 매개변수가 없는 S.E.E.D.(Semantic Energy-Entropy Density) 기준을 사용하여 최적의 문맥 뷰를 선택하고, 추가적인 사용자 상호 작용 없이 밀집된 기초 특징으로부터 기하학적으로 정확한 보조 포인트 프롬프트를 합성하는 스케일 공간 의미 프롬프팅(Scale-Space Semantic Prompting)을 제안합니다. 전파 드리프트를 줄이기 위해, 신뢰도가 낮은 예측에 대해 분할기의 메모리 뱅크를 선택적으로 동결하는 신뢰도 기반 메모리 업데이트(Reliability-Gated Memory update)를 추가하여 오류 누적을 방지합니다. 또한, 우리는 671개의 주석이 달린 프레임으로 구성된 최초의 동적 태아 양막 초음파 비디오 분할 데이터셋을 제공합니다. 세 개의 초음파 비디오 데이터셋에 대한 실험 결과, EchoPilot은 희소한 상호 작용 환경에서 최첨단 성능을 보여주며, 학습이 필요 없는 기본 모델 및 미세 조정된 전문 시스템보다 일관되게 우수한 성능을 보입니다.
Ultrasound video segmentation is clinically valuable yet difficult due to speckle noise, weak boundaries, and rapid anatomical deformation. Recent promptable foundation models enable point-guided segmentation, but their direct deployment in ultrasound remains unreliable: a single point provides insufficient spatial context to resolve scale ambiguity, and greedy memory updates amplify early errors into severe temporal drift. We present EchoPilot, a training-free framework for ultrasound video segmentation under sparse first-frame interaction, requiring only a single point click and an anatomical category name. EchoPilot orchestrates a frozen medical vision-language model (VLM) for semantic localization, a vision foundation model (VFM) for dense geometric feature extraction, and a promptable video segmentor for mask prediction and propagation. To resolve initialization ambiguity, we propose Scale-Space Semantic Prompting, which first selects an optimal contextual view via a parameter-free S.E.E.D. (Semantic Energy-Entropy Density) criterion, and then synthesizes geometrically precise auxiliary point prompts from dense foundation features without additional user interaction. To reduce propagation drift, a Reliability-Gated Memory update is further introduced to selectively freeze the segmentor's memory bank under uncertain predictions, preventing error accumulation. We also contribute the first dynamic fetal placenta ultrasound video segmentation dataset with 671 annotated frames. Across three ultrasound video datasets, EchoPilot achieves state-of-the-art performance under the sparse-interactive setting, consistently outperforming training-free baselines and finetuned specialists.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.