기초 모델을 활용한 시각-언어 모델 기반 반복 정제: 수술 영상 분할
VLM-Guided Iterative Refinement for Surgical Image Segmentation with Foundation Models
수술 영상 분할은 로봇 보조 수술 및 수술 중 안내에 필수적입니다. 그러나 기존 방법은 미리 정의된 범주에 제한되며, 적응적 정제 없이 단일 예측을 수행하고, 임상의와의 상호 작용 메커니즘이 부족합니다. 본 연구에서는 자연어 설명을 입력으로 받아 수술 영상 분할을 위한 반복 정제 시스템인 IR-SIS를 제안합니다. IR-SIS는 미세 조정된 SAM3을 사용하여 초기 분할을 수행하고, 시각-언어 모델을 활용하여 수술 기구를 감지하고 분할 품질을 평가하며, 에이전트 기반 워크플로우를 통해 적응적으로 정제 전략을 선택합니다. 또한, 본 시스템은 자연어 피드백을 통한 임상의의 참여를 지원합니다. 또한, EndoVis2017 및 EndoVis2018 벤치마크 데이터셋을 기반으로 다단계의 언어 주석 데이터셋을 구축했습니다. 실험 결과, IR-SIS는 동일 데이터셋 및 외부 데이터셋 모두에서 최첨단 성능을 보였으며, 임상의의 참여는 추가적인 성능 향상을 제공했습니다. 본 연구는 적응적인 자체 정제 기능을 갖춘 최초의 언어 기반 수술 분할 프레임워크를 제시합니다.
Surgical image segmentation is essential for robot-assisted surgery and intraoperative guidance. However, existing methods are constrained to predefined categories, produce one-shot predictions without adaptive refinement, and lack mechanisms for clinician interaction. We propose IR-SIS, an iterative refinement system for surgical image segmentation that accepts natural language descriptions. IR-SIS leverages a fine-tuned SAM3 for initial segmentation, employs a Vision-Language Model to detect instruments and assess segmentation quality, and applies an agentic workflow that adaptively selects refinement strategies. The system supports clinician-in-the-loop interaction through natural language feedback. We also construct a multi-granularity language-annotated dataset from EndoVis2017 and EndoVis2018 benchmarks. Experiments demonstrate state-of-the-art performance on both in-domain and out-of-distribution data, with clinician interaction providing additional improvements. Our work establishes the first language-based surgical segmentation framework with adaptive self-refinement capabilities.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.