2608.04472v1 Aug 05, 2026 cs.CV

EndoVLM: 해부학적 지침 기반 희소성 및 점진적 정렬을 통한 내시경 영상-언어 사전 학습 모델

EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment

Zhongwei Qiu
Zhongwei Qiu
Citations: 92
h-index: 3
Yingda Xia
Yingda Xia
Citations: 127
h-index: 6
Ling Zhang
Ling Zhang
Citations: 1,029
h-index: 15
Sijing Li
Sijing Li
Citations: 133
h-index: 3
Zhenyu Yi
Zhenyu Yi
Citations: 25
h-index: 2
Bin Lv
Bin Lv
Citations: 0
h-index: 0
Jianwei Xu
Jianwei Xu
Citations: 5
h-index: 1
Yue Hu
Yue Hu
Citations: 6
h-index: 1
Liang Huang
Liang Huang
Citations: 5
h-index: 1

내시경 이미지 분석 발전을 위해서는 기초 모델(Foundation Models, FMs)의 개발이 매우 중요합니다. 그러나 기존 내시경 기초 모델은 주로 단일 모드 이미지 또는 비디오에 대한 자기 지도 학습에 의존하며, 임상 보고서에 포함된 풍부한 의미 정보를 간과합니다. 또한, 이러한 기록을 효과적으로 활용하는 것은 근본적인 양식(modality) 간의 격차로 인해 어려움을 겪습니다. 즉, 구조화된 해부학적 설명이 높은 중복성을 가진 비정형 시각 데이터 내의 특정 프레임에 자연스럽게 매핑되지 않기 때문입니다. 본 논문에서는 348,000건 이상의 내시경 검사에 대해 사전 학습된 새로운 영상-언어 기초 모델인 EndoVLM을 제안합니다. 각 검사는 임상 보고서와 해당 이미지 컬렉션으로 구성됩니다. 해부학적 지침 기반 희소 풀링(Anatomy-Guided Sparse Pooling) 메커니즘은 텍스트 설명을 쿼리로 사용하여 희소 어텐션을 유도하며, 이를 통해 중복된 이미지 세트 내에서 의미적으로 중요한 프레임을 효율적으로 집계하여 해부학적 특징에 특화된 시각적 표현을 생성합니다. 다음으로, 점진적인 의미 인식 정렬(Progressive Semantic-Aware Alignment) 전략은 구조화된 소프트 타겟을 사용하여 임상 분류 체계(해부학과 병리 상태)를 모델링함으로써, 전체 환자 수준의 매칭에서부터 세밀한 위치 기반 정렬까지 양식 간 격차를 해소합니다. 마지막으로, 의미적으로 풍부한 프레임에 Semantic-Concentrated Masked Autoencoder를 적용하여 저수준 시각적 정확성과 견고한 고수준 의미 표현을 통합합니다. 다양한 하위 작업에서의 광범위한 실험 결과는 EndoVLM이 기존 기초 모델보다 우수한 성능을 보이며, 특정 작업에 특화된 방법과도 경쟁력이 있음을 보여줍니다. 주목할 만한 점은 EndoVLM이 뛰어난 제로샷 일반화 능력을 보여주어, 더 넓은 임상 적용 가능성을 제시한다는 것입니다.

Original Abstract

The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-supervised learning from uni-modal images or videos, overlooking the rich semantic knowledge contained in clinical reports. Furthermore, effectively leveraging these records is hindered by a fundamental modality gap: structured anatomical descriptions are not naturally mapped to specific frames within the high-redundancy, uncurated visual streams. In this paper, we present EndoVLM, a novel vision-language FM pre-trained on over 348K endoscopic examinations, each pairing a clinical report with its corresponding image collection. An Anatomy-Guided Sparse Pooling mechanism utilizes textual descriptions as queries to drive sparse attention, efficiently aggregating semantically salient frames into anatomy-specific visual representations across redundant image-sets. Next, a Progressive Semantic-Aware Alignment strategy models clinical taxonomy (anatomy and pathological status) via structured soft targets, bridging the gap from global patient-level matching to fine-grained localized alignment. Finally, a Semantic-Concentrated Masked Autoencoder is applied exclusively to these semantic-rich frames, integrating low-level visual precision with robust high-level semantic representation. Extensive experiments across various downstream tasks demonstrate that EndoVLM outperforms existing foundation models and remains competitive with task-specific methods. Remarkably, EndoVLM also exhibits robust zero-shot generalization capabilities, highlighting its potential for broader clinical application.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!