2607.28466v1 Jul 30, 2026 cs.AI

28만 건의 정기 검사 보고서를 기반으로 한 대장내시경 영상-언어 기초 모델

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

Yan Zhu
Yan Zhu
Citations: 22
h-index: 3
Zilong Wang
Zilong Wang
Citations: 825
h-index: 13
Xinyang Jiang
Xinyang Jiang
Citations: 257
h-index: 8
Fei Wu
Fei Wu
Citations: 290
h-index: 9
Peiyao Fu
Peiyao Fu
Citations: 19
h-index: 3
Xian Yang
Xian Yang
Citations: 8
h-index: 2
Pinghong Zhou
Pinghong Zhou
Citations: 73
h-index: 5
Shuo Wang
Shuo Wang
Citations: 68
h-index: 4
Ruijie Yang
Ruijie Yang
Citations: 7
h-index: 2
Zhihua Wang
Zhihua Wang
Citations: 8
h-index: 2
Jia Yu
Jia Yu
Citations: 29
h-index: 3
Yili He
Yili He
Citations: 15
h-index: 1
Tianyi Chen
Tianyi Chen
Citations: 5
h-index: 1
Siyuan Li
Siyuan Li
Citations: 0
h-index: 0
Quanlin Li
Quanlin Li
Citations: 21
h-index: 3

대장내시경 분야에서, 풍부한 전문가 설명이 기록된 정기 보고서에도 불구하고, 영상-언어 모델은 아직 활용도가 낮은 실정입니다. 이러한 보고서는 병변의 모양, 크기 및 위치를 설명하지만, 전체 절차를 요약하는 방식으로 작성되어 개별 프레임에 대한 설명을 제공하지 않으므로, 임상적 발견물과 해당 이미지 간의 연결이 미약합니다. 본 연구에서는 280,476건의 정기 대장내시경 기록에서 추출한 125,756개의 병변 수준의 이미지-텍스트 쌍을 사용하여 학습된 대장내시경 영상-언어 기초 모델인 EndoCLIP을 개발했습니다. EndoCLIP은 병변 수준의 이미지-텍스트 검색, 구조화된 보고서 생성 및 6가지 다기관 임상 분류 작업에서 일반적인 용도 및 생물 의학 영상-언어 인코더보다 뛰어난 성능을 보였습니다. 특히 양성-악성 분류 작업에서는 전문가 독자가 수행한 검사 결과와 유사한 성능을 보이는 선형 탐지 기능을 보여주었습니다. 이러한 결과는 보고서 내의 발견물과 프레임 간의 대응 관계를 파악함으로써, 정기적인 기록을 확장 가능한 지도 학습 데이터로 활용할 수 있으며, 이를 통해 임상적 목표를 개별적으로 주석 처리하는 대신 언어적으로 명시할 수 있음을 시사합니다.

Original Abstract

Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and six multi-centre clinical classification tasks, EndoCLIP outperforms general-purpose and biomedical vision-language encoders in both zero-shot and linear-probe settings. On benign-versus-malignant classification, its linear probe approaches the performance of expert readers in a blinded study involving 12 endoscopists. These results suggest that recovering finding-to-frame correspondence can transform routine documentation into scalable supervision, enabling clinical targets to be specified in language rather than separately annotated for each task.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!