2608.14262v1 Aug 14, 2026 cs.CV

수술 내시경 비디오에 대한 시간 기반 시각-언어 모델의 강건성 연구

On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Imran Razzak
Imran Razzak
Citations: 7
h-index: 2
Mohammad Yaqub
Mohammad Yaqub
Citations: 353
h-index: 11
Ufaq Khan
Ufaq Khan
Citations: 9
h-index: 2
Muhammad Bilal
Muhammad Bilal
Citations: 31
h-index: 3
Muhammad Haris Khan
Muhammad Haris Khan
Citations: 12
h-index: 2
B. Lall
B. Lall
Citations: 628
h-index: 14
Dwarikanath Mahapatra
Dwarikanath Mahapatra
Citations: 92
h-index: 5
Darakshan Rashid
Darakshan Rashid
Citations: 0
h-index: 0
Raza Imam
Raza Imam
Citations: 49
h-index: 4
Shazad Ashraf
Shazad Ashraf
Citations: 306
h-index: 2
Lena Maier-Hein
Lena Maier-Hein
Citations: 32
h-index: 3

시간 기반 시각-언어 모델(TVLM)은 수술 비디오 이해를 위한 재사용 가능하고 프롬프트 기반 인터페이스를 제공하지만, 실제 임상 환경에서 발생하는 내시경 영상의 문제로 인한 TVLM의 강건성은 충분히 연구되지 않았습니다. 실제로 흐림, 안개, 움직임 블러, 노이즈, 전기 소작 연기, 패킷 손실과 같은 요소들은 비디오-텍스트 정렬을 저해할 수 있는 구조적인 분포 변화를 야기합니다. 본 연구에서는 클립 프레임의 왜곡으로 인해 발생하는 이러한 변화에 대한 TVLM의 강건성을 분석합니다. 우리는 Endo-C6이라는 6가지 실제 내시경 환경에서 발생 가능한 문제점을 포함하는 데이터셋을 개발하고, 이를 사용하여 공개된 위장관(GI) 내시경 및 복강경 담낭 절제술 비디오를 평가합니다. 표준화된 프롬프트 프로토콜을 사용하여 최근 개발된 3개의 TVLM 모델을 비교 분석하고, 평균 성능과 최악의 경우 성능 모두에서 294개의 데이터셋 수준으로 강건성을 평가했습니다. 그 결과, VeRA를 활용한 소량의 파라미터 조정만으로 RobustEndoCLIP이라는 모델을 개발하여 기존 TVLM 모델보다 우수한 성능을 달성했습니다. 연구 결과는 상용 TVLM 모델이 내시경 영상에 특화된 문제로 인해 심각한 성능 저하를 보일 수 있지만, 프롬프트 기반 인터페이스를 변경하지 않고도 경량의 소량 데이터 학습(few-shot)을 통해 왜곡된 영상에서의 성능과 강건성을 크게 향상시킬 수 있음을 보여줍니다. Endo-C6 데이터셋은 표준화된 강건성 보고를 지원하고 더 신뢰할 수 있는 임상 시각-언어 시스템 개발에 기여할 것으로 기대됩니다.

Original Abstract

Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition artifacts in endoscopy remains insufficiently characterized. In practice, degradations such as defocus, haze, motion blur, noise, cautery smoke, and packet loss introduce structured distribution shifts which may compromise video-text alignment. We study the robustness of temporal VLMs under such shifts caused by corruptions in clip frames. We introduce Endo-C6, a compact corruption benchmark of six endoscopy-realistic perturbations evaluated at a fixed high severity, and apply it to public Gastrointestinal (GI) endoscopy and laparoscopic cholecystectomy videos. Under a standardized prompt protocol, we benchmark 3 recent surgical TVLM baselines and analyze robustness in both mean and worst-case settings, spanning 294 dataset-level evaluations. Finally, we present RobustEndoCLIP, obtained by few-shot parameter-efficient tuning with VeRA, outperforming existing TVLM baselines. Our findings show that off-the-shelf TVLMs can exhibit severe worst-case collapse under endoscopy-specific corruptions, whereas lightweight few-shot adaptation can substantially improve corrupted performance and robustness without changing the prompt-based interface. We expect Endo-C6 to support standardized robustness reporting and promote more reliable clinical vision-language systems.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!