2603.23501v1 Mar 24, 2026 cs.CV

MedObvious: 임상 트riage를 통한 VLM의 의료 Moravec 역설 분석

MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage

L. S. Teja
L. S. Teja
Citations: 5
h-index: 2
Mohammad Yaqub
Mohammad Yaqub
Citations: 353
h-index: 11
Ufaq Khan
Ufaq Khan
Citations: 9
h-index: 2
Umair Nawaz
Umair Nawaz
Citations: 35
h-index: 3
Numaan Saeed
Numaan Saeed
Citations: 2
h-index: 1
Muhammad Bilal
Muhammad Bilal
Citations: 31
h-index: 3
Yutong Xie
Yutong Xie
Citations: 109
h-index: 5
Muhammad Haris Khan
Muhammad Haris Khan
Citations: 12
h-index: 2

비전 언어 모델(VLM)은 의료 보고서 생성 및 시각적 질의 응답과 같은 작업에 점점 더 많이 사용되고 있습니다. 그러나 유창한 진단 텍스트가 안전한 시각적 이해를 보장하지는 않습니다. 임상 환경에서 해석은 사전 진단 검사를 통해 시작됩니다. 즉, 입력 데이터가 유효한지 확인하는 단계입니다(정확한 모달리티 및 해부학적 구조, 타당한 시점 및 방향, 그리고 눈에 띄는 무결성 위반 여부). 기존의 벤치마크는 이 단계가 해결되었다고 가정하는 경향이 있으며, 따라서 중요한 오류 모드를 간과합니다. 즉, 모델이 입력 데이터가 일관성이 없거나 유효하지 않더라도 타당한 설명을 생성할 수 있습니다. 본 논문에서는 입력 데이터 유효성 검사를 작은 다중 패널 이미지 세트에 대한 집합 수준의 일관성 능력으로 분리하는 1,880개의 작업으로 구성된 벤치마크인 MedObvious를 소개합니다. 모델은 패널 중 하나라도 예상되는 일관성을 위반하는지 식별해야 합니다. MedObvious는 기본적인 방향/모달리티 불일치부터 임상적으로 동기 부여된 해부학적 구조/시점 확인 및 triage 스타일의 단서에 이르기까지 다섯 가지 단계로 구성되어 있으며, 다양한 인터페이스에서의 견고성을 테스트하기 위해 다섯 가지 평가 형식을 포함합니다. 17개의 서로 다른 VLM을 평가한 결과, 사전 검사가 여전히 신뢰할 수 없는 것으로 나타났습니다. 여러 모델이 정상(음성 제어) 입력에서 이상 현상을 환각하거나, 더 큰 이미지 세트로 확장할 때 성능이 저하되며, 객관식 및 자유 응답 형식 간에 측정된 정확도가 크게 다릅니다. 이러한 결과는 사전 진단 검증이 의료 VLM에게는 여전히 해결되지 않은 문제이며, 배포 전에 안전에 중요한 별도의 기능으로 간주되어야 함을 보여줍니다.

Original Abstract

Vision Language Models (VLMs) are increasingly used for tasks like medical report generation and visual question answering. However, fluent diagnostic text does not guarantee safe visual understanding. In clinical practice, interpretation begins with pre-diagnostic sanity checks: verifying that the input is valid to read (correct modality and anatomy, plausible viewpoint and orientation, and no obvious integrity violations). Existing benchmarks largely assume this step is solved, and therefore miss a critical failure mode: a model can produce plausible narratives even when the input is inconsistent or invalid. We introduce MedObvious, a 1,880-task benchmark that isolates input validation as a set-level consistency capability over small multi-panel image sets: the model must identify whether any panel violates expected coherence. MedObvious spans five progressive tiers, from basic orientation/modality mismatches to clinically motivated anatomy/viewpoint verification and triage-style cues, and includes five evaluation formats to test robustness across interfaces. Evaluating 17 different VLMs, we find that sanity checking remains unreliable: several models hallucinate anomalies on normal (negative-control) inputs, performance degrades when scaling to larger image sets, and measured accuracy varies substantially between multiple-choice and open-ended settings. These results show that pre-diagnostic verification remains unsolved for medical VLMs and should be treated as a distinct, safety-critical capability before deployment.

2 Citations
0 Influential
4 Altmetric
22.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!