2604.27720v1 Apr 30, 2026 cs.AI

신뢰성 있는 의료 질의응답을 위한 최첨단 시각-언어 모델 감사: 오류 발생 원인, 형식 붕괴 및 도메인 적응

Auditing Frontier Vision-Language Models for Trustworthy Medical VQA: Grounding Failures, Format Collapse, and Domain Adaptation

Binbin Shi
Binbin Shi
Citations: 8
h-index: 2
Chenqian Le
Chenqian Le
Citations: 38
h-index: 3
Ran Gong
Ran Gong
Citations: 110
h-index: 3
Qifu Yin
Qifu Yin
Citations: 54
h-index: 1
Haowei Ni
Haowei Ni
Citations: 171
h-index: 5
Lang Lin
Lang Lin
Citations: 12
h-index: 2
Panfeng Li
Panfeng Li
Citations: 18
h-index: 2
Xupeng Chen
Xupeng Chen
Citations: 481
h-index: 12

임상 환경에서 시각-언어 모델(VLM)을 사용할 때 현실적인 오류 조건에서도 검증 가능한 성능이 요구되지만, 특정 의료 데이터에 대한 최첨단 VLM의 오류 현상에 대한 이해는 부족합니다. 본 연구에서는 최근 개발된 5가지 최첨단 VLM(Gemini~2.5~Pro, GPT-5, o3, GLM-4.5V, Qwen~2.5~VL)을 의료 질의응답(VQA) 시스템에서 신뢰성과 관련된 두 가지 측면에서 분석했습니다. 첫째, 인지 능력 측면에서 모든 모델이 해부학적 및 병리학적 특징을 정확하게 식별하는 데 어려움을 겪으며, 최고 성능 모델조차 평균 IoU 0.23, 정확도 19.1%를 기록했으며, 임상적으로 위험할 수 있는 좌우 혼동 현상을 보였습니다. 둘째, 파이프라인 통합 측면에서 동일한 모델이 위치 정보 추출 후 답변을 생성하는 방식은 모든 모델의 VQA 정확도를 저하시키는데, 이는 부정확한 위치 정보 추출과 함께 두 단계 프롬프트에서 발생하는 형식 준수 오류(VQA-RAD 데이터셋에서 70%~99%에 이르는 높은 오류율) 때문입니다. 예측된 위치 정보를 실제 위치 정보로 대체하면 VQA 정확도가 회복되고 향상되는데, 이는 오류가 위치 정보 추출 모듈에 있다는 것을 시사합니다. 이러한 분석 결과는 본 연구에서 사용하는 SLAKE 바운딩 박스 설정에서 위치 정보 품질이 신뢰성을 저해하는 주요 요인임을 보여줍니다. 추가적으로, Qwen~2.5~VL 모델을 다양한 의료 데이터로 미세 조정했을 때, 동등한 수준의 다른 모델보다 높은 SLAKE 오픈 엔드 재현율(85.5%)을 달성했는데, 이는 의료 VQA 수준의 성능 격차는 도메인 적응을 통해 해결할 수 있음을 시사합니다. 그러나 이러한 개선이 인지 능력 및 신뢰성 문제를 해결하는 데 도움이 되는지는 향후 연구에서 검토할 필요가 있습니다.

Original Abstract

Deploying vision-language models (VLMs) in clinical settings demands auditable behavior under realistic failure conditions, yet the failure landscape of frontier VLMs on specialized medical inputs is poorly characterized. We audit five recent frontier and grounding-aware VLMs (Gemini~2.5~Pro, GPT-5, o3, GLM-4.5V, Qwen~2.5~VL) on Medical VQA along two trust-relevant axes. Perception: all models localize anatomical and pathological targets poorly -- the best model reaches only 0.23 mean IoU and 19.1% Acc@0.5 -- and exhibit clinically dangerous laterality confusion. Pipeline integration: a self-grounding pipeline, where the same model localizes then answers, degrades VQA accuracy for every model -- driven by both inaccurate localization and format-compliance failures under the two-step prompt (parse failure rises to 70%--99% for Gemini and GPT-5 on VQA-RAD). Replacing predicted boxes with ground-truth annotations recovers and improves VQA accuracy, consistent with the failure residing in the perception module rather than in the decomposition itself. These observational findings identify grounding quality as a primary trustworthiness bottleneck in our SLAKE bounding-box setting. As a complementary fine-tuning follow-up, supervised fine-tuning of Qwen~2.5~VL on combined Med-VQA training data attains the highest reported SLAKE open-ended recall (85.5%) among comparable methods, suggesting that the VQA-level gap is tractable with domain adaptation; whether this also closes the perception/trustworthiness bottleneck is left to future work.

1 Citations
0 Influential
6 Altmetric
31.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!