2606.17710v1 Jun 16, 2026 cs.CV

흉부 방사선 촬영을 위한 비전-언어 모델은 항상 이미지가 필요하지 않다.

Vision-language models for chest radiography do not always need the image

Soroosh Tayebi Arasteh
Soroosh Tayebi Arasteh
Citations: 1,026
h-index: 12
D. Truhn
D. Truhn
Citations: 8,091
h-index: 48
Mahshad Lotfinia
Mahshad Lotfinia
Citations: 231
h-index: 9
S. Ziegelmayer
S. Ziegelmayer
Citations: 668
h-index: 15
L. Adams
L. Adams
Citations: 780
h-index: 15
Andreas K. Maier
Andreas K. Maier
Citations: 154
h-index: 5

의료 분야의 비전-언어 모델들은 흉부 방사선 사진 분석에서 높은 정확도를 보이는 것으로 보고되지만, 이는 모델이 실제로 이미지를 활용한다는 증거로 해석되는 경우가 많습니다. 하지만 이러한 추론은 위험합니다. 특정 질병명에 대한 사전 지식을 활용하는 모델과 실제 스캔 데이터를 읽는 모델의 성능을 구별할 수 있는 표준적인 벤치마크가 존재하지 않기 때문입니다. 본 연구에서는 인과 관계 분석을 통해 이미지가 결과에 미치는 영향을 평가하기 위해, 이미지 영역을 가리고, 관련 없는 영역을 가리고, 다른 환자의 동일한 질병명으로 분류된 스캔 이미지를 교체하는 실험을 수행했습니다. 또한, 세 가지 행동 지표를 결합하여 모델이 올바른 답변을 내리는 데 이미지의존성이 있는지 테스트했습니다. 9개의 시스템을 분석한 결과, 이미지에 접근할 수 없는 텍스트 기반 모델이 가장 성능이 좋은 멀티모달 모델과 5.7개 이상의 정확도 포인트 차이를 벌리지 않았으며, 1190억 개의 파라미터를 가진 멀티모달 모델은 70억 개의 파라미터를 가진 텍스트 기반 모델과 통계적으로 유의미한 차이가 없었습니다. 본 연구는 시스템을 세 가지 그룹으로 분류했습니다: 이미지를 무시하는 모델, 불안정한 모델, 그리고 특정 질병에 대해서만 이미지를 선택적으로 사용하는 모델. 이러한 구분은 다른 데이터셋, 해상도 및 프롬프트 표현 방식에서도 유지되었습니다. 숙련된 방사선 전문의와 비교했을 때, 텍스트 기반 모델은 방사선 전문의의 정확도와 통계적으로 유의미한 차이가 없었으며, 이미지 활용 모델들은 방사선 전문의 수준과 비슷한 성능을 보였습니다. 모델이 이미지를 사용하는 경우에만 신뢰도 플래그가 설정되어 부정확한 답변을 나타냈습니다. 임상 적용 시에는 정확도가 아닌, 데이터 근거 여부를 기준으로 결정해야 합니다.

Original Abstract

Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image. That inference is unsafe: a model exploiting finding-name priors scores like one that reads the scan, and no standard benchmark separates them. We introduce a causal audit that intervenes on the image, occluding the relevant region, occluding an irrelevant one, and swapping in another patient's same-label scan, and combines three behavioral metrics to test whether a correct answer depends on the image. Across nine systems, a text-only model with no image access reaches within 5.7 accuracy points of the best multimodal one, and a 119-billion-parameter multimodal model is statistically indistinguishable from a 7-billion text-only baseline. The audit splits the cohort into three models that ignore the image, one that is unstable, and five that use it selectively, for a subset of findings; the categories hold across a second dataset, resolution, and prompt phrasing. Against board-certified radiologists, a text-only model is statistically indistinguishable from a radiologist's accuracy while grounding at zero, whereas the image-using models ground at radiologist-comparable rates. Reported confidence flags ungrounded answers only when a model uses the image. Grounding audits, not accuracy, should gate clinical deployment.

1 Citations
0 Influential
24 Altmetric
121.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!