시각적 증거 분석을 넘어서: 합성 의료 이미지 탐지를 위한 다중 모드 강건성 감사
Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection
생성형 인공지능의 급속한 발전으로 인해 합성 의료 이미지는 진단 오류 및 보험 사기 등 심각한 위험을 초래합니다. 기존 연구에서는 비전-언어 모델(VLM) 기반의 합성 이미지 탐지를 다루었지만, 이러한 평가는 일반적으로 개별 이미지를 대상으로 합니다. 그러나 실제 임상 환경에서는 이미지가 구조화된 기록 및 메타데이터와 함께 해석되며, VLM은 점점 더 많은 경우 이미지와 기록을 결합한 입력으로 사용됩니다. 본 연구에서는 아직 충분히 조사되지 않은 다중 모드 취약점을 발견했습니다. 즉, 두 가지 모드를 모두 사용할 때 VLM은 진위 판단 시 기록의 맥락에 지나치게 의존할 수 있으며, 이로 인해 동일한 이미지라도 동반되는 텍스트가 변경됨에 따라 예측 결과가 달라질 수 있습니다. 이는 실제 환경에서의 적용 가능성에 대한 우려를 불러일으킵니다. 이러한 현상을 체계적으로 분석하기 위해, 본 연구는 합성 의료 이미지 탐지를 이미지-기록 인터페이스에서의 다중 모드 강건성 감사로 재정의하고, 이미지는 고정한 채 메타데이터 변형을 교환하는 페어링 벤치마크를 도입했습니다. 다양한 영상 모달리티에서 다양한 공개 모델 및 최첨단 API 기반 VLM을 평가하여 메타데이터만 변경했을 때 진위 예측이 어떻게 달라지는지 정량적으로 분석했습니다. 본 벤치마크는 이미지만을 고려한 환경을 넘어 다중 모드 강건성을 평가하고 개선하는 데 유용한 표준 도구를 제공합니다. 관련 코드는 https://github.com/chiuhaohao/Beyond-Visual-Forensics 에서 확인할 수 있습니다.
With the rapid adoption of generative AI, synthetic medical images pose growing risks, including diagnostic deception and insurance fraud. Although prior work has explored vision-language model (VLM)-based synthetic image detection, these evaluations typically consider images in isolation. In clinical practice, however, images are interpreted alongside structured records and metadata, and VLMs are increasingly deployed under joint image-record inputs. We uncover a previously underexamined multimodal vulnerability: when given both modalities, VLMs may overweight record context in authenticity judgments, such that the same image receives different predictions solely due to changes in its accompanying text. This raises concerns about robustness in real-world deployment. To systematically characterize this effect, we reformulate synthetic medical image detection as an audit of multimodal robustness at the image-record interface and introduce a paired benchmark that holds the image fixed while swapping controlled metadata variants. Across multiple imaging modalities, we evaluate diverse open-weight and frontier API VLMs and quantify how metadata alone shifts authenticity predictions. Our benchmark provides a standardized tool for assessing and improving multimodal robustness beyond image-only settings. The code is available at https://github.com/chiuhaohao/Beyond-Visual-Forensics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.