의료 영상 언어 모델의 모델 편집 기술 평가 및 이해
Evaluating and Understanding Model Editing for Medical Vision Language Models
모델 편집은 의료 영상 언어 모델(VLM)에서 배포 후 발생하는 오류를 비용이 많이 드는 재학습 없이 빠르고 정확하게 수정할 수 있는 방법입니다. 그러나 기존 다중 모드 모델 편집 벤치마크는 일반적인 작업에 초점을 맞추고 있으며, 실제 임상 분야의 요구 사항과 변동성을 제대로 반영하지 못합니다. 이를 해결하기 위해, 우리는 이미지와 텍스트의 변화, 모달리티 및 프로토콜 변경, 임상 지식 구성, 시간적 진행 등 다양한 어려움 하에서 편집된 모델이 얼마나 신뢰성 있고 정확하며 일반화 가능한지를 평가하는 임상 기반 다중 모드 모델 편집 벤치마크인 M3Bench를 소개합니다. M3Bench는 다양한 해부학 구조, 모달리티 및 전문 분야에 걸쳐 16,276개의 질문을 포함하고 있으며, 단일 편집과 순차적 편집 모두를 지원합니다. 우리는 6개의 의료 및 일반 VLM에서 4가지 대표적인 편집 기술을 평가한 결과, 어떤 방법도 모든 기준에서 뛰어난 성능을 보이지 않는다는 것을 확인했습니다. 기울기 기반 편집기는 높은 전이율을 달성하지만, 심각한 지역성 위반 문제를 일으키는 반면, 메모리 기반 방법은 지역성을 유지하지만, 구성적인 일반화 능력이 부족하고 백본 모델에 따라 하이퍼파라미터 민감도가 높습니다. 우리는 이러한 실패의 원인을 VLM의 잠재 공간 기하학적 구조와 다양한 편집 기술이 이 구조를 어떻게 변화시키는지 분석했습니다. 전반적으로, M3Bench는 다중 모드 모델 편집에 대한 엄격한 임상 기반 테스트를 제공하며, 보다 안전한 배포 후 적응을 위한 실질적인 지침을 제시합니다. 벤치마크는 https://github.com/BioMed-AI-Lab-U-Michgan/M3Bench 에서 공개적으로 이용할 수 있습니다.
Model editing promises a fast, targeted way to correct post-deployment mistakes in medical vision-language models (VLMs) without costly retraining. However, existing multimodal model editing benchmarks focus on general-purpose tasks and do not reflect realistic clinical domain requirements and variability. To address this, we introduce M3Bench, a clinically grounded benchmark for multimodal model editing that evaluates whether an edit remains reliable, precise, and generalizable under the challenges of image and text variation, modality and protocol shifts, clinical knowledge composition, and temporal progression. M3Bench contains 16,276 questions spanning diverse anatomy, modalities, and specialties, and supports both single and sequential edits. By evaluating 4 representative editors across 6 medical and general VLMs, we find that no method excels across all criteria. Gradient-based editors achieve strong transfer but suffer from catastrophic locality violations, whereas memory-based methods preserve locality but lack compositional generality and exhibit high backbone-dependent hyperparameter sensitivity. We further attribute these failures to the latent space geometry of VLMs and how different editing methods shift its landscape. Overall, M3Bench establishes a rigorous clinical stress test for multimodal model editing and offers actionable guidance for safer post-deployment adaptation. The benchmark is publicly available at https://github.com/BioMed-AI-Lab-U-Michgan/M3Bench .
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.