어려운 실제 임상 사례에 대한 다중 라운드 멀티모달 진단 추론 평가
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
임상 진단 평가는 모델이 정확한 진단을 내릴 수 있는지 뿐만 아니라, 점진적인 다중 정보 공개, 역동적인 진단 가설 업데이트 및 지속적인 임상 추론 개선을 포함하는 실제 임상 환경의 현실을 반영해야 합니다. 그러나 기존의 멀티모달 대규모 언어 모델(MLLM) 평가 방법은 주로 단일 라운드 또는 독립적인 작업에 의존하여, 실제 임상 진단의 복잡성을 완전히 파악하기 어렵습니다. 이러한 격차를 해소하기 위해, 현재까지 가장 큰 규모의 다중 라운드 멀티모달 임상 진단 평가 벤치마크인 ClinMM-Bench를 개발했습니다. ClinMM-Bench는 8개의 전문 분야에 걸쳐 1,089개의 어려운 실제 임상 사례와 3,760개의 의료 이미지를 포함합니다. 우리는 두 가지 수준의 평가 프레임워크를 사용하여 15개의 대표적인 MLLM을 체계적으로 평가하여 진단 정확도와 진단 추론 품질 모두를 측정했습니다. 결과는 독점 모델이 전반적인 진단 정확도가 가장 높았지만, 모든 모델에서 완전히 정확한 진단을 내리는 비율은 여전히 제한적이라는 것을 보여주었습니다. 진단 추론 품질 측면에서, 현재의 모델은 타당한 진단 방향을 식별할 수 있지만 신뢰할 수 있는 진단 추론을 생성하는 데는 상당한 한계가 있습니다. 추가적인 오류 분석 결과, 정보 통합 실패, 지식 매핑 오류, 인식 오류, 조기 결론 도출 및 시각적 환각의 5가지 대표적인 오류 유형이 확인되었습니다.
Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on single-turn or isolated tasks, making it difficult to fully capture the complexity of real-world clinical diagnosis. To bridge this gap, we developed ClinMM-Bench, the largest multi-turn multimodal clinical diagnostic evaluation benchmark to date. ClinMM-Bench contains 1,089 challenging real-world clinical cases and 3,760 medical images across eight specialties. We systematically evaluated 15 representative MLLMs using a two-level evaluation framework that assessed both diagnostic accuracy and diagnostic reasoning quality. Results showed that proprietary models achieved the highest overall diagnostic accuracy, but the proportion of completely correct diagnoses remained limited across all models. In terms of diagnostic reasoning quality, current models can identify plausible diagnostic directions but still have considerable limitations in generating reliable diagnostic reasoning. Error analysis further identified five representative failure modes: information synthesis failure, knowledge mapping error, perception error, premature closure, and visual hallucination.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.