의료 LLM 진단에서의 증거 활용 감사
Auditing Evidence Use in Medical LLM Diagnosis
의료 LLM은 종종 올바른 진단을 선택하는지 여부를 기준으로 평가받지만, 진단 정확도만으로는 모델이 사례 증거를 적절하게 사용했는지 알 수 없습니다. 본 연구에서는 의료 진단에서 증거 활용에 대한 행동 감사(behavioral audit)를 제시합니다. 각 사례에 대해 환자 정보를 증거 단위로 분해하고, 통제된 증거 집합 하에서 후보 진단을 평가하며, 진단 경계면에서의 낮은 수준의 상호 작용을 분석합니다. 의료 증거는 진단과 관련이 있으므로, 본 감사는 상호 작용 발견과 오류 할당을 분리합니다. 큰 또는 부정적인 상호 작용은 합리적인 감별 진단을 반영할 수 있으며, 의심스러운 상호 작용은 견고성 검사 및 임상 검토가 필요합니다. 우리는 DDXPlus, CupCase, MedCase 데이터셋에서 공개된 5개의 LLM을 평가했습니다. 데이터셋 전반적으로, 신뢰할 수 있는 지지와 차별적인 충돌 또는 상쇄가 대부분의 상호 작용 강도를 설명하며, 이는 많은 증거 상호 작용이 실패라기보다는 임상적으로 타당하다는 것을 보여줍니다. DDXPlus에 초점을 맞춘 익명 5명의 검토자가 참여한 130개 항목으로 구성된 풍부한 검토 샘플에서, 잘못되거나 단순화된 사례는 주로 부정적이거나 부재인 소견과 임상적으로 제한적인 증거와 관련되어 있습니다. 이러한 결과는 정확도가 잠재적인 증거 활용 실패를 숨길 수 있으며, 의료 LLM 평가를 위한 역할 기반 감사(role-aware audits)의 필요성을 강조합니다.
Medical LLMs are often evaluated by whether they select the correct diagnosis, but diagnostic accuracy alone does not show whether the model used the case evidence appropriately. We present a behavioral audit of evidence use in medical diagnosis. For each case, we decompose patient information into evidence units, score candidate diagnoses under controlled evidence subsets, and mine low-order interactions in diagnostic margins. Because medical evidence is diagnosis-relative, the audit separates interaction discovery from failure assignment: large or negative interactions can reflect plausible differential diagnosis, while suspicious interactions require robustness checks and clinical review. We evaluate five open-weight LLMs on DDXPlus, CupCase, and MedCase. Across datasets, faithful support and differential conflict or cancellation account for most interaction strength, showing that many evidence interactions are clinically plausible rather than failures. In a DDXPlus-focused blinded five-reviewer 130-item enriched review sample, invalid or shortcut-like cases concentrate in negated or absent findings and clinically local evidence. These results show that accuracy can hide candidate evidence-use failures and motivate role-aware audits for medical LLM evaluation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.