AIPatient Arena: 전자의무기록(EHR) 기반 대규모 언어 모델의 종단 간 임상 상담 워크플로우 평가
AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows
대규모 언어 모델(LLM)은 임상 상담 작업에 점점 더 많이 활용되고 있지만, 대부분의 의료 평가는 정적인 단일 회전 방식으로 이루어지거나 제한된 결과 지표에만 초점을 맞춰 실제 의료 환경의 순차적이고 불확실하며 상호작용적인 특성을 제대로 반영하지 못합니다. 본 연구에서는 전자의무기록(EHR) 기반으로 LLM의 임상 유용성을 8가지 임상 역량 차원에서 평가하는 AIPatient Arena라는 평가 프레임워크를 제안합니다. 이 프레임워크는 EHR 데이터를 활용하여 환자별 지식 그래프를 구축하고, 이를 통해 다회원 의료진-환자 상호작용을 가능하게 합니다. 우리는 437명의 주요 환자 그룹과 119명 및 67명의 추가 검증 그룹에 AIPatient Arena를 적용했습니다. 그 결과, LLM은 의학적 면담 질문 능력(QS; 평균 점수 4.43-4.99/5), 윤리적 및 전문적 태도(ET; 4.38-4.93/5), 임상 설명의 명확성과 투명성(EX; 3.80-4.72/5) 측면에서 우수한 성능을 보였습니다. 정보 통합 능력(II; 3.19-4.21/5) 및 약물 안전 및 근거 제시 능력(MS; 3.13-3.78/5)은 중간 수준이었지만, 모호한 환자 응답 처리(HR; 2.57-3.32/5), 정보 포괄성(IC; 2.08-3.02/5), 그리고 진단 정확도 및 추론 능력(Dx; 2.63-3.55/5)에서는 지속적인 약점을 보였습니다. 프로세스 기반 평가는 반복적인 질문, 과거 병력 누락, 불확실성 처리 부족 등과 같은 상호작용 실패 사례를 드러냈습니다. 풍부한 대화 맥락은 진단 추론 능력을 향상시키는 데 도움이 되었지만, 치료 계획 수립에는 큰 효과가 없었습니다. 이러한 결과는 최종 답변 정확도만으로는 임상 적용 준비 상태를 평가하는 데 충분하지 않으며, 모델이 상담 과정 전반에 걸쳐 정보를 어떻게 수집하고 해석하며 전달하는지를 평가하는 것이 중요하다는 점을 시사합니다. AIPatient Arena는 의료 LLM의 워크플로우 기반 사전 배포 평가를 위한 EHR 기반 프레임워크를 제공합니다.
Large language models (LLMs) are increasingly considered for use in clinical consultation tasks, yet most medical evaluations remain static, single-turn, or narrowly outcome-based, limiting their ability to reflect the sequential, uncertain, and interactive nature of real-world care. Here, we propose AIPatient Arena, an EHRs-grounded evaluation framework for assessing the clinical utility of LLMs across eight dimensions of clinical competence. The framework integrates EHR data into patient-specific knowledge graphs, enabling multi-turn physician-patient interactions. We applied AIPatient Arena on a primary cohort of 437 patients and two out-of-distribution validation cohorts of 119 and 67 patients. We observe that LLMs performed well in medical interview questioning skills (QS; mean scores, 4.43-4.99/5), ethical and professional conduct (ET; 4.38-4.93/5), and clarity and transparency of clinical explanations (EX; 3.80-4.72/5). Performance was moderate in information integration (II; 3.19-4.21/5) and medication safety and justification (MS; 3.13-3.78/5), but persistent weaknesses were observed in handling of ambiguous patient responses (HR; 2.57-3.32/5), information coverage (IC; 2.08-3.02/5), and diagnostic accuracy and reasoning (Dx; 2.63-3.55/5). Process-based evaluation revealed recurrent interaction failures, including repetitive questioning, omission of past medical history, and inadequate handling of uncertainty. Richer conversational context improved diagnostic reasoning but yielded limited gains in treatment planning. These findings indicate that final-answer accuracy alone is insufficient for evaluating clinical readiness and highlight the importance of assessing how models gather, interpret, and communicate information throughout a consultation. AIPatient Arena provides an EHR-grounded framework for workflow-oriented pre-deployment evaluation of medical LLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.