CliniCARE-Bench: EHR 데이터 기반 의료 추론의 임상적 정확성 평가
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
대규모 언어 모델은 의료 지식 벤치마크에서 뛰어난 성능을 보이지만, 실제 임상 환경에 적용하기 위해서는 시스템이 다양한 형태의 과거 기록들을 바탕으로 신뢰할 수 있는 분석을 수행해야 합니다. 이는 필요한 증거를 결정하고, 구조화된 데이터와 자유 형식 데이터를 검색 및 통합하며, 결론을 검증 가능한 근거로 뒷받침하고, 명확하게 해결되지 않는 경우에는 판단을 유보하는 것을 포함합니다. 본 논문에서는 CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR)라는 임상 감사 벤치마크를 소개합니다. 이는 실제 환자 데이터를 기반으로 한 MIMIC-IV 데이터셋을 활용하여, 의료 전문가가 검증한 25개의 시나리오를 750개의 환자별 사례로 구현했습니다. 시스템은 기록 검색, 계산 및 정책 접근을 위한 관리되고 로깅된 환경에서 각 사례를 분석하고, 'Yes', 'No', 'Indeterminate: Lack of Data', 또는 'Indeterminate: Medically Ambiguous'의 네 가지 판정 중 하나를 제시합니다. 여기서 후자의 두 가지는 결론을 내릴 수 있는 충분한 증거가 없을 때를 나타냅니다. 우리는 단순히 판정 정확도 외에도, 환자-증거 연결성 및 정책 기반 판단, 프로세스 준수 여부, 신뢰성 있는 판단 유보(abstention), 그리고 효율성을 평가합니다. 이러한 평가는 독립적인 다중 모델의 판정과 임상 검토 위원회의 검토를 통해 결정된 사례별 기준 판정을 기준으로 이루어집니다. 모든 검색, 계산 및 보고는 재현 가능하도록 설계되어 있어 분석 과정을 검토하고 점수를 매길 수 있습니다. CliniCARE-Bench는 우리가 알고 있는 한, 실제 장기적인 EHR 데이터 분석, 증거 기반 판단, 정책 활용, 프로세스 준수 및 신뢰성 있는 판단 유보를 종합적으로 평가하는 최초의 임상 에이전트 벤치마크입니다. 16개의 시스템을 대상으로 테스트한 결과, 4가지 판정의 정확도는 65.3%에서 76.1% 사이로 나타났지만, 단순히 정확도만으로는 분석 품질을 제대로 반영하지 못합니다. 금지된 방법(shortcuts)을 사용하지 않고 정확하게 판단했을 때의 정확도는 4.8점에서 14.8점 낮으며, 이는 시스템 순위를 재조정하게 합니다.
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably. We introduce CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR), a benchmark for retrospective clinical audit: 25 clinician-validated scenarios instantiated as 750 patient-specific cases over real-patient-derived MIMIC-IV data. Systems investigate each case through a governed, logged tool environment for record retrieval, computation, and policy access, and return one of four verdicts---Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous---the last two separating missing evidence from residual medical ambiguity. Beyond verdict accuracy, we score patient-evidence and policy grounding, process adherence, calibrated abstention, reliability, and efficiency against case-level reference verdicts produced by independent multi-model adjudication and calibrated against Clinical Board review. Every retrieval, computation, and report is replayable, so the investigation trace is inspectable and scorable. To our knowledge, CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework. Across 16 agentic systems, four-way accuracy spans 65.3-76.1%, but raw accuracy overstates investigation quality. Defect-free accuracy, which credits a verdict only when correct and free of prohibited shortcuts, is 4.8-14.8 points lower and reorders the leaderboard.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.