2608.07796v1 Aug 07, 2026 cs.AI

CliniCARE-Bench: EHR 데이터 기반 의료 추론의 임상적 정확성 평가

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

Jingxuan Fan
Jingxuan Fan
Citations: 6
h-index: 2
George Pu
George Pu
Citations: 388
h-index: 3
Y. Xue
Y. Xue
Citations: 98
h-index: 2
Soham Dan
Soham Dan
Citations: 185
h-index: 8
Daniel Zhang
Daniel Zhang
Citations: 290
h-index: 4
Varun Ursekar
Varun Ursekar
Citations: 18
h-index: 3
Apaar Shanker
Apaar Shanker
Citations: 74
h-index: 5
Vijay S. Kalmath
Vijay S. Kalmath
Citations: 0
h-index: 0
Veronica Chatrath
Veronica Chatrath
Citations: 85
h-index: 5
Bryan Zhu
Bryan Zhu
Citations: 0
h-index: 0
Anahita Sharma
Anahita Sharma
Citations: 0
h-index: 0
Jason Qin
Jason Qin
Citations: 13
h-index: 1
Keqi Han
Keqi Han
Citations: 10
h-index: 2
Soham Dinesh Tiwari
Soham Dinesh Tiwari
Carnegie Mellon University
Citations: 39
h-index: 4
Yuan Li
Yuan Li
Citations: 0
h-index: 0
Chenguang Wang
Chenguang Wang
Citations: 0
h-index: 0
Zainab Doctor
Zainab Doctor
Citations: 0
h-index: 0
Zhijun Yin
Zhijun Yin
Citations: 31
h-index: 3
Nigam H.Shah
Nigam H.Shah
Citations: 0
h-index: 0

대규모 언어 모델은 의료 지식 벤치마크에서 뛰어난 성능을 보이지만, 실제 임상 환경에 적용하기 위해서는 시스템이 다양한 형태의 과거 기록들을 바탕으로 신뢰할 수 있는 분석을 수행해야 합니다. 이는 필요한 증거를 결정하고, 구조화된 데이터와 자유 형식 데이터를 검색 및 통합하며, 결론을 검증 가능한 근거로 뒷받침하고, 명확하게 해결되지 않는 경우에는 판단을 유보하는 것을 포함합니다. 본 논문에서는 CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR)라는 임상 감사 벤치마크를 소개합니다. 이는 실제 환자 데이터를 기반으로 한 MIMIC-IV 데이터셋을 활용하여, 의료 전문가가 검증한 25개의 시나리오를 750개의 환자별 사례로 구현했습니다. 시스템은 기록 검색, 계산 및 정책 접근을 위한 관리되고 로깅된 환경에서 각 사례를 분석하고, 'Yes', 'No', 'Indeterminate: Lack of Data', 또는 'Indeterminate: Medically Ambiguous'의 네 가지 판정 중 하나를 제시합니다. 여기서 후자의 두 가지는 결론을 내릴 수 있는 충분한 증거가 없을 때를 나타냅니다. 우리는 단순히 판정 정확도 외에도, 환자-증거 연결성 및 정책 기반 판단, 프로세스 준수 여부, 신뢰성 있는 판단 유보(abstention), 그리고 효율성을 평가합니다. 이러한 평가는 독립적인 다중 모델의 판정과 임상 검토 위원회의 검토를 통해 결정된 사례별 기준 판정을 기준으로 이루어집니다. 모든 검색, 계산 및 보고는 재현 가능하도록 설계되어 있어 분석 과정을 검토하고 점수를 매길 수 있습니다. CliniCARE-Bench는 우리가 알고 있는 한, 실제 장기적인 EHR 데이터 분석, 증거 기반 판단, 정책 활용, 프로세스 준수 및 신뢰성 있는 판단 유보를 종합적으로 평가하는 최초의 임상 에이전트 벤치마크입니다. 16개의 시스템을 대상으로 테스트한 결과, 4가지 판정의 정확도는 65.3%에서 76.1% 사이로 나타났지만, 단순히 정확도만으로는 분석 품질을 제대로 반영하지 못합니다. 금지된 방법(shortcuts)을 사용하지 않고 정확하게 판단했을 때의 정확도는 4.8점에서 14.8점 낮으며, 이는 시스템 순위를 재조정하게 합니다.

Original Abstract

Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably. We introduce CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR), a benchmark for retrospective clinical audit: 25 clinician-validated scenarios instantiated as 750 patient-specific cases over real-patient-derived MIMIC-IV data. Systems investigate each case through a governed, logged tool environment for record retrieval, computation, and policy access, and return one of four verdicts---Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous---the last two separating missing evidence from residual medical ambiguity. Beyond verdict accuracy, we score patient-evidence and policy grounding, process adherence, calibrated abstention, reliability, and efficiency against case-level reference verdicts produced by independent multi-model adjudication and calibrated against Clinical Board review. Every retrieval, computation, and report is replayable, so the investigation trace is inspectable and scorable. To our knowledge, CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework. Across 16 agentic systems, four-way accuracy spans 65.3-76.1%, but raw accuracy overstates investigation quality. Defect-free accuracy, which credits a verdict only when correct and free of prohibited shortcuts, is 4.8-14.8 points lower and reorders the leaderboard.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!