2606.23301v1 Jun 22, 2026 cs.AI

EHR-Complex: 복잡한 임상 추론을 위한 의료 에이전트 성능 평가

EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning

Jinjie Gu
Jinjie Gu
Citations: 476
h-index: 12
Jian Wang
Jian Wang
Citations: 46
h-index: 3
Zhixuan Chu
Zhixuan Chu
Citations: 113
h-index: 5
Yue Shen
Yue Shen
Citations: 514
h-index: 10
Yitong Qiao
Yitong Qiao
Citations: 25
h-index: 3
Kui Ren
Kui Ren
Citations: 81
h-index: 5
Lei Liu
Lei Liu
Citations: 5
h-index: 1

임상 에이전트는 전자 건강 기록(EHR)에 대한 접근성을 높일 잠재력을 가지고 있지만, 기존의 성능 평가 기준은 실제 EHR 분석의 복잡성을 제대로 반영하지 못합니다. 예를 들어, 많은 평가가 이상적인 상태의 깨끗한 EHR 데이터를 사용하며, 정적인 SQL 생성 방식을 통해 분석을 수행하는 경우가 많습니다. 본 연구에서는 대규모 임상 데이터베이스 추론을 위한 상호 작용형 벤치마크인 EHR-Complex를 소개합니다. EHR-Complex는 365,000명의 환자 데이터를 포함하고 있으며, 31개의 테이블과 5억 개 이상의 레코드를 가진 MIMIC-IV 데이터셋을 기반으로 구축되었습니다. EHR-Complex는 약 52,000개의 작업으로 구성되어 있으며, 이는 여섯 가지 임상적 목표를 지원하며, 개별 환자 수준 및 전체 인구 수준의 질의를 모두 포함합니다. 각 작업에서 에이전트는 SQL 쿼리 또는 Python 코드를 실행하여 격리된 환경과 상호 작용해야 합니다. 특히 EHR-Complex는 장기적인 다중 테이블 집계 및 합성 추론을 위한 실제 SQL 작업의 복잡성을 고려하며, 평균적으로 쿼리당 약 31.93개의 SQL 구조적 요소가 사용됩니다. EHR-Complex에 대한 평가 결과는 이러한 EHR 추론 시나리오의 임상적 어려움을 보여줍니다. 최고 성능 모델조차도 정확 일치율이 62.3%에 불과했습니다. Pass^k 일관성은 거의 모든 평가된 모델에서 k=4일 때 50% 미만으로 떨어졌으며, 이는 광범위한 확률적 취약성을 드러냅니다. 대표적인 LLM의 3,800건 이상의 실패 사례에 대한 세부 분석 결과, SQL 논리 오류, 의료 코드 조회 실패 및 의미 이해 부족이 주요 원인으로 나타났습니다. EHR-Complex는 임상 에이전트를 위한 엄격한 테스트 환경을 제공하며, 대규모 EHR 분석을 위한 견고한 추론 능력의 개선 필요성을 강조합니다.

Original Abstract

Clinical agents promise to democratize access to electronic health records (EHRs), yet existing benchmarks fail to reflect the complexity of practical EHR analysis, e.g., often operating on idealized, clean EHRs via static SQL generation rather than interactive execution. In this work, we introduce EHR-Complex, a large-scale benchmark designed for interactive clinical database reasoning. Built on the large MIMIC-IV substrate (365K patients, 31 tables, 500M+ records), EHR-Complex comprises about 52K tasks spanning six clinical intents, supporting both patient-level and population-level queries, where each task requires an agent to interact with a sandboxed environment by executing SQL queries or Python code. Notably, EHR-Complex considers the real-world SQL task complexity for longitudinal multi-table aggregation and compositional reasoning, resulting in 31.93 SQL structural components per query on average. Evaluation results on EHR-Complex reveal the clinical difficulty of these EHR reasoning scenarios, with the top-performing model achieving only 62.3% exact-match accuracy. Pass^k consistency drops below 50% for nearly all evaluated models at k=4, exposing broad stochastic fragility. A fine-grained analysis of more than 3,800 failed trajectories for representative LLMs reveals three dominant failure modes: SQL logic errors, medical-code lookup failures, and semantic misunderstandings. EHR-Complex provides a rigorous testbed for clinical agents and highlights remaining gaps in robust reasoning for large-scale EHR analysis.

2 Citations
0 Influential
6 Altmetric
32.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!