ART: 의료 AI 에이전트를 위한 행동 기반 추론 작업 벤치마킹
ART: Action-based Reasoning Task Benchmarking for Medical AI Agents
신뢰할 수 있는 임상 의사 결정 지원을 위해서는 구조화된 전자의무기록(EHR)에 대해 안전하고 다단계의 추론을 수행할 수 있는 의료 AI 에이전트가 필요하다. 대규모 언어 모델(LLM)이 헬스케어 분야에서 유망한 모습을 보이고 있지만, 기존 벤치마크들은 임계값 평가, 시간적 집계, 조건부 논리를 포함하는 행동 기반 작업에 대한 성능을 제대로 평가하지 못하고 있다. 우리는 의료 AI 에이전트를 위한 행동 기반 임상 추론 작업 벤치마크인 ART를 소개한다. ART는 실제 EHR 데이터를 마이닝하여 알려진 추론 약점을 겨냥한 고난도 작업들을 생성한다. 기존 벤치마크 분석을 통해 우리는 검색 실패, 집계 오류, 조건부 논리 판단 착오라는 세 가지 주요 오류 범주를 식별했다. 시나리오 식별, 작업 생성, 품질 감사, 평가로 구성된 우리의 4단계 파이프라인은 실제 환자 데이터에 기반하여 임상적으로 검증된 다양한 작업들을 생성한다. 600개의 작업에 대해 GPT-4o-mini와 Claude 3.5 Sonnet을 평가한 결과, 프롬프트 개선 후 검색 능력은 거의 완벽했으나 집계(28~64%)와 임계값 추론(32~38%)에서는 상당한 격차가 나타났다. 행동 지향적 EHR 추론에서의 실패 양상을 드러냄으로써, ART는 보다 신뢰할 수 있는 임상 에이전트로의 발전을 도모한다. 이는 수요가 높은 의료 환경에서 인지적 부하와 행정적 부담을 줄여 의료 인력의 역량을 지원하는 AI 시스템을 위한 필수적인 단계이다.
Reliable clinical decision support requires medical AI agents capable of safe, multi-step reasoning over structured electronic health records (EHRs). While large language models (LLMs) show promise in healthcare, existing benchmarks inadequately assess performance on action-based tasks involving threshold evaluation, temporal aggregation, and conditional logic. We introduce ART, an Action-based Reasoning clinical Task benchmark for medical AI agents, which mines real-world EHR data to create challenging tasks targeting known reasoning weaknesses. Through analysis of existing benchmarks, we identify three dominant error categories: retrieval failures, aggregation errors, and conditional logic misjudgments. Our four-stage pipeline -- scenario identification, task generation, quality audit, and evaluation -- produces diverse, clinically validated tasks grounded in real patient data. Evaluating GPT-4o-mini and Claude 3.5 Sonnet on 600 tasks shows near-perfect retrieval after prompt refinement, but substantial gaps in aggregation (28--64%) and threshold reasoning (32--38%). By exposing failure modes in action-oriented EHR reasoning, ART advances toward more reliable clinical agents, an essential step for AI systems that reduce cognitive load and administrative burden, supporting workforce capacity in high-demand care settings
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.