2607.25485v1 Jul 28, 2026 cs.AI

PatientAgentBench: 환자 대상 의료 AI 에이전트 평가를 위한 벤치마크 프레임워크

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

Ashutosh Joshi
Ashutosh Joshi
Citations: 27
h-index: 2
Korosh Vatanparvar
Korosh Vatanparvar
Citations: 1,392
h-index: 19
Maria Xenochristou
Maria Xenochristou
Citations: 17
h-index: 2
M. Hashemi
M. Hashemi
Citations: 0
h-index: 0
Prasad Kasu
Prasad Kasu
Citations: 0
h-index: 0
D. Bansal
D. Bansal
Citations: 7
h-index: 1
Daniel Lopez-Martinez
Daniel Lopez-Martinez
Citations: 279
h-index: 3
Anchal Nema
Anchal Nema
Citations: 13
h-index: 1
Ramya Ganesan
Ramya Ganesan
Citations: 0
h-index: 0
Will Kimbrough
Will Kimbrough
Citations: 0
h-index: 0
Alex Woody
Alex Woody
Citations: 507
h-index: 12
Yadunandana Rao
Yadunandana Rao
Citations: 103
h-index: 3
Dilek Hakkani-Tur
Dilek Hakkani-Tur
Citations: 17
h-index: 3
Wilko Schulz-Mahlendorf
Wilko Schulz-Mahlendorf
Citations: 0
h-index: 0

의료 AI는 질문에 답변하는 수준을 넘어, 환자와 대화하고 건강 기록을 분석하며 사용자를 대신하여 행동하는 에이전트 시스템으로 발전하고 있습니다. 일차 의료는 진단 오류 및 안전하지 않은 치료를 예방하는데, 이 분야에서 활동하는 에이전트는 동일한 위험에 대한 평가가 필요합니다. 현재 벤치마크는 주로 의학 지식을 평가하며, 이는 독립적인 질문-응답 또는 임상의 대상 작업으로 이루어집니다. PatientAgentBench는 환자를 대상으로 하는 에이전트 기반 의료 시스템을 평가하는 벤치마크입니다. 이 프레임워크는 파운데이션 모델을 기반으로 구축되며, 다양한 의료 도구를 활용하여 시뮬레이션된 환자와 대화하는 에이전트를 포함합니다. 각 대화는 LLM(Large Language Model)을 기반으로 하는 심사 위원이 100개 이상의 대화-독립적이고 임상 전문가의 의견을 반영한 기준에 따라 6가지 측면에서 평가합니다. 모델의 적합성을 검증하기 위해, 자격을 갖춘 의료진이 동일한 대화를 평가하였으며, 심사위원과 전문가 간의 일치도가 79~93%로 나타났습니다. 이는 임상의 간의 평가 일치도와 유사하거나 그 이상입니다. 저희는 4가지 모델 패밀리에 속하는 10개의 모델을 동일한 1,200개 시나리오에서 벤치마킹하였으며, 상당한 수준의 임상적 격차를 발견했습니다. 삼각 분류(triage) 품질이 가장 큰 차이를 보이는 요소이며, 성능이 낮은 모델은 32%의 성공률을 보이는 반면, 성능이 우수한 모델은 88%의 성공률을 보입니다. 이 때, 일부 에이전트는 임상적 검토 없이 단순히 관리적인 요청에 응답하는 경향이 있습니다. 임상 안전 및 워크플로우 정확성 또한 유사한 패턴을 보이며, 성능이 낮은 모델은 종종 존재하지 않는 작업을 수행하거나 잘못된 정보를 제공하는 반면, 최첨단 모델은 1~3%의 경우에만 오류를 발생시킵니다. 이는 주로 검증되지 않은 도구 출력 또는 응급 상황 시 누락된 정보 때문입니다. 더 발전된 모델은 이러한 격차를 줄이지만 완전히 해소하지 못하며, 전체적으로 5점 만점에 4.25점을 받습니다. 이러한 실패는 실제 환자 기록을 바탕으로 하는 지속적인 대화 및 도구 사용 과정에서만 나타나며, 이는 정적인 벤치마크가 의료 에이전트 시스템의 자율성이 증가함에 따라 충분하지 않음을 시사합니다. 저희는 이 프레임워크를 재현 가능한 형태로 공개하여 임상 전문가의 검증을 거친 평가 표준으로 활용하고, 이 분야의 발전에 기여하고자 합니다.

Original Abstract

Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation model, wrapped in an agent with a sandbox of healthcare tools, conversing with a simulated patient. Each conversation is scored by an LLM-as-a-Jury across six dimensions via over a hundred conversation-agnostic, clinician-grounded criteria. To validate alignment, licensed clinicians annotated shared conversations, yielding 79-93% adjacent agreement between jury and expert raters, on par with or exceeding clinician inter-rater agreement. We benchmarked 10 models across four families on the same 1,200 scenarios and found clinical gaps. Triage quality is the most discriminating dimension: pass rates rise from 32% for the weakest models to 88% for the strongest, with agents often acting on administrative requests without clinical screening. Clinical safety and workflow accuracy follow the same pattern: the weakest models fail often, fabricating unexecuted actions, while frontier models fail on only 1-3% of cases, from unverified tool outputs and omitted crisis resources in an emergency. More capable models narrow these gaps but do not close them; the strongest scores only 4.25 of 5 overall. These failures surface only in sustained, tool-using conversations against realistic patient records, confirming that static benchmarks are insufficient as healthcare agentic systems gain autonomy. We release the framework as a reproducible, clinician-validated evaluation standard to help the field close this gap.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!