2606.31179v1 Jun 30, 2026 cs.AI

HealthAgentBench: 현실적인 의료 환경 시뮬레이션을 기반으로 한 통합 벤치마크 스위트 - 최첨단 AI 에이전트의 성능을 평가하기 위한 도구

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

N. Usuyama
N. Usuyama
Citations: 7,850
h-index: 22
Tristan Naumann
Tristan Naumann
Citations: 1,579
h-index: 13
H. Poon
H. Poon
Citations: 435
h-index: 8
Wen-wai Yim
Wen-wai Yim
Citations: 63
h-index: 4
Maximilian R. Rokuss
Maximilian R. Rokuss
Citations: 443
h-index: 10
Qianchu Liu
Qianchu Liu
Citations: 145
h-index: 6
Sheng Zhang
Sheng Zhang
Citations: 64
h-index: 4
Guanghui Qin
Guanghui Qin
Citations: 12
h-index: 2
Jeya Maria Jose Valanarasu
Jeya Maria Jose Valanarasu
Citations: 1,093
h-index: 11
Ming Lu
Ming Lu
Citations: 0
h-index: 0
Timothy Ossowski
Timothy Ossowski
Citations: 78
h-index: 5
Juan Manuel Zambrano Chaves
Juan Manuel Zambrano Chaves
Citations: 82
h-index: 2
Cliff Wong
Cliff Wong
Citations: 4,177
h-index: 17
Peniel N. Argaw
Peniel N. Argaw
Citations: 103
h-index: 3
Yashna Hasija
Yashna Hasija
Citations: 0
h-index: 0
Mu Wei
Mu Wei
Citations: 9
h-index: 1
Qin Liu
Qin Liu
Citations: 15
h-index: 3
Zilin Jing
Zilin Jing
Citations: 9
h-index: 2
Jason Entenmann
Jason Entenmann
Citations: 12
h-index: 2

인공지능 에이전트가 복잡하고 장기적인 추론 능력을 갖추게 되면서, 실제 의료 분야 적용 가능성을 측정하기 위해서는 엄격하고 포괄적인 평가가 필수적입니다. 본 논문에서는 HealthAgentBench를 소개합니다. HealthAgentBench는 7가지 범주로 구성된 54개의 다양한 의료 에이전트 관련 작업들을 포함하는 벤치마크 스위트로, 각 작업은 고유한 환경을 가지고 있습니다. 이 벤치마크는 환자 여정 전반에 걸친 다양한 워크플로우와 광범위한 데이터 유형을 포괄합니다. 각 작업은 전체적인 임상 워크플로우를 모방하도록 설계되었으며, 에이전트는 최소한의 지침만 주어지고, 원시 의료 데이터를 탐색하고 복잡한 환경에서 작동하며, 단순한 프롬프트 이상의 다단계 솔루션을 실행해야 합니다. HealthAgentBench의 전반적인 성능을 나타내는 단일 해석 가능한 지표로, 각 에이전트의 최종 작업 성공률을 보고합니다. HealthAgentBench를 사용하여 최첨단 에이전트를 평가한 결과, 전체 작업 성공률은 여전히 낮은 수준이며, 이는 이 벤치마크의 어려움을 보여줍니다. 가장 강력하고 비용 효율적인 에이전트인 Codex GPT-5.5는 약 42%의 성공률을 기록했습니다. 집계된 성능 외에도 HealthAgentBench는 작업 범주별로 미묘한 강점과 약점을 드러냅니다. 최첨단 에이전트는 EHR 데이터를 기반으로 연구 모델링 파이프라인을 자동으로 개발하는 데 유망한 결과를 보이지만, 의료 영상 분야는 특히 어려운 것으로 나타났습니다. Claude Code 모델의 경우 이러한 어려움이 두드러지며, Codex GPT-5.5는 일부 개선된 능력을 보여줍니다. 대규모 검색 공간과 조립적 추론 요구 사항을 결합하는 작업은 현재 모든 에이전트에게 여전히 어렵습니다. 종합적으로 볼 때, HealthAgentBench는 도전적이고 현실적인 벤치마크를 제공하며, 향후 발전의 잠재력이 매우 높습니다. 본 벤치마크는 https://github.com/microsoft/HealthAgentBench 에서 공개됩니다.

Original Abstract

As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories each with its unique environment. The benchmark suite spans diverse workflows throughout the patient journey and a broad range of modalities. Each task is designed to replicate an end-to-end clinical workflow: given minimal instructions, an agent must explore raw healthcare data, operate within a complex environment, and execute multi-step solutions that go beyond naive prompting. A final task success rate is reported to provide a single, interpretable metric for HealthAgentBench overall performance for each agent. Evaluating frontier agents on HealthAgentBench, we find that overall task success rate remains low, underscoring the difficulty of the suite. The strongest and the most cost effective agent, Codex GPT-5.5, achieves only approximately 42% success rate. Beyond aggregate performance, HealthAgentBench reveals nuanced strengths and weaknesses across task categories. Frontier agents show promise in automatically developing research modeling pipelines over EHR data, but medical imaging remains especially challenging, particularly for Claude Code models, while Codex GPT-5.5 shows emerging capability. Tasks that combine large search spaces with compositional reasoning requirements remain difficult for all current agents. Together, these results suggest that HealthAgentBench provides a challenging and realistic benchmark with substantial room for future progress. We release our benchmark at https://github.com/microsoft/HealthAgentBench.

5 Citations
0 Influential
34.4657359028 Altmetric
17.9 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!