HealthAgentBench: 현실적인 의료 환경 시뮬레이션을 기반으로 한 통합 벤치마크 스위트 - 최첨단 AI 에이전트의 성능을 평가하기 위한 도구
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
인공지능 에이전트가 복잡하고 장기적인 추론 능력을 갖추게 되면서, 실제 의료 분야 적용 가능성을 측정하기 위해서는 엄격하고 포괄적인 평가가 필수적입니다. 본 논문에서는 HealthAgentBench를 소개합니다. HealthAgentBench는 7가지 범주로 구성된 54개의 다양한 의료 에이전트 관련 작업들을 포함하는 벤치마크 스위트로, 각 작업은 고유한 환경을 가지고 있습니다. 이 벤치마크는 환자 여정 전반에 걸친 다양한 워크플로우와 광범위한 데이터 유형을 포괄합니다. 각 작업은 전체적인 임상 워크플로우를 모방하도록 설계되었으며, 에이전트는 최소한의 지침만 주어지고, 원시 의료 데이터를 탐색하고 복잡한 환경에서 작동하며, 단순한 프롬프트 이상의 다단계 솔루션을 실행해야 합니다. HealthAgentBench의 전반적인 성능을 나타내는 단일 해석 가능한 지표로, 각 에이전트의 최종 작업 성공률을 보고합니다. HealthAgentBench를 사용하여 최첨단 에이전트를 평가한 결과, 전체 작업 성공률은 여전히 낮은 수준이며, 이는 이 벤치마크의 어려움을 보여줍니다. 가장 강력하고 비용 효율적인 에이전트인 Codex GPT-5.5는 약 42%의 성공률을 기록했습니다. 집계된 성능 외에도 HealthAgentBench는 작업 범주별로 미묘한 강점과 약점을 드러냅니다. 최첨단 에이전트는 EHR 데이터를 기반으로 연구 모델링 파이프라인을 자동으로 개발하는 데 유망한 결과를 보이지만, 의료 영상 분야는 특히 어려운 것으로 나타났습니다. Claude Code 모델의 경우 이러한 어려움이 두드러지며, Codex GPT-5.5는 일부 개선된 능력을 보여줍니다. 대규모 검색 공간과 조립적 추론 요구 사항을 결합하는 작업은 현재 모든 에이전트에게 여전히 어렵습니다. 종합적으로 볼 때, HealthAgentBench는 도전적이고 현실적인 벤치마크를 제공하며, 향후 발전의 잠재력이 매우 높습니다. 본 벤치마크는 https://github.com/microsoft/HealthAgentBench 에서 공개됩니다.
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduce HealthAgentBench, a suite of 54 agentic healthcare tasks across 7 categories each with its unique environment. The benchmark suite spans diverse workflows throughout the patient journey and a broad range of modalities. Each task is designed to replicate an end-to-end clinical workflow: given minimal instructions, an agent must explore raw healthcare data, operate within a complex environment, and execute multi-step solutions that go beyond naive prompting. A final task success rate is reported to provide a single, interpretable metric for HealthAgentBench overall performance for each agent. Evaluating frontier agents on HealthAgentBench, we find that overall task success rate remains low, underscoring the difficulty of the suite. The strongest and the most cost effective agent, Codex GPT-5.5, achieves only approximately 42% success rate. Beyond aggregate performance, HealthAgentBench reveals nuanced strengths and weaknesses across task categories. Frontier agents show promise in automatically developing research modeling pipelines over EHR data, but medical imaging remains especially challenging, particularly for Claude Code models, while Codex GPT-5.5 shows emerging capability. Tasks that combine large search spaces with compositional reasoning requirements remain difficult for all current agents. Together, these results suggest that HealthAgentBench provides a challenging and realistic benchmark with substantial room for future progress. We release our benchmark at https://github.com/microsoft/HealthAgentBench.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.