2606.11702v1 Jun 10, 2026 cs.CV

MedCTA: 임상 도구 에이전트를 위한 벤치마크

MedCTA: A Benchmark for Clinical Tool Agents

Fida Mohammad Thoker
Fida Mohammad Thoker
Citations: 357
h-index: 6
Tajamul Ashraf
Tajamul Ashraf
Citations: 140
h-index: 4
H. Jeong
H. Jeong
Citations: 440
h-index: 5
Bernard Ghanem
Bernard Ghanem
Citations: 1,894
h-index: 8

임상적으로 타당한 의사 결정을 내리기 위해, 의료 AI 에이전트는 단순한 인식 능력을 넘어 도구 검색, 증거 수집 및 통합 기능을 갖춰야 합니다. 기존의 벤치마크는 주로 개별적인 인식 능력이나 단일 질문-응답 방식을 평가하며, 따라서 계획 실패, 도구 활용 실패, 그리고 안정성 문제에 대한 제한적인 정보를 제공합니다. 본 연구에서는 실제 임상 환경에서 사용되는 다양한 형태의 의료 데이터(예: 영상, 병리 슬라이드, 보고서)를 기반으로 한 다단계 작업을 평가하기 위한 벤치마크인 MedCTA를 소개합니다. MedCTA는 실제 임상 작업 107개를 포함하며, 각 작업은 임상의가 검증한 실행 가능한 단계로 구성되어 있으며, 5개의 도구 활용을 지원합니다. MedCTA는 도구 선택의 적절성, 논리적 타당성, 실행 안정성, 경로 충실성 및 결과 품질과 같은 과정을 고려한 평가를 가능하게 합니다. 오픈 소스 및 상용 모델 18개를 벤치마킹한 결과, 최첨단 시스템조차도 다단계 임상 도구 사용에서 불안정성을 보이는 것으로 나타났습니다. 자율적인 실행 과정에서는 프로토콜 오류, 조기 종료, 그리고 잘못된 도구 활용이 빈번하게 발생하며, 이상적인 도구 연결 방식은 상당한 개선 효과를 가져다주지만 여전히 완벽하지 않습니다. 이러한 결과는 강력한 인식 능력이 임상 환경에서 신뢰할 수 있는 에이전트 행동으로 이어지지 않는다는 것을 보여줍니다. MedCTA는 신뢰할 수 있는 의료 AI 에이전트를 평가, 진단 및 발전시키는 데 필요한 엄격한 테스트 환경을 제공합니다. 데이터셋과 평가 도구는 https://ivul-kaust.github.io/MedCTA/ 에서 이용 가능합니다.

Original Abstract

To make clinically grounded decisions, medical AI agents are expected to go beyond simple recognition and be capable of tool retrieval, evidence acquisition, and integration. Existing benchmarks largely evaluate isolated perception or single-turn question answering, and therefore provide limited visibility into failures of planning, tool recruitment, and rollout reliability. We introduce MedCTA, a benchmark for evaluating medical tool agents on clinician-validated, step-implicit tasks grounded in realistic multimodal clinical inputs, including radiology images, pathology slides, and reports. MedCTA comprises 107 real-world clinical tasks with clinician-verified executable trajectories over 5 deployed tools, and supports process-aware evaluation of tool selection, argument validity, execution stability, trajectory fidelity, and outcome quality. We benchmark 18 open- and closed-source multimodal models and find that even frontier systems remain brittle in multi-step clinical tool use: autonomous rollouts are dominated by protocol failures, premature stopping, and incorrect tool recruitment, while gold-standard tool routing yields large but still incomplete gains. These results show that strong backbone perception does not translate into reliable agentic behavior in clinical settings. MedCTA provides a rigorous testbed for auditing, diagnosing, and advancing trustworthy medical AI agents. The dataset and evaluation suite are available at https://ivul-kaust.github.io/MedCTA/

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!