2601.23049v1 Jan 30, 2026 cs.AI

MedMCP-Calc: MCP 통합을 통한 실제 의료 계산기 시나리오에서의 LLM 벤치마킹

MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration

Yutong Huang
Yutong Huang
Citations: 13
h-index: 1
Shengqian Qin
Shengqian Qin
Citations: 14
h-index: 2
Zhongzhen Huang
Zhongzhen Huang
Citations: 260
h-index: 7
Shaoting Zhang
Shaoting Zhang
Citations: 1,256
h-index: 14
Xiaofan Zhang
Xiaofan Zhang
Citations: 109
h-index: 7
Yakun Zhu
Yakun Zhu
Citations: 48
h-index: 4

의료 계산기는 정량적이고 근거 기반의 임상 진료에 필수적입니다. 그러나 기존 벤치마크가 명시적인 지시가 있는 정적인 단일 단계 계산에만 초점을 맞추는 반면, 실제 환경에서의 사용은 능동적인 EHR 데이터 획득, 시나리오에 따른 계산기 선택, 다단계 연산이 요구되는 적응형 다단계 과정입니다. 이러한 한계를 해결하기 위해, 본 논문에서는 모델 컨텍스트 프로토콜(MCP) 통합을 통해 실제 의료 계산기 시나리오에서 LLM을 평가하는 최초의 벤치마크인 MedMCP-Calc를 소개합니다. MedMCP-Calc는 4개 임상 분야에 걸친 118개 시나리오 작업으로 구성되어 있으며, 자연어 질의를 모방한 모호한 작업 설명, 구조화된 EHR 데이터베이스 상호작용, 외부 참조 검색, 그리고 프로세스 수준 평가를 특징으로 합니다. 23개의 선도적인 모델을 평가한 결과 중대한 한계가 드러났습니다. Claude Opus 4.5와 같은 최상위 모델조차 모호한 질의가 주어지는 엔드투엔드 워크플로우에서 적절한 계산기를 선택하는 데 어려움을 겪고, 반복적인 SQL 기반 데이터베이스 상호작용에서 성능이 저조하며, 수치 연산을 위해 외부 도구를 활용하는 데 있어 뚜렷한 소극성을 보이는 등 상당한 결점을 보였습니다. 또한 임상 분야에 따라 성능 편차가 상당히 컸습니다. 이러한 결과를 바탕으로, 저자들은 시나리오 계획 및 도구 증강을 통합한 파인 튜닝 모델인 CalcMate를 개발하여 오픈 소스 모델 중 최고 수준(SOTA)의 성능을 달성했습니다. 벤치마크와 코드는 https://github.com/SPIRAL-MED/MedMCP-Calc 에서 확인할 수 있습니다.

Original Abstract

Medical calculators are fundamental to quantitative, evidence-based clinical practice. However, their real-world use is an adaptive, multi-stage process, requiring proactive EHR data acquisition, scenario-dependent calculator selection, and multi-step computation, whereas current benchmarks focus only on static single-step calculations with explicit instructions. To address these limitations, we introduce MedMCP-Calc, the first benchmark for evaluating LLMs in realistic medical calculator scenarios through Model Context Protocol (MCP) integration. MedMCP-Calc comprises 118 scenario tasks across 4 clinical domains, featuring fuzzy task descriptions mimicking natural queries, structured EHR database interaction, external reference retrieval, and process-level evaluation. Our evaluation of 23 leading models reveals critical limitations: even top performers like Claude Opus 4.5 exhibit substantial gaps, including difficulty selecting appropriate calculators for end-to-end workflows given fuzzy queries, poor performance in iterative SQL-based database interactions, and marked reluctance to leverage external tools for numerical computation. Performance also varies considerably across clinical domains. Building on these findings, we develop CalcMate, a fine-tuned model incorporating scenario planning and tool augmentation, achieving state-of-the-art performance among open-source models. Benchmark and Codes are available in https://github.com/SPIRAL-MED/MedMCP-Calc.

3 Citations
0 Influential
42.222612188617 Altmetric
13.9 Score
Original PDF
20

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!