MyoCardBench: 임상적으로 현실적인 심혈관 치료 시나리오에서 대규모 언어 모델을 평가하기 위한 실세계 데이터 벤치마크
MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
배경: 대부분의 의료 대규모 언어 모델(LLM) 벤치마크는 시험 지식 또는 개별 작업에 초점을 맞추며, 심혈관 치료의 장기적이고 다중 모달이며 안전이 중요한 워크플로우를 제대로 반영하지 못할 수 있습니다. 목적: 본 연구에서는 심혈관 치료 전반을 포괄하는 실세계 벤치마크인 MyoCardBench를 개발하고, LLM의 성능을 다양한 임상 측면 및 전문 분야 작업에 걸쳐 평가하고자 합니다. 방법: MyoCardBench는 익명화된 심혈관 기록 및 시험 데이터에서 파생된 13개의 작업별 데이터 세트의 2,263개 항목으로 구성되어 있습니다. 16명의 순환기내과 의사가 주석 작업을 수행하고 참조 자료를 구축했으며, 이후 두 명의 선임 순환기내과 의사의 교차 검토가 이루어졌습니다. 7개의 LLM이 표준화된 제로샷 환경에서 15,841개의 출력을 생성했습니다. 개방형 작업은 핵심 내용 포함 여부와 전반적인 임상 품질을 기준으로 평가되었으며, CardioEthics는 정확도를 기준으로 점수 매겨졌습니다. 결과: GPT-5.4가 가장 높은 거시 평균(62.55)과 항목 가중 평균(62.19)을 기록했으며, 그 뒤를 Gemini 3.1 Pro (59.95)와 Qwen 3.6 27B (59.72)이 따랐습니다. GPT-5.4는 세 가지 모든 측면에서 1위를 차지했습니다. CardioAuxReport가 가장 우수한 성능(86.38)을 보인 반면, CardioECGRead (17.25)와 CardioEthics (17.34)가 가장 낮은 점수를 기록했습니다. 전반적인 임상 품질과 핵심 내용 포함 여부 간의 가장 큰 격차는 CardioComm (52.71), CardioEmergRescue (52.05) 및 CardioTreatPlan (48.80)에서 나타났습니다. 결론: MyoCardBench는 현재까지 개발된 심혈관 치료 전반에 걸친 LLM 평가를 위한 가장 큰 실세계, 다중 작업 벤치마크이며, 지금까지 보고된 임상적으로 현실적인 순환기내과 시나리오에 대한 가장 광범위한 적용 범위를 제공합니다. 이는 모델의 강점, 임상적으로 중요한 누락 사항 및 향후 개발 우선순위를 파악하기 위한 엄격한 프레임워크를 제공합니다.
Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To develop MyoCardBench, a real-world benchmark spanning the cardiovascular care continuum, and assess LLM performance across clinical dimensions and specialist tasks. Methods: MyoCardBench includes 2,263 items from 13 task-specific datasets derived from de-identified cardiovascular records and examination data. Sixteen cardiology physicians conducted annotation and reference construction, followed by cross-review from two senior cardiologists. Seven LLMs generated 15,841 outputs under standardized zero-shot settings. Open-ended tasks were evaluated using key-point coverage and holistic clinical quality, while CardioEthics was scored by accuracy. Results: GPT-5.4 achieved the highest macro-average (62.55) and item-weighted mean (62.19), followed by Gemini 3.1 Pro (59.95) and Qwen 3.6 27B (59.72). GPT-5.4 ranked first in all three dimensions. CardioAuxReport performed best (86.38), whereas CardioECGRead (17.25) and CardioEthics (17.34) were lowest. The largest gaps between holistic clinical quality and key-point coverage occurred in CardioComm (52.71), CardioEmergRescue (52.05), and CardioTreatPlan (48.80). Conclusions: To our knowledge, MyoCardBench is the largest real-world, multi-task benchmark for LLM evaluation across the cardiovascular care continuum and offers the broadest coverage of clinically authentic cardiology scenarios reported to date. It provides a rigorous framework for identifying model strengths, clinically important omissions, and priorities for future development.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.