2601.12661v1 Jan 19, 2026 cs.AI

MedConsultBench: 의료 상담 에이전트를 위한 전주기, 세밀한, 과정 인식 벤치마크

MedConsultBench: A Full-Cycle, Fine-Grained, Process-Aware Benchmark for Medical Consultation Agents

Chuhan Qiao
Chuhan Qiao
Citations: 7
h-index: 2
Daxing Zhao
Daxing Zhao
Citations: 0
h-index: 0
Wei Lin
Wei Lin
Citations: 31
h-index: 2
Ziding Liu
Ziding Liu
Citations: 14
h-index: 2
Yan-Bin Shen
Yan-Bin Shen
Citations: 1,139
h-index: 7
Bing Cheng
Bing Cheng
Citations: 15
h-index: 2
Kai Wu
Kai Wu
Citations: 42
h-index: 4
Jianghua Huang
Jianghua Huang
Citations: 257
h-index: 7

현재의 의료 상담 에이전트 평가는 종종 결과 지향적 과제에 우선순위를 두며, 실제 진료에 필수적인 종단간 프로세스 무결성과 임상적 안전성을 간과하는 경우가 많다. 최근 대화형 벤치마크들이 동적 시나리오를 도입했지만, 여전히 파편화되고 대략적인 수준에 머물러 있어 전문적인 상담에 요구되는 구조화된 문진 논리와 진단의 엄격함을 포착하지 못하고 있다. 이러한 격차를 해소하기 위해, 우리는 병력 청취와 진단에서부터 치료 계획 및 후속 질의응답에 이르는 전체 임상 워크플로우를 포괄하여 온라인 상담의 전 과정을 평가하도록 설계된 포괄적 프레임워크인 MedConsultBench를 제안한다. 우리의 방법론은 임상 정보 획득을 서브 턴(sub-turn) 수준에서 추적하기 위해 '원자적 정보 단위(AIUs)'를 도입했으며, 22가지 세부 지표를 통해 핵심 사실들이 어떻게 도출되는지 정밀하게 모니터링할 수 있게 한다. 온라인 상담에 내재된 정보 부족과 모호성을 다룸으로써, 이 벤치마크는 불확실성을 인지하면서도 간결한 문진 능력을 평가한다. 또한 약물 처방의 호환성을 강조하고, 제약 조건을 준수하는 계획 수정을 통해 현실적인 처방 후 후속 질의응답을 처리하는 능력을 중시한다. 19개 거대언어모델(LLM)에 대한 체계적인 평가 결과, 높은 진단 정확도가 정보 수집 효율성과 약물 안전성 측면의 중대한 결함을 가리는 경우가 많음이 밝혀졌다. 이러한 결과는 이론적 의학 지식과 임상 진료 능력 사이의 결정적인 격차를 강조하며, MedConsultBench가 의료 AI를 실제 임상 현장의 세밀한 요구사항에 부합하도록 만드는 데 있어 엄격한 기반이 됨을 입증한다.

Original Abstract

Current evaluations of medical consultation agents often prioritize outcome-oriented tasks, frequently overlooking the end-to-end process integrity and clinical safety essential for real-world practice. While recent interactive benchmarks have introduced dynamic scenarios, they often remain fragmented and coarse-grained, failing to capture the structured inquiry logic and diagnostic rigor required in professional consultations. To bridge this gap, we propose MedConsultBench, a comprehensive framework designed to evaluate the complete online consultation cycle by covering the entire clinical workflow from history taking and diagnosis to treatment planning and follow-up Q\&A. Our methodology introduces Atomic Information Units (AIUs) to track clinical information acquisition at a sub-turn level, enabling precise monitoring of how key facts are elicited through 22 fine-grained metrics. By addressing the underspecification and ambiguity inherent in online consultations, the benchmark evaluates uncertainty-aware yet concise inquiry while emphasizing medication regimen compatibility and the ability to handle realistic post-prescription follow-up Q\&A via constraint-respecting plan revisions. Systematic evaluation of 19 large language models reveals that high diagnostic accuracy often masks significant deficiencies in information-gathering efficiency and medication safety. These results underscore a critical gap between theoretical medical knowledge and clinical practice ability, establishing MedConsultBench as a rigorous foundation for aligning medical AI with the nuanced requirements of real-world clinical care.

2 Citations
0 Influential
3.5 Altmetric
19.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!