2606.18613v1 Jun 17, 2026 cs.CL

LLM은 의사를 지원할 준비가 되었는가? 대화형 의사-환자-전자의료기록 지원을 위한 PhysAssistBench

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

Peijie Yu
Peijie Yu
Citations: 34
h-index: 3
Shaoxiong Ji
Shaoxiong Ji
Citations: 59
h-index: 4
D. Rueckert
D. Rueckert
Citations: 11
h-index: 1
Jiazhen Pan
Jiazhen Pan
Citations: 570
h-index: 13
T. Du
T. Du
Citations: 191
h-index: 3
Sihan Shang
Sihan Shang
Citations: 11
h-index: 2
Danli Shi
Danli Shi
Citations: 97
h-index: 4
My L. T. Nguyen
My L. T. Nguyen
Citations: 14
h-index: 2
Shengbo Gao
Shengbo Gao
Citations: 0
h-index: 0
Guangyuan Li
Guangyuan Li
Citations: 7
h-index: 1
Yi Yu
Yi Yu
INRIA Paris
Citations: 0
h-index: 0
Yan Jiang
Yan Jiang
Citations: 0
h-index: 0
Qianlong Zhao
Qianlong Zhao
Citations: 0
h-index: 0
Behzad Bozorgtabar
Behzad Bozorgtabar
Citations: 2,611
h-index: 27
Jiancheng Yang
Jiancheng Yang
Citations: 36
h-index: 2

의료 LLM의 가장 현실적인 단기적 역할은 의사를 대체하는 것이 아니라 돕는 것입니다. 그러나 현재 평가에서는 종종 개별적인 능력, 즉 임상 지식, 전자의료 기록 시스템과의 상호 작용 또는 환자와의 소통을 테스트합니다. 반면, 실제 의료 환경에서의 지원은 이러한 능력을 하나의 상호 작용 내에서 조정해야 합니다. 의사는 명확하게 정의되지 않은 요청을 하고, 환자는 모호하게 증상을 설명하며, 전자의료 기록 시스템은 정확한 도구 사용을 요구하기 때문입니다. 본 논문에서는 대화형 의사-환자-전자의료 기록 지원을 위한 벤치마크인 PhysAssistBench를 소개합니다. 실제 MIMIC-IV 데이터를 기반으로 구축된 PhysAssistBench는 확장 가능한 파이프라인을 사용하여 능동적인 환자를 생성합니다. 이 환자는 상호 작용하며, 기록에 기반하여 정적 전자의료 기록을 다중 턴의 임상 시나리오로 변환하면서 임상적 사실성을 유지합니다. PhysAssistBench는 수동으로 검토 및 의사가 검증한 1,296개의 대화 턴으로 구성된 양방향 평가 데이터 세트를 제공합니다. 선도적인 LLM을 사용한 실험 결과, 현재 모델은 이러한 환경에서 여전히 신뢰성이 낮다는 것을 보여줍니다. 이는 임상용 LLM의 핵심적인 문제점을 드러냅니다. 즉, 신뢰할 수 있는 지원을 위해서는 지식, 소통 및 시스템 간의 조화가 필요하며, 개별 능력에서의 단순한 발전으로는 충분하지 않습니다.

Original Abstract

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PhysAssistBench, a benchmark for interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, PhysAssistBench uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality. PhysAssistBench provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.

0 Citations
0 Influential
13.5 Altmetric
67.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!