2608.04514v1 Aug 05, 2026 cs.CL

RESPClinBench: 호흡기 전문 분야의 다중 모드 임상 의사 결정 및 장기 질환 관리를 위한 벤치마킹

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

Mouxiao Bian
Mouxiao Bian
Citations: 17
h-index: 3
Lu Lu
Lu Lu
Citations: 108
h-index: 6
Jingru Ding
Jingru Ding
Citations: 0
h-index: 0
Yun Zhong
Yun Zhong
Citations: 25
h-index: 3
Jie Xu
Jie Xu
Citations: 5
h-index: 1
Chao Huang
Chao Huang
Citations: 0
h-index: 0
Yueming Su
Yueming Su
Citations: 0
h-index: 0
Z. Chen
Z. Chen
Citations: 0
h-index: 0
Ruiyao Chen
Ruiyao Chen
Citations: 76
h-index: 5
Hengrui Liang
Hengrui Liang
Citations: 11,878
h-index: 32
Yiluo Lin
Yiluo Lin
Citations: 3
h-index: 1

배경: 호흡기 전문 진료는 다중 정보 해석, 장기 위험 평가, 가이드라인에 따른 중재, 그리고 전체 치료 과정을 포함하며, 이러한 요소들은 기존의 시험 중심 의료 벤치마크로는 제대로 반영하기 어렵습니다. 목표: 실제 임상 데이터를 기반으로 호흡기 임상 의사 결정을 위한 벤치마크인 RESPClinBench를 개발하고, AECOPD-PIM 및 PNBIM 데이터셋을 사용하여 7개의 최신 거대 언어 모델의 성능을 평가했습니다. 방법론: RESPClinBench 사례는 익명화된 호흡기 임상 데이터를 기반으로 작성되었으며, 3명의 호흡기 전문의가 사례, 표준 답안, 그리고 세부적인 임상적 조치 사항을 검토하고 수정했습니다. 또한, 1명의 선임 호흡기 전문의가 교차 검토를 통해 최종 결정을 내렸습니다. AECOPD-PIM은 427개의 개방형 COPD 사례로 구성되었으며, PNBIM은 흉부 CT 영상과 구조화된 임상 정보를 결합한 196개의 다중 모드 폐종양 사례로 구성되었습니다. 7개의 모델이 표준 API를 사용하여 (temperature = 0, 최대 출력 길이 = 8192 토큰) 총 4,361개의 응답을 생성했습니다. 자동화된 프레임워크는 세부 조치 사항의 재현율과 rubrics 기반 LLM-as-a-Judge 평가 결과를 산술 평균하여 최종 점수를 계산했습니다. 결과: 623건의 사례에서 평균 최종 점수는 68.58점이었습니다. Qwen3.6-27B 모델이 전체적으로 가장 높은 순위를 차지했으며 (71.22점), Qwen3.5-397B-A17B 모델이 PNBIM 데이터셋에서 가장 높은 성능을 보였으며 (72.48점), Qwen3.6-27B 모델이 AECOPD-PIM 데이터셋에서 가장 높은 성능을 보였습니다 (71.11점). PNBIM 응답의 31.85%에서 영상 왜곡 현상이 발생했고, 8.16%에서 심각한 의료적 위험이 관찰되었습니다. AECOPD-PIM 응답의 26.93%에서 약물 안전 관련 문제가 발생했고, 1.44%에서 심각한 의료적 위험이 관찰되었습니다. 결론: RESPClinBench는 다중 모드 폐종양 평가 및 장기 COPD 관리와 관련된 특정 과제를 식별합니다. 명확한 임상적 조치 범위, 종합적인 평가, 그리고 독립적인 안전성 지표를 결합함으로써 모델 선택 및 잠재적 검증을 위한 임상적으로 타당한 기반을 제공합니다.

Original Abstract

Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-course management, which are poorly represented by examination-oriented medical benchmarks. Objective: To develop RESPClinBench, a real-world scenario-based benchmark for respiratory clinical decision-making, and evaluate seven contemporary large language models across AECOPD-PIM and PNBIM. Methods: RESPClinBench cases were adapted from de-identified respiratory clinical data. Three attending-level respiratory physicians revised cases, reference answers, and atomic clinical-action points, while one senior respiratory specialist performed cross-review and final adjudication. AECOPD-PIM comprised 427 open-ended COPD cases, and PNBIM comprised 196 multimodal pulmonary nodule cases combining chest CT with structured clinical information. Seven models generated 4,361 responses through standardized API inference with temperature 0 and a maximum output length of 8192 tokens. An automated framework calculated the final score as the arithmetic mean of atomic-action recall and rubric-based LLM-as-a-Judge assessment. Results: Across 623 cases, the mean final score was 68.58. Qwen3.6-27B ranked first overall at 71.22, Qwen3.5-397B-A17B led PNBIM at 72.48, and Qwen3.6-27B led AECOPD-PIM at 71.11. Imaging hallucination and serious medical risk occurred in 31.85% and 8.16% of PNBIM responses; medication-safety risk and serious medical risk occurred in 26.93% and 1.44% of AECOPD-PIM responses. Conclusions: RESPClinBench identifies task-specific limitations in multimodal pulmonary nodule assessment and longitudinal COPD management. Combining explicit clinical-action coverage, holistic evaluation, and independent safety flags provides a clinically grounded basis for model selection and prospective validation.

0 Citations
0 Influential
16 Altmetric
80.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!