MedGuideX: 실행 가능한 임상 지침에서 추출한 의사 결정 로직을 대규모 언어 모델에 통합하여 임상 추론 성능 향상
MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning
임상 실천 지침(CPG)은 의료 전문가가 환자 변수, 조건부 기준 및 권장 규칙을 평가하여 적용하는 근거 기반의 의사 결정 논리를 포함합니다. 그러나 기존 방법들은 종종 CPG를 자유 텍스트 데이터 또는 검색 자료로 활용하며, 그 안에 내재된 절차적 의사 결정 구조를 충분히 활용하지 못합니다. 이러한 구조를 더욱 효과적으로 활용하기 위해, 우리는 CPG 권장 사항을 실행 가능한 임상 의사 결정 논리로 변환하고, 이를 이용하여 사실 및 반사실 질문-응답 데이터를 생성하는 지침 기반 학습 파이프라인을 소개합니다. 이 데이터는 모델에게 지침에 따른 의사 결정뿐만 아니라 다양한 환자 상태에서 의사 결정이 어떻게 변화하는지를 가르쳐줍니다. 이렇게 생성된 데이터를 사용하여 의료 LLM을 추가 훈련한 결과, MedGuideX가 탄생했습니다. 네 가지 임상 추론 벤치마크에서 MedGuideX는 평균 정확도 측면에서 10.28%의 상대적인 성능 향상을 달성했습니다. 또한, 의료 전문가 평가 결과, MedGuideX는 의료진이 작성한 의사 결정 과정을 더 잘 복원하고, 신뢰성, 타당성, 완전성 및 명확성을 기준으로 의료진이 선호하는 논리적 근거를 제시하는 것으로 나타났습니다. 전반적으로, 본 연구의 결과는 CPG에서 추출된 실행 가능한 의사 결정 로직을 확장 가능한 지도 학습으로 변환하여 안정적인 의료 LLM을 구축할 수 있음을 보여줍니다.
Clinical practice guidelines (CPGs) encode evidence-based decision logic that clinicians apply by evaluating patient variables, conditional criteria, and recommendation rules. However, existing methods often use CPGs as free-text training data or retrieval sources, underutilizing their procedural decision structure. To better exploit this structure, we introduce a guideline-derived training pipeline that transforms CPG recommendations into executable clinical decision logic and uses it to generate factual and counterfactual question-answering data. Theses data teach models both guideline-supported decisions and how decisions change under different patient conditions. Post-training a medical LLM on the generated data yields MedGuideX. Across four clinical reasoning benchmarks, MedGuideX achieves a 10.28% relative improvement in average accuracy. Physician evaluation further shows that MedGuideX better recovers clinician authored reasoning steps and produces physician-preferred rationales in faithfulness, validity, completeness, and clarity. Overall, our results show that executable decision logic from CPGs can be transformed into scalable supervision for building reliable medical LLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.