GuideSkill: 임상 지침 기반 추론을 위한 실행 가능한 LLM 에이전트 기술의 진화
GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning
임상 실천 지침(CPGs)은 진단 기준을 포함하고 있지만, 일반적으로 LLM 시스템은 이러한 지침 텍스트를 검색하거나 학습을 통해 이를 활용하는 반면, 실제 규칙을 실행하지는 않습니다. 본 연구에서는 질병별 기준을 실행 가능한 함수로 컴파일하여 순위 기반 진단 지원 점수를 반환하는 외부 추론 레이어인 GuideSkill을 소개합니다. GuideSkill-Zero는 지침에서 초기화되며, GuideSkill-Evo는 사례-진단 쌍을 사용하여 포함된 기술을 개선하고 누락된 진단을 추가합니다. 추론 과정에서 LLM은 감별 진단을 제안하고, 일치하는 각 기술에 필요한 특징을 기반으로 이를 활용하며, 실행된 기술 점수와 LLM의 순위를 통합합니다. 네 가지 벤치마크 및 네 가지 모델 구조를 사용하여 GuideSkill-Zero는 평균적으로 지침 검색 증강 생성(RAG) 방식보다 정확도가 13.45% 향상되었습니다. GuideSkill-Evo는 모든 모델 구조에서 가장 높은 평균 정확도를 달성했으며, 직접 추론 방식에 비해 상대적으로 18.49% 향상되었으며, 정답 레이블 기술 적용 범위를 56.5%에서 99.5%로 증가시켰습니다. 또한 Qwen3.5-9B 모델에서는 백본 모델을 업데이트하지 않고도 가장 강력한 파라미터 업데이트 기반 모델보다 11.16% 더 높은 성능을 보였습니다. 전문가 평가 결과, GuideSkill은 임상적으로 타당하고 광범위하게 수용 가능한 기술을 생성하며, 이는 초기화 및 진화된 규칙이 신뢰할 수 있고 실질적인 의미를 갖는다는 것을 시사합니다. 이러한 결과는 실행 가능한 기술이 지침에서 파생된 절차와 사례에서 파생된 진단 패턴을 결합하는 모델에 독립적인 메커니즘임을 뒷받침합니다.
Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather than execute its rules. We introduce GuideSkill, an external reasoning layer that compiles disease-specific criteria into executable functions returning ordinal diagnostic-support scores. GuideSkill-Zero is initialized from guidelines, while GuideSkill-Evo uses case--diagnosis pairs to refine covered skills and add missing diagnoses. At inference, an LLM proposes a differential diagnosis, grounds the features required by each matched skill, and fuses its ranking with the executed skill scores. Across four benchmarks and four backbones, GuideSkill-Zero improves macro-average accuracy over guideline RAG by 13.45% on average. GuideSkill-Evo achieves the highest macro-average for every backbone, improves over direct inference by 18.49% relatively, and increases gold-label skill coverage from 56.5% to 99.5%. On Qwen3.5-9B, it also exceeds the strongest parameter-update baseline by 11.16% without updating the backbone. Expert evaluation further indicates that GuideSkill produces clinically sound and broadly acceptable skills, suggesting that its initialized and evolved rules are reliable and practically meaningful. These results support executable skills as a model-agnostic mechanism for combining guideline-derived procedures with case-derived diagnostic patterns.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.