LingLanMiDian: 중의학 지식 및 임상 추론에 대한 LLM의 체계적 평가
LingLanMiDian: Systematic Evaluation of LLMs on TCM Knowledge and Clinical Reasoning
거대 언어 모델(LLM)이 의료 NLP 분야에서 빠르게 발전하고 있지만, 독특한 온톨로지, 용어 및 추론 패턴을 가진 중의학(TCM)은 해당 도메인에 충실한 평가를 필요로 합니다. 기존의 TCM 벤치마크는 범위와 규모 면에서 파편화되어 있고, 통일되지 않거나 생성 중심의 채점 방식에 의존하여 공정한 비교를 저해했습니다. 이에 우리는 지식 회상, 멀티 홉 추론, 정보 추출 및 실제 임상 의사 결정을 아우르는 통합 평가를 위해 전문가가 엄선한 대규모 멀티 태스크 스위트인 LingLanMiDian(LingLan) 벤치마크를 제안합니다. LingLan은 일관된 메트릭 설계, 임상 레이블에 대한 동의어 허용 프로토콜, 데이터셋별 400개 항목의 고난도(Hard) 서브셋, 그리고 진단 및 치료 권고를 단일 선택형 의사 결정 인식으로 재구성하는 방식을 도입했습니다. 우리는 14개의 주요 오픈 소스 및 독점 LLM을 대상으로 포괄적인 제로 샷 평가를 수행하여, TCM 상식 지식 이해, 추론 및 임상 의사 결정 지원에 있어서 모델들의 강점과 한계에 대한 통합된 관점을 제공합니다. 특히, 고난도 서브셋에 대한 평가는 TCM 전문 추론 영역에서 현재 모델과 인간 전문가 사이에 상당한 격차가 있음을 보여줍니다. 표준화된 평가를 통해 기초 지식과 응용 추론을 연결함으로써, LingLan은 TCM LLM 및 도메인 특화 의료 AI 연구의 발전을 위한 통합적이고 정량적이며 확장 가능한 기반을 마련합니다. 모든 평가 데이터와 코드는 https://github.com/TCMAI-BJTU/LingLan 및 http://tcmnlp.com 에서 공개됩니다.
Large language models (LLMs) are advancing rapidly in medical NLP, yet Traditional Chinese Medicine (TCM) with its distinctive ontology, terminology, and reasoning patterns requires domain-faithful evaluation. Existing TCM benchmarks are fragmented in coverage and scale and rely on non-unified or generation-heavy scoring that hinders fair comparison. We present the LingLanMiDian (LingLan) benchmark, a large-scale, expert-curated, multi-task suite that unifies evaluation across knowledge recall, multi-hop reasoning, information extraction, and real-world clinical decision-making. LingLan introduces a consistent metric design, a synonym-tolerant protocol for clinical labels, a per-dataset 400-item Hard subset, and a reframing of diagnosis and treatment recommendation into single-choice decision recognition. We conduct comprehensive, zero-shot evaluations on 14 leading open-source and proprietary LLMs, providing a unified perspective on their strengths and limitations in TCM commonsense knowledge understanding, reasoning, and clinical decision support; critically, the evaluation on Hard subset reveals a substantial gap between current models and human experts in TCM-specialized reasoning. By bridging fundamental knowledge and applied reasoning through standardized evaluation, LingLan establishes a unified, quantitative, and extensible foundation for advancing TCM LLMs and domain-specific medical AI research. All evaluation data and code are available at https://github.com/TCMAI-BJTU/LingLan and http://tcmnlp.com.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.