T2D-Bench: 다층 임상-생활 습관 지식 그래프를 활용한 LLM 출력에 대한 증거 기반 평가 - 제2형 당뇨병
T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph
대규모 언어 모델(LLM)은 제2형 당뇨병 환자를 위한 임상적으로 유창한 권장 사항을 생성할 수 있지만, 동시에 가이드라인의 제약을 충족하지 못하거나 생활 습관과 관련된 혈당 조절 주장을 명확하게 뒷받침하지 못하는 경우가 있습니다. 본 연구에서는 LLM 출력이 명시적인 증거 요건을 충족하는지 테스트하기 위한 재현 가능한 벤치마크 및 증거 기반 평가 프레임워크인 T2D-Bench를 제시합니다. T2D-Bench는 생체의학 데이터(UMLS, DrugBank, SIDER), 계산 가능한 ADA 치료 지침, 그리고 기계적 연관성을 통해 혈당 관련 실험실 결과와 연결된 생활 습관 지식을 결합한 다층 임상-생활 습관 지식 그래프를 기반으로 구축되었습니다. 진단, 약물 안전성 및 잠재적인 생활 습관 충돌을 포함하는 100개의 구조화된 시나리오에서, GPT-4o-mini 모델은 35%, GPT-4o 모델은 33%의 경우에 벤치마크 정의된 증거 경로 검증에 실패했습니다. 이 증거 게이트는 근거 없는 누락을 감지하고 제한적인 수정 과정을 통해 출력을 벤치마크에서 정의한 증거 요건에 부합하도록 만듭니다. 이러한 결과는 계산 가능한 증거 제약 조건을 통해 당뇨병 관련 LLM 출력에서 근거 없는 임상적 누락을 명확하게 드러내고 측정 가능하며 수정할 수 있음을 보여줍니다.
Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explicitly justify lifestyle-related glycemic claims. We present T2D-Bench, a reproducible benchmark and evidence-gated evaluation framework for testing whether LLM outputs satisfy explicit, graph-checkable evidence requirements. T2D-Bench is built on a multi-layer clinical-lifestyle knowledge graph that combines a biomedical spine (UMLS, DrugBank, SIDER), computable ADA Standards of Care rules, and lifestyle knowledge connected through a mechanistic bridge to glycemic laboratory effects. Across 100 structured vignettes spanning diagnosis, medication safety, and adversarial lifestyle conflicts, baseline outputs failed benchmark-defined evidence-path checks in 35% of cases for GPT-4o-mini and 33% for GPT-4o. The evidence gate detects unsupported omissions and uses constrained revision to bring outputs into verifier-level compliance with benchmark-defined evidence requirements. These results show that computable evidence constraints can make unsupported clinical omissions explicit, measurable, and correctable in diabetes-focused LLM outputs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.