TRIAGE: 설명 가능한 의료 시계열 데이터 예측을 위한 대화형 추론 - LLM 활용, 불규칙 샘플링 기반
TRIAGE: Dialectical Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series with LLMs
전자 건강 기록(EHR)에 기반한 임상 조기 경고 시스템은 환자 분류를 위한 정확한 위험 점수와 더불어, 의료진이 검증할 수 있는 해석 가능한 근거를 제공해야 합니다. 대규모 언어 모델(LLM)은 이러한 작업에 활용되어 왔지만, 미세한 임상적 위험을 과도하게 확신하는 이분법적인 예측으로 축소시키는 경향이 있습니다. 이러한 위험 양극화는 정확성과 환자 간 비교 가능성을 저해합니다. 이를 해결하기 위해, 우리는 LLM이 경쟁적인 임상 결과를 바탕으로 결과별 근거를 도출하여 대화형 추론을 생성하도록 학습하는 프레임워크인 TRIAGE를 제안합니다. 이 대화형 구조는 위험 양극화를 완화하고, 하나의 LLM이 명시적인 임상적 추론에 기반한 연속적인 위험 점수를 제공할 수 있도록 합니다. 세 가지 ISMTS(불규칙 샘플링 의료 시계열) 벤치마크에서 TRIAGE를 평가한 결과, 평균적으로 AUPRC 성능이 3.3% 향상되었고, 교정 오차가 81% 감소했습니다. 또한, LLM을 활용한 판단 기준으로 평가한 결과, 우리 모델의 근거가 기준 모델의 사후 설명보다 임상적 추론 품질 측면에서 20% 더 우수했습니다. 소스 코드는 https://github.com/HyeongWon-Jang/TRIAGE 에서 확인할 수 있습니다.
Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time series (ISMTS), must deliver both calibrated risk scores for patient triage and interpretable rationales that clinicians can verify. Large Language Models (LLMs) have been explored for this task, yet they collapse graded clinical risk into overconfident binary predictions. This risk polarization undermines both calibration and cross-patient comparability. To address this, we propose TRIAGE, a framework that trains an LLM to generate dialectical reasoning over competing clinical outcomes by eliciting outcome-specific rationales. This dialectical formulation mitigates risk polarization, enabling a single LLM to yield continuous risk scores grounded in explicit clinical reasoning. Evaluated on three ISMTS benchmarks, TRIAGE achieves an average AUPRC improvement of 3.3% and reduces calibration error by 81% compared to the competitive baselines. An LLM-as-a-judge assessment further shows that our rationales surpass post-hoc explanations from the baseline by 20% in clinical reasoning quality. The source code is available at https://github.com/HyeongWon-Jang/TRIAGE .
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.