종단 간 LLM 기반의 검열(censoring)을 고려한 생존 분석 연구
Towards end-to-end LLM-based censoring-aware survival analysis
목표: 생존 분석은 의료 예측에서 중요한 역할을 하지만, 대규모 언어 모델(LLM)이 종단 간 생존 모델로 거의 사용되지 않는데, 이는 검열(censoring) 때문에 직접적인 지도 학습 방식의 미세 조정이 어렵기 때문입니다. 본 연구에서는 LLMSurvival이라는 프레임워크를 제시합니다. 이 프레임워크는 수정되지 않은 LLM을 사용하여 테이블 형태의 임상 데이터에 직접 적용하여 검열을 고려한 생존 분석을 가능하게 합니다. 방법: LLMSurvival은 시간-사건 예측 문제를 훈련 코호트에서 선택된 기준 대상과의 비교를 통해 이루어지는 쌍대 순위 결정 문제로 재구성하고, 테스트 시 위험도를 이러한 비교 결과들을 종합하여 산출합니다. 결과: 두 가지 임상 과제(MIMIC-IV 데이터셋을 사용한 중환자실 사망 예측 및 NewYork-Presbyterian/Weill Cornell Medicine 코호트를 사용한 골다공증성 골절 예측)에서, LLMSurvival은 Cox 비례 위험 모델 대비 전체적으로 3.1% 더 나은 결과를 보였으며(중환자실 사망), 골절 위험의 경우 0.5% 더 나은 결과를 보였습니다. 또한, 기존의 세 가지 심층 학습 기반 생존 분석 모델 대비 평균적으로 중환자실 사망 예측에서 2.1%, 골절 위험 예측에서 2.8% 더 우수한 성능을 나타냈습니다. 논의: 본 연구 결과는 검열을 고려한 생존 모델링이 비교 기반의 재구성을 통해 LLM 미세 조정과 호환될 수 있음을 보여줍니다. 또한, 이 프레임워크는 다양한 임상 환경에서 전문가가 직접 설계한 점수(예: SAPS-II 및 FRAX)보다 높은 휴대성과 우수한 성능을 제공합니다. 게다가, 이 프레임워크는 로컬 배포를 지원하는데, 이는 크기가 작고 공개적으로 사용 가능한 기본 모델이 충분한 성능을 제공하기 때문입니다. 결론: LLMSurvival 프레임워크는 LLM을 활용한 통합적이고 검열을 고려한 생존 분석 접근 방식의 실현 가능성을 보여주는 연구입니다.
Objective: Survival analysis is central to medical prediction, yet large language models (LLMs) are rarely used as end-to-end survival models because censoring prevents straightforward supervised fine-tuning. Here we present LLMSurvival, a framework that enables censoring-aware survival analysis with unmodified LLMs operating directly on tabular clinical data. Materials and Methods: LLMSurvival reformulates time-to-event prediction as pairwise ranking among comparable subjects, and derives test-time risk by aggregating comparisons against anchor individuals from the training cohort. Results: Across two clinical tasks (ICU mortality prediction in MIMIC-IV and fragility fracture prediction in a NewYork-Presbyterian/Weill Cornell Medicine cohort), LLMSurvival improves overall concordance over Cox proportional hazards modeling by 3.1% for ICU mortality and 0.5% for fracture risk, 2.1% on average for ICU mortality and 2.8% for fracture risk over three established deep learning survival models. Discussion: The results show that survival modeling with censoring can be made compatible with LLM fine-tuning through comparison-based reformulation. The framework demonstrates high portability and superior performance over expert curated scores like SAPS-II and FRAX scores across diverse clinical context. Furthermore, the framework supports local deployment, as compact, publicly available base models provide sufficient performance. Conclusion: The LLMSurvival framework serves as a proof of concept for an integrated, censoring-conscious approach to survival analysis via LLMs.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.