불규칙적인 임상 시계열 데이터에 대한 질문 응답을 위한 비용 효율적인 다중 모드 LLM 추론 프레임워크
A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series
불규칙적인 임상 시계열(ICTS) 데이터를 이용한 질의응답(QA)은 광범위한 의료 분야에서 중요한 역할을 합니다. 최근 개발된 다중 모드 시계열 대규모 언어 모델(LLM)이 일반적인 시계열 QA에서 상당한 가능성을 보여주었지만, 임상 관찰의 희소성, 비동기성 및 불규칙 샘플링 패턴을 효과적으로 모델링하는 데는 여전히 부족합니다. 이러한 격차를 해소하기 위해, 본 논문에서는 ICTS 데이터에 대한 질문 응답을 위한 비용 효율적인 다중 모드 LLM 추론 프레임워크인 ClinPRISM을 제안합니다. 먼저, 다양한 시간 척도에서 희소한 임상 증거를 포착할 수 있는 불규칙성 인지 다중 척도 인코더를 설계했습니다. 다음으로, 이러한 척도 간의 표현을 통합하고 LLM과 호환되는 소수의 토큰으로 압축하는 Temporal Evidence Distiller를 제안합니다. 또한, 불규칙한 시계열 데이터를 순차적으로 LLM의 텍스트 임베딩 공간에 정렬하는 Progressive Alignment 전략을 도입했습니다. 학습 편의성을 위해, 다중 척도 설명을 포함한 30,000개의 임상 시계열 데이터와 함께 11가지 작업에 걸쳐 41,000개의 instruction-tuning 인스턴스를 구축했습니다. 40억 개의 파라미터를 가진 LLM을 기반으로 ClinPRISM은 기존의 평가 벤치마크에서 최첨단 성능을 달성했으며, 단 16개의 시계열 토큰만을 사용하고 평균 추론 지연 시간이 0.15초/질문이라는 놀라운 효율성을 보여줍니다.
Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable promise in general-purpose time-series QA, they remain poorly equipped to model the sparsity, asynchrony, and irregular sampling patterns of clinical observations. To fill this gap, we propose ClinPRISM, a cost-effective multimodal LLM reasoning framework for question answering over ICTS data. First, we devise an irregularity-aware multi-scale encoder to capture sparse clinical evidence at diverse temporal scales. Then, we propose a temporal evidence distiller to integrate representations across these scales and compress them into a small number of LLM-compatible tokens. Moreover, we introduce a progressive alignment strategy that sequentially aligns the irregular trajectories with the LLM's textual embedding space. To facilitate training, we construct 30,000 clinical time series paired with multi-scale descriptions, together with 41,000 instruction-tuning instances spanning 11 tasks. Using a 4-billion-parameter LLM backbone, ClinPRISM achieves state-of-the-art performance on the held-out evaluation benchmark while using only 16 time-series tokens and achieving an average inference latency of 0.15 seconds per question.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.