LLM을 활용한 스키마 적응형 표형 데이터 표현 학습: 일반화 가능한 다중 모달 임상 추론을 위한 방법
Schema-Adaptive Tabular Representation Learning with LLMs for Generalizable Multimodal Clinical Reasoning
표형 데이터에 대한 머신러닝은 일반적으로 스키마 일반화 능력 부족으로 인해 제약이 있으며, 이는 구조화된 변수에 대한 의미론적 이해 부족에서 비롯된 문제입니다. 이러한 문제는 특히 전자 건강 기록(EHR) 스키마가 크게 다른 임상 의학과 같은 분야에서 더욱 심각합니다. 이 문제를 해결하기 위해, 우리는 대규모 언어 모델(LLM)을 활용하여 전이 가능한 표형 임베딩을 생성하는 새로운 방법인 스키마 적응형 표형 데이터 표현 학습을 제안합니다. 우리 방법은 구조화된 변수를 의미론적 자연어 문장으로 변환하고, 사전 훈련된 LLM을 사용하여 이를 인코딩함으로써, 수동적인 특징 엔지니어링이나 재훈련 없이도 새로운 스키마에 대한 제로샷 정렬을 가능하게 합니다. 우리는 제안하는 인코더를 다중 모달 프레임워크에 통합하여 치매 진단을 수행하며, 표형 데이터와 MRI 데이터를 결합합니다. NACC 및 ADNI 데이터 세트에 대한 실험 결과, 우리 방법은 최첨단 성능을 보여주며, 새로운 스키마에 대한 성공적인 제로샷 전이를 달성했습니다. 또한, 후향적 진단 작업에서 신경학 전문의를 포함한 기존 임상 모델보다 훨씬 뛰어난 성능을 보였습니다. 이러한 결과는 우리 LLM 기반 접근 방식이 이질적인 실제 데이터에 대한 확장 가능하고 강력한 솔루션임을 입증하며, LLM 기반 추론을 구조화된 영역으로 확장할 수 있는 방법을 제시합니다.
Machine learning for tabular data remains constrained by poor schema generalization, a challenge rooted in the lack of semantic understanding of structured variables. This challenge is particularly acute in domains like clinical medicine, where electronic health record (EHR) schemas vary significantly. To solve this problem, we propose Schema-Adaptive Tabular Representation Learning, a novel method that leverages large language models (LLMs) to create transferable tabular embeddings. By transforming structured variables into semantic natural language statements and encoding them with a pretrained LLM, our approach enables zero-shot alignment across unseen schemas without manual feature engineering or retraining. We integrate our encoder into a multimodal framework for dementia diagnosis, combining tabular and MRI data. Experiments on NACC and ADNI datasets demonstrate state-of-the-art performance and successful zero-shot transfer to unseen schemas, significantly outperforming clinical baselines, including board-certified neurologists, in retrospective diagnostic tasks. These results validate our LLM-driven approach as a scalable, robust solution for heterogeneous real-world data, offering a pathway to extend LLM-based reasoning to structured domains.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.