2608.04180v1 Aug 04, 2026 cs.LG

EHR 진단 코드 기반 오피오이드 사용 장애 예측을 위한 특징 선택 방법 비교 연구

A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction

Zihan Ding
Zihan Ding
Citations: 36
h-index: 3
Yinan Liu
Yinan Liu
Citations: 13
h-index: 2
Tengfei Ma
Tengfei Ma
Citations: 13
h-index: 2
Rachel Wong
Rachel Wong
Citations: 143
h-index: 3
George S. Leibowitz
George S. Leibowitz
Citations: 15
h-index: 3
Benjamin Littenberg
Benjamin Littenberg
Citations: 3,340
h-index: 8
Richard N Rosenthal
Richard N Rosenthal
Citations: 14
h-index: 2
Xiaohong Zheng
Xiaohong Zheng
Citations: 0
h-index: 0
Fusheng Wang
Fusheng Wang
Citations: 835
h-index: 15

전자 건강 기록(EHR) 기반 예측 모델링에서, 입력 변수는 종종 고차원적이고 희소하며 노이즈가 많고 중복되는 경향이 있습니다. 큰 특징 집합은 계산 부담과 과적합 위험을 증가시킬 뿐만 아니라 모델 해석을 어렵게 만들어 임상 환경에서의 유용성을 제한합니다. 본 연구에서는 진단 관련 특징에 초점을 맞추어, 오피오이드 사용 장애(OUD) 예측을 위한 다섯 가지 특징 선택 방법을 비교 분석합니다: 재발 빈도 풍부화, NTK 기반 초기 기울기 민감성, LightGBM-SHAP, 엘라스틱 넷, 그리고 대규모 언어 모델(LLM) 기반 의미론적 선택. 우리는 통일된 전처리 및 평가 프레임워크를 사용하고, 각 방법을 예측 성능, 재샘플링 안정성 및 빈도가 낮은 진단 코드의 표현력을 기준으로 평가합니다. 연구 결과는 특징 집합 크기가 증가함에 따라 성능이 향상되지만, 어느 정도 수준을 넘어서면 그 효과가 감소하는 것을 보여줍니다. NTK 민감성은 정확성과 안정성의 균형이 가장 우수하며, LLM 기반 선택은 독립적인 성능은 낮지만 임상적으로 의미 있는 추가 정보를 제공합니다.

Original Abstract

Feature selection is a critical step in electronic health record (EHR)-based predictive modeling, where input variables are often high-dimensional, sparse, noisy, and redundant. Large feature sets not only increase computational burden and overfitting risk, but also make model interpretation difficult, leading to limited usefulness in clinical settings. In this study, we focus on diagnosis-related features and compare five feature selection paradigms for opioid use disorder (OUD) prediction: recurrence enrichment, NTK-motivated early gradient sensitivity, LightGBM-SHAP, Elastic Net, and large language model (LLM)-guided semantic selection. We use a unified preprocessing and evaluation framework and assess each method by downstream predictive performance, resampling stability, and representation of infrequent diagnosis codes. Our results demonstrate that performance improves with larger feature budgets with diminishing returns beyond a moderate size. NTK sensitivity provides the best overall balance of accuracy and stability, and LLM-guided selection contributes complementary clinically meaningful signals despite lower standalone performance.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!