2602.11360v1 Feb 11, 2026 cs.LG

부트스트래핑 기반 정규화를 통한 임상 위험 예측 모델의 개별 예측 불안정성 감소

Bootstrapping-based Regularisation for Reducing Individual Prediction Instability in Clinical Risk Prediction Models

C. Yau
C. Yau
Citations: 238
h-index: 8
Sara Matijevic
Sara Matijevic
Citations: 12
h-index: 2

임상 예측 모델은 환자 치료를 지원하는 데 점점 더 많이 사용되고 있지만, 많은 딥러닝 기반 접근 방식은 여전히 불안정하며, 동일 인구에서 추출된 서로 다른 데이터셋으로 훈련할 때 예측 값이 크게 달라질 수 있습니다. 이러한 불안정성은 신뢰성을 저해하고 임상 적용을 제한합니다. 본 연구에서는 딥 신경망 훈련 과정에 부트스트래핑 프로세스를 직접 통합하는 새로운 부트스트래핑 기반 정규화 프레임워크를 제안합니다. 이 접근 방식은 재샘플링된 데이터셋 간의 예측 변동성을 제한하여, 내재적인 안정성을 갖는 단일 모델을 생성합니다. 제안된 정규화 방식을 사용하여 구축된 모델을 시뮬레이션 데이터와 GUSTO-I, Framingham, SUPPORT의 세 가지 임상 데이터셋을 사용하여 기존 모델 및 앙상블 모델과 비교 평가했습니다. 모든 데이터셋에서 제안된 모델은 예측 안정성이 향상되었으며, 평균 절대 차이(예: GUSTO-I에서 0.019 vs. 0.059; Framingham에서 0.057 vs. 0.088)가 낮고, 유의미하게 벗어나는 예측의 수가 현저히 감소했습니다. 또한, 판별 성능과 특징 중요성 일관성이 유지되었으며, 모델 간의 높은 SHAP 상관 관계(예: GUSTO-I에서 0.894; Framingham에서 0.965)를 보였습니다. 앙상블 모델은 더 높은 안정성을 달성했지만, 이는 각 구성 모델이 예측 변수를 서로 다른 방식으로 사용했기 때문에 해석 가능성이 저하되는 결과를 초래했습니다. 제안된 접근 방식은 부트스트래핑 분포에 맞춰 예측을 정규화함으로써, 해석 가능성을 희생하지 않고 더 큰 견고성과 재현성을 갖춘 예측 모델을 개발할 수 있도록 합니다. 이 방법은 특히 데이터가 제한적인 의료 환경에서 더욱 신뢰할 수 있고 임상적으로 신뢰할 수 있는 딥러닝 모델을 개발하는 데 실질적인 경로를 제공합니다.

Original Abstract

Clinical prediction models are increasingly used to support patient care, yet many deep learning-based approaches remain unstable, as their predictions can vary substantially when trained on different samples from the same population. Such instability undermines reliability and limits clinical adoption. In this study, we propose a novel bootstrapping-based regularisation framework that embeds the bootstrapping process directly into the training of deep neural networks. This approach constrains prediction variability across resampled datasets, producing a single model with inherent stability properties. We evaluated models constructed using the proposed regularisation approach against conventional and ensemble models using simulated data and three clinical datasets: GUSTO-I, Framingham, and SUPPORT. Across all datasets, our model exhibited improved prediction stability, with lower mean absolute differences (e.g., 0.019 vs. 0.059 in GUSTO-I; 0.057 vs. 0.088 in Framingham) and markedly fewer significantly deviating predictions. Importantly, discriminative performance and feature importance consistency were maintained, with high SHAP correlations between models (e.g., 0.894 for GUSTO-I; 0.965 for Framingham). While ensemble models achieved greater stability, we show that this came at the expense of interpretability, as each constituent model used predictors in different ways. By regularising predictions to align with bootstrapped distributions, our approach allows prediction models to be developed that achieve greater robustness and reproducibility without sacrificing interpretability. This method provides a practical route toward more reliable and clinically trustworthy deep learning models, particularly valuable in data-limited healthcare settings.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!