2606.10279v1 Jun 09, 2026 cs.AI

합성적인 근거 데이터로 한 supervised fine-tuning은 실제 세계의 질병 예측 성능을 저하시킨다

Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction

Cheng Qian
Cheng Qian
Citations: 243
h-index: 9
Yiwei Wang
Yiwei Wang
Citations: 490
h-index: 13
Buxin Su
Buxin Su
University of Pennsylvania
Citations: 45
h-index: 4
Bingxin Zhao
Bingxin Zhao
Citations: 7
h-index: 1
Bingxuan Li
Bingxuan Li
Citations: 260
h-index: 4
Jinran Jin
Jinran Jin
Citations: 7
h-index: 1

합성적인 근거 데이터를 활용한 supervised fine-tuning은 모델이 무엇을 예측해야 하는지 뿐만 아니라 왜 예측해야 하는지를 학습시켜 임상 예측 작업에서 언어 모델의 성능을 향상시킨다는 것이 널리 알려져 있습니다. 본 연구에서는 장기간 건강 기록으로부터 5년 알츠하이머병 및 관련 치매(ADRD)를 예측하는 문제에 대해 이러한 가설을 검증합니다. 대규모로 설계된 504개의 실험 구성에서, 근거 기반의 SFT는 항상 일관적이고 현저하게 예측 성능을 저하시킨다는 것을 발견했습니다. 이러한 성능 저하 현상은 모델 종류 및 데이터 규모에 관계없이 지속되며, 추론 중심의 기본 모델을 사용하더라도 해결되지 않습니다. 중요한 점은, 이러한 실패가 부실한 근거 품질 때문이 아니라는 것입니다. 인간 전문가의 주석 결과, 생성된 근거는 의학적으로 정확하며 환자별 증거에 기반하고 있는 것으로 확인되었으며, few-shot 실험에서는 동일한 근거가 학습 목표로 사용될 때보다 추론 시 데모 자료로 사용될 때 성능을 향상시키는 것을 보여줍니다. 우리는 이러한 실패의 근본 원인이 서술적 타당성과 판별력 최적화 간의 구조적인 충돌이라고 밝히고자 합니다. 본 연구는 언제 그리고 어떻게 근거 기반의 감독 학습이 도움이 되는지, 그리고 그렇지 않을 때를 더 정확하게 이해하는 데 기여하며, 고위험 임상 예측을 위한 언어 모델의 책임 있는 개발을 위한 지침을 제공할 수 있기를 바랍니다.

Original Abstract

Supervised fine-tuning with synthetic rationale data is widely assumed to improve language model performance on clinical prediction tasks by teaching models not just what to predict but why. We test this assumption on five-year Alzheimer's disease and related dementias (ADRD) prediction from longitudinal health histories. Across a large-scale controlled experiment of 504 configurations, we find that rationale-based SFT consistently and substantially hurts prediction performance relative to label-only fine-tuning. The degradation persists across model families and data scales, and is not resolved by using a reasoning-oriented base model. Crucially, the failure is not explained by poor rationale quality: human expert annotation confirms that the generated rationales are medically accurate and faithfully grounded in patient-specific evidence, and few-shot experiments show that the same rationales improve performance when used as inference-time demonstrations rather than training targets. We identify the root cause as a structural conflict between narrative plausibility and discriminative optimization. We hope our work paves the path toward a more precise understanding of when and how rationale-based supervision helps and when it does not, guiding the responsible development of language models for high-stakes clinical prediction.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!