2606.11189v1 Jun 09, 2026 cs.LG

타겟 분포 설계 관점에서 본 지도 학습 미세 조정에 대한 통합적 접근

A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design

Yihang Chen
Yihang Chen
Citations: 3
h-index: 1
Yuanhao Ban
Yuanhao Ban
Citations: 232
h-index: 5
Yunqi Hong
Yunqi Hong
Citations: 15
h-index: 2
Sohyun An
Sohyun An
Citations: 65
h-index: 4
Tong Xie
Tong Xie
Citations: 37
h-index: 3
Cho-Jui Hsieh
Cho-Jui Hsieh
Citations: 8
h-index: 1

지도 학습 미세 조정(SFT)은 일반적으로 주어진 시퀀스 내의 모든 토큰에 대한 가능도를 최대화합니다. 그러나 관찰된 토큰이 고유하지 않거나, 노이즈를 포함하거나, 모델의 사전 지식과 일치하지 않을 수 있습니다. 이러한 특정 토큰만을 엄격하게 학습하는 것은 최적 이하일 수 있으며, 특히 사전 훈련된 모델이 풍부한 지식 기반을 가지고 있을 때 더욱 그렇습니다. 본 연구에서는 SFT를 타겟 분포 설계 문제로 재해석합니다. 손실 함수의 목표 값 자체뿐만 아니라, 해당 손실 함수가 모델에게 학습하도록 요구하는 토큰 수준의 타겟 분포에 주목하여 분석합니다. 우리는 Q-타겟 프레임워크를 도입하며, 이는 SFT의 감독 신호를 두 가지 명확한 선택으로 분해합니다: (1) 관찰된 토큰을 얼마나 강하게 활용할 것인가, 그리고 (2) 남은 확률 질량을 대안적인 토큰에 어떻게 할당할 것인가. 이러한 관점은 기존의 다양한 SFT 방법들을 타겟 분포 Q에 대한 암묵적인 선택으로 통합합니다. 이러한 관점을 바탕으로, 우리는 원하는 타겟 분포로부터 직접 학습 목표를 구성하는 Target-SFT 방법을 제안합니다. 이 방법은 평가된 10개의 추론 데이터셋-모델 설정에서 일관되게 우수한 성능을 보이며, 본 연구의 타겟 기반 접근 방식의 효과성을 입증합니다. 종합적으로, 본 연구는 SFT 학습에 대한 더욱 근본적인 설계 원리를 제시하며, SFT 목표 함수 탐색 공간을 확장합니다.

Original Abstract

Supervised fine-tuning (SFT) typically maximizes the likelihood of every token in a demonstrated trajectory. However, an observed token can be non-unique, noisy, or misaligned with the model prior. Strictly fitting toward this one-hot target may be suboptimal, especially when the pretrained model encodes a rich knowledge prior. In this work, we reinterpret SFT as target distribution design: instead of studying only the loss objective, we analyze the token-level target that the loss drives the model to match. We introduce the Q-target framework, which decomposes SFT supervision into two explicit choices: (1) how strongly to rely on the observed token, and (2) how to allocate the remaining probability mass over alternatives. This perspective unifies many existing SFT variants as implicit choices of the target distribution Q. Building on this view, we propose Target-SFT which constructs the training objective directly from the desired target distribution. This method consistently outperforms across the ten reasoning dataset-model settings evaluated, showing the effectiveness of this target-based approach. Overall, our formulation reveals a more fundamental design principle for SFT training and opens a broader search space for SFT objectives.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!