기울기는 그 영향력을 얻어야 합니다: 일반화된 엔트로피 목적 함수를 이용한 지도 미세 조정의 통합
Gradients Must Earn Their Influence: Unifying SFT with Generalized Entropic Objectives
지도 미세 조정(SFT)에 사용되는 표준적인 음의 로그 우도(NLL)는 토큰 수준에서 균일한 가중치를 적용합니다. 이러한 경직성은 다음과 같은 두 가지 문제점을 야기합니다. (i) 낮은 확률의 목표에 대한 과도한 강조는 노이즈가 많은 지도 데이터에 대한 기울기를 증폭시켜 견고한 사전 지식을 방해할 수 있고, (ii) 균일한 가중치는 모델이 이미 확신을 가지고 있을 때 효과적인 성능 향상을 제공하지 못합니다. 기존 방법들은 이러한 불안정성-안정성 딜레마를 해결하지 못하며, 종종 유용한 학습 신호를 억제하는 동시에 해로운 신호도 함께 억제합니다. 이러한 문제를 해결하기 위해, 우리는 토큰 수준의 SFT 목적 함수를 일반화된 변형된 로그 함수 내에서 통합하고, 모델이 현재 예측에 얼마나 신뢰를 부여하는지를 제어하는 보편적인 게이트 × 오차 기울기 구조를 제시합니다. 케일리 변환을 사용하여 모델의 지속적으로 변화하는 불확실성을 연속적인 집중 경로로 매핑함으로써, 불확실한 새로운 개념과 확립된 지식 간의 시나리오 간에 원활하게 보간할 수 있습니다. 또한, 우리는 동적 엔트로피 미세 조정(DEFT)이라는 파라미터가 없는 목적 함수를 도입합니다. DEFT는 모델의 예측 상태를 나타내는 실제적인 지표로서 분포 집중도(레니-2 엔트로피)를 사용하여 신뢰 게이트를 조절합니다. 광범위한 실험과 분석 결과, DEFT는 탐색과 활용 사이의 균형을 개선하여 전반적인 성능을 향상시키는 것으로 나타났습니다.
Standard negative log-likelihood (NLL) for Supervised Fine-Tuning (SFT) applies uniform token-level weighting. This rigidity creates a two-fold failure mode: (i) overemphasizing low-probability targets can amplify gradients on noisy supervision and disrupt robust priors, and (ii) uniform weighting provides weak sharpening when the model is already confident. Existing methods fail to resolve the resulting plasticity--stability dilemma, often suppressing necessary learning signals alongside harmful ones. To address this issue, we unify token-level SFT objectives within a generalized deformed-log family and expose a universal gate $\times$ error gradient structure, where the gate controls how much the model trusts its current prediction. By employing the Cayley transform, we map the model's continuously evolving uncertainty onto a continuous focus trajectory, which enables seamless interpolation between scenarios involving uncertain novel concepts and those involving well-established knowledge. We then introduce Dynamic Entropy Fine-Tuning (DEFT), a parameter-free objective that modulates the trust gate using distribution concentration (Rényi-2 entropy) as a practical proxy for the model's predictive state. Extensive experiments and analyses demonstrate that DEFT achieves a better balance between exploration and exploitation, leading to improved overall performance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.