2603.21606v1 Mar 23, 2026 cs.LG

mSFT: 다중 작업 지도 미세 조정에서 데이터셋 혼합에 의한 과적합 문제 해결

mSFT: Addressing Dataset Mixtures Overfiting Heterogeneously in Multi-task SFT

Woosung Koh
Woosung Koh
Citations: 16
h-index: 3
Se-young Yun
Se-young Yun
Citations: 17
h-index: 3
Jaehyeon Choi
Jaehyeon Choi
Citations: 41
h-index: 3
Jeyoung Jeon
Jeyoung Jeon
Citations: 0
h-index: 0
Y. Cheon
Y. Cheon
Citations: 0
h-index: 0
Youngjin Song
Youngjin Song
Citations: 18
h-index: 3
Soowon Oh
Soowon Oh
Citations: 0
h-index: 0

현재의 언어 모델 학습은 일반적으로 모든 하위 데이터셋에 대해 균일한 컴퓨팅 예산을 사용하는 다중 작업 지도 미세 조정(SFT)을 적용합니다. 이러한 접근 방식은 근본적으로 최적이 아니며, 학습 속도가 빠른 작업은 초기 단계에서 과적합되는 반면, 학습 속도가 느린 작업은 여전히 과소적합됩니다. 이러한 문제를 해결하기 위해, 우리는 다중 작업 데이터셋 혼합에 대한 반복적이고 과적합 인지 검색 알고리즘인 mSFT를 소개합니다. mSFT는 활성 데이터셋 혼합을 사용하여 모델을 학습시키고, 가장 먼저 과적합되는 하위 데이터셋을 식별하여 제외한 후, 해당 최적의 체크포인트로 되돌아가 학습을 계속합니다. 광범위한 실험 결과는 mSFT가 10개의 벤치마크와 6개의 기본 모델에서 4가지 기준 모델보다 일관되게 우수한 성능을 보인다는 것을 보여줍니다. 추가 분석 결과, mSFT는 다양한 데이터셋 크기, 작업 세분성에서 견고한 성능 향상을 유지하며, 새로운 하이퍼파라미터(컴퓨팅 예산)에 둔감합니다. 특히, 낮은 컴퓨팅 예산에서도 mSFT는 성능을 향상시키면서 학습 FLOPs를 줄일 수 있습니다. 궁극적으로, mSFT는 다중 작업 SFT를 위한 실용적인 과적합 인지 알고리즘을 제시하며, 다양한 데이터셋 혼합에서 모델의 잠재력을 극대화합니다.

Original Abstract

Current language model training commonly applies multi-task Supervised Fine-Tuning (SFT) using a homogeneous compute budget across all sub-datasets. This approach is fundamentally sub-optimal: heterogeneous learning dynamics cause faster-learning tasks to overfit early while slower ones remain under-fitted. To address this, we introduce mSFT, an iterative, overfitting-aware search algorithm for multi-task data mixtures. mSFT trains the model on an active mixture, identifies and excludes the earliest overfitting sub-dataset, and reverts to that specific optimal checkpoint before continuing. Extensive evaluations demonstrate that mSFT consistently outperforms 4 baselines across 10 benchmarks and 6 base models. Further analysis confirms mSFT maintains robust gains across diverse dataset sizes, task granularities, and is insensitive to its single new hyperparameter (compute budget). Notably, at low compute budget, mSFT can improve performance while lowering training FLOPs. Ultimately, mSFT establishes a practical overfitting-aware algorithm for multi-task SFT that maximizes the potential of models across diverse data mixtures.

1 Citations
0 Influential
1.5 Altmetric
8.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!