2608.03573v1 Aug 04, 2026 cs.CL

SFT 충돌과 RL 공존: LLM을 위한 다중 작업 학습에 대한 이론적 및 경험적 분석

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Hongbang Yuan
Hongbang Yuan
Citations: 561
h-index: 8
Shangqing Tu
Shangqing Tu
Citations: 1,574
h-index: 12
Kejian Zhu
Kejian Zhu
Citations: 30
h-index: 3

지도 미세 조정(SFT)과 강화 학습(RL)은 대규모 언어 모델(LLM)의 다중 작업 추론 능력을 향상시키는 데 있어 근본적으로 다른 동작 방식을 보입니다. 초기 실험 결과, SFT는 다단계 훈련 과정에서 심각한 작업 간 충돌을 겪는 반면, RL은 다양한 작업에 걸쳐 안정적인 공존을 가능하게 하는 현상을 발견했습니다. 경험적으로, 이러한 현상은 파라미터 수준에서 관찰되며, RL은 작업 간에 희소하고 거의 직교하는 업데이트를 유도합니다. 다중 작업 기울기 간섭을 분석하여 이 메커니즘에 대한 이론적 설명을 제공합니다. 우리의 결과는 SFT의 간섭이 노름-제한되어 절대 기울기 크기에 비례하는 반면, RL의 간섭은 분산-제한되어 어드밴티지 정규화 및 온폴리시 최적화로 인해 발생하는 기울기 분산에 의해 제한된다는 것을 보여줍니다. 이러한 작은 분산 경계는 작업 전체에서 거의 직교하는 최적화 방향을 제공합니다. 이러한 통찰력을 바탕으로, 다중 작업 훈련을 분리하여 효율성과 유연성을 크게 향상시키는 새로운 패러다임인 Parallel-RL을 제안합니다.

Original Abstract

Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage training, whereas RL enables stable coexistence across diverse tasks. Empirically, we trace this to the parameter level, observing that RL induces sparse and approximately orthogonal updates across tasks. We provide a theoretical explanation for this mechanism by analyzing multi-task gradient interference. Our results reveal a distinction: interference in SFT is norm-limited, scaling with the absolute gradient magnitude, whereas interference in RL is variance-limited, bounded by the gradient variance induced by advantage normalization and on-policy optimization. This small variance bound yields near-orthogonal optimization directions across tasks. Leveraging this insight, we propose Parallel-RL, a paradigm that decouples multi-task training, significantly improving efficiency and flexibility.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!