2604.16259v1 Apr 17, 2026 cs.LG

분포 정교화 너머: 태스크 보상의 중요성

Beyond Distribution Sharpening: The Importance of Task Rewards

Sarthak Mittal
Sarthak Mittal
Citations: 1,037
h-index: 16
L'eo Gagnon
L'eo Gagnon
Citations: 29
h-index: 3
Guillaume Lajoie
Guillaume Lajoie
Citations: 55
h-index: 4

최첨단 모델들은 학습 과정에 태스크-보상 기반 강화 학습(RL)을 통합함으로써 탁월한 성능을 보여주었으며, 이를 통해 순수한 추론 모델에서 정교한 에이전트로 발전할 수 있게 되었습니다. 그러나 RL이 기본 모델에 실제로 새로운 기술을 부여하는 것인지, 아니면 단순히 기존 분포를 정교화하여 잠재적인 능력을 발휘하도록 하는 것인지에 대한 논쟁이 계속되고 있습니다. 이러한 이분법을 해결하기 위해, 우리는 분포 정교화와 태스크-보상 기반 학습을 명시적으로 비교하며, RL을 사용하여 두 가지 패러다임을 모두 구현합니다. 우리의 분석은 분포 정교화의 고유한 한계를 드러내며, 최적값이 어떻게 그리고 왜 불리할 수 있는지, 그리고 이 접근 방식이 근본적으로 불안정하다는 것을 원리적으로 설명합니다. 또한, Llama-3.2-3B-Instruct, Qwen2.5-3B-Instruct 및 Qwen3-4B-Instruct-2507을 사용하여 수학 데이터 세트에서 수행한 실험 결과, 정교화는 제한적인 성능 향상만을 가져오는 반면, 태스크 기반 보상 신호를 통합하면 강력하고 안정적인 성능 향상을 달성하는 데 크게 기여한다는 것을 확인했습니다.

Original Abstract

Frontier models have demonstrated exceptional capabilities following the integration of task-reward-based reinforcement learning (RL) into their training pipelines, enabling systems to evolve from pure reasoning models into sophisticated agents. However, debate persists regarding whether RL genuinely instills new skills within a base model or merely sharpens its existing distribution to elicit latent capabilities. To address this dichotomy, we present an explicit comparison between distribution sharpening and task-reward-based learning, utilizing RL as a tool to implement both paradigms. Our analysis reveals the inherent limitations of distribution sharpening, demonstrating from first principles how and why the optima can be unfavorable and the approach fundamentally unstable. Furthermore, our experiments using Llama-3.2-3B-Instruct, Qwen2.5-3B-Instruct and Qwen3-4B-Instruct-2507 on math datasets confirm that sharpening yields limited gains, whereas incorporating task-based reward signal can greatly help achieve robust performance improvements and stable learning.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!