2602.14868v1 Feb 16, 2026 cs.LG

골디락스 강화학습: 추론을 위한 희소 보상을 극복하기 위한 작업 난이도 조정

Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning

Ilia Mahrooghi
Ilia Mahrooghi
Citations: 0
h-index: 0
Aryo Lotfi
Aryo Lotfi
Citations: 194
h-index: 6
Emmanuel Abbe
Emmanuel Abbe
Citations: 134
h-index: 4

강화학습은 대규모 언어 모델에서 추론 능력을 발휘하는 강력한 패러다임으로 부상했습니다. 그러나 희소 보상에 의존하는 것은 이 과정을 매우 샘플 비효율적으로 만듭니다. 모델은 최소한의 피드백으로 방대한 탐색 공간을 탐색해야 하기 때문입니다. 기존의 커리큘럼 학습은 복잡성에 따라 데이터를 정렬하여 이를 완화하려고 하지만, 특정 모델에 대한 적절한 정렬 순서는 종종 불분명합니다. 이러한 문제를 해결하기 위해, 우리는 '골디락스(Goldilocks)'라는 새로운 교사 기반 데이터 샘플링 전략을 제안합니다. 이 전략은 학생 모델의 각 질문의 난이도를 예측하는 것을 목표로 합니다. 교사 모델은 학생 모델에 적절한 난이도의 질문을 선택합니다. 즉, 너무 쉽거나 너무 어려운 질문이 아닌, '골디락스' 원칙에 따른 질문을 선택합니다. 동시에, 학생 모델은 GRPO(Guided Reinforcement Proximal Optimization)를 사용하여 학습합니다. 교사 모델은 학생 모델이 학습한 샘플에 대한 학생 모델의 성능을 활용하여 학생 모델의 변화하는 능력에 지속적으로 적응합니다. OpenMathReasoning 데이터셋에서, 골디락스 데이터 샘플링은 동일한 컴퓨팅 예산 하에서 표준 GRPO로 학습된 모델의 성능을 향상시킵니다.

Original Abstract

Reinforcement learning has emerged as a powerful paradigm for unlocking reasoning capabilities in large language models. However, relying on sparse rewards makes this process highly sample-inefficient, as models must navigate vast search spaces with minimal feedback. While classic curriculum learning aims to mitigate this by ordering data based on complexity, the right ordering for a specific model is often unclear. To address this, we propose Goldilocks, a novel teacher-driven data sampling strategy that aims to predict each question's difficulty for the student model. The teacher model selects questions of appropriate difficulty for the student model, i.e., questions that are neither too easy nor too hard (Goldilocks principle), while training the student with GRPO. By leveraging the student's performance on seen samples, the teacher continuously adapts to the student's evolving abilities. On OpenMathReasoning dataset, Goldilocks data sampling improves the performance of models trained with standard GRPO under the same compute budget.

1 Citations
0 Influential
3 Altmetric
16.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!