2608.01522v1 Aug 02, 2026 cs.LG

질은 또 다른 질문을 낳는다: 강화 학습 기반 미세 조정에 대한 자기 진화형 교육 과정 - 경쟁 수학

Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics

Jianyou Wang
Jianyou Wang
Citations: 214
h-index: 5
Youze Zheng
Youze Zheng
UC San Diego
Citations: 14
h-index: 2
Longtian Bao
Longtian Bao
UC San Diego
Citations: 14
h-index: 2
R. Paturi
R. Paturi
Citations: 6,500
h-index: 30
Yang Zhang
Yang Zhang
Citations: 0
h-index: 0

언어 모델에게 아직 숙달하지 못한 기술을 가르치는 것은 세 가지 주요 어려움 때문에 방해받습니다. 첫째, 훈련 데이터가 부족합니다. 둘째, 정확한 추론 과정을 알 수 없는 경우가 많습니다. 셋째, 모델은 종종 개선이 더 이상 이루어지지 않는 성능 한계에 도달합니다. 우리는 이러한 어려움을 통제된 환경에서 연구하며, Qwen2.5-Math-7B 모델을 경쟁 수학(AIME) 문제에 대해 미세 조정했습니다. 이 모델은 처음에는 5.6%의 문제만 해결할 수 있었습니다(pass@1). 데이터 부족 문제를 해결하기 위해, 기존 문제를 활용하여 동일한 핵심 기술을 평가하는 다양한 변형 문제를 생성하는 확장 가능한 방법인 Question-begets-Question (QbQ)를 도입했습니다. 또한, 정답 추론 과정이 없는 상황을 모델링하기 위해, 문제 설명과 최종 답변만을 사용하여 강화 학습으로만 훈련했으며, 가이드(선생님)의 추론 과정을 사용하지 않았습니다. 그러나 이러한 방식으로 훈련했을 때 성능은 목표 수준에 훨씬 미치지 못했습니다. 실제 데이터와 합성 데이터를 결합하거나, QbQ를 통해 생성된 합성 데이터로 훈련했음에도 불구하고, pass@1은 각각 12.5% 및 14.5%로 제한되었습니다. 이러한 결과는 모델 자체의 고유한 한계 때문이 아니라는 것을 보여줍니다. 우리는 자기 진화형 교육 과정을 제안합니다. 이 과정은 각 단계에서 현재 모델의 성능을 평가하고, 모델이 비교적 잘 해결할 수 있는 문제들을 기반으로 QbQ를 실행하여 변형 문제를 생성하고, 이를 사용하여 훈련합니다. 동일한 데이터 예산 하에서도 이러한 방식은 성능 한계를 극복하고 pass@1을 16.5%까지 향상시켰으며, 20번의 반복 후에도 개선이 멈추는 기미가 보이지 않았습니다. 놀랍게도, 모델은 자신이 비교적 잘 해결할 수 있는 문제들의 변형된 버전으로 훈련될 때 성능이 향상되며, 이렇게 훈련된 모델은 훈련 중에 한 번도 본 적 없는 더 어려운 문제를 해결하는 능력을 갖추는 것을 확인했습니다.

Original Abstract

Teaching a language model a skill it has not mastered is obstructed by three recurring difficulties: training data is scarce, ground-truth reasoning traces are usually unavailable, and models often exhibit an apparent ceiling beyond which additional data yields no further improvement. We study these difficulties in a controlled setting, fine-tuning Qwen2.5-Math-7B on competition mathematics (AIME), a task on which it initially solves only 5.6\% of problems (pass@1). To address data scarcity, we introduce Question-begets-Question (QbQ), a scalable procedure in which a teacher transforms existing problems into diverse variants that probe the same underlying skills; to model the absence of oracle reasoning, we train exclusively via reinforcement learning on problem statements and final answers, never on teacher reasoning traces. Static training on such data, however, plateaus well short of the task: real-plus-synthetic augmentation and non-curriculum QbQ generated synthetic data training cap pass@1 at 12.5\% and 14.5\% respectively, despite large increases in data. Our central finding is that this ceiling is not intrinsic to the model. We propose a self-evolving curriculum that, each round, evaluates the current checkpoint, seeds QbQ from the problems it can mostly get right, and trains on the resulting variants; under an identical data budget, this breaks the ceiling and lifts pass@1 to 16.5\% with no sign of saturation after 20 rounds. Counterintuitively, we find that models improve when trained on variants of problems they can mostly get right, and that models trained this way go on to solve harder problems never seen during training.

0 Citations
0 Influential
15 Altmetric
75.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!