2605.28421v1 May 27, 2026 cs.AI

DenoiseRL: 약한 모델의 오류로부터 회복하기 위한 추론 모델 학습 방법

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes

Changyi Xiao
Changyi Xiao
Citations: 43
h-index: 3
Yixin Cao
Yixin Cao
Citations: 59
h-index: 4
Caijun Xu
Caijun Xu
Citations: 4
h-index: 1
Z. Peng
Z. Peng
Citations: 643
h-index: 10

강화 학습은 거대 언어 모델의 추론 능력을 향상시키는 핵심 패러다임으로 자리 잡았지만, 대부분의 기존 방법은 여전히 성능이 더 뛰어난 가이드 모델이나 엄선된 어려운 데이터 세트에 의존하여 확장 가능한 개선을 제한합니다. 본 논문에서는 DenoiseRL이라는 강화 학습 프레임워크를 소개합니다. DenoiseRL은 외부 감독 대신 약한 모델에서 발생하는 오류로부터 회복에 중점을 둔 최적화 방식을 사용합니다. DenoiseRL은 더 강력한 감독이나 정교하게 설계된 데이터에 의존하는 대신, 잘못된 추론 과정을 개선의 기회로 전환하여 직접 학습하므로, 학습을 더욱 확장 가능하고 외부 리소스에 대한 의존성을 줄입니다. 이를 통해 풍부하고 다양한 학습 신호를 얻어 불완전한 모델 동작으로부터 탐색 효율성을 향상시킵니다. 그 결과, DenoiseRL은 추론 성능과 전체적인 학습 효율성을 향상시키면서 동시에 비용이 많이 드는 데이터 큐레이션이나 더 강력한 가이드 모델의 필요성을 줄입니다. 실험적으로, DenoiseRL은 경쟁력 있는 수학 및 일반 추론 벤치마크에서 강력한 온라인 강화 학습 기준 모델보다 일관되게 우수한 성능을 보이며, 학습 난이도가 증가함에 따라 더욱 강력한 자체 수정 기능을 촉진합니다. 이는 거대 언어 모델의 추론 능력을 향상시키는 효과적이고 확장 가능한 대안적인 방법을 제시합니다.

Original Abstract

Reinforcement learning has become a central paradigm for advancing reasoning in large language models, yet most existing methods still depend on stronger teacher models or heavily curated difficult datasets, limiting scalable capability improvement. In this paper, we introduce DenoiseRL, a reinforcement learning framework that substitutes external supervision with recovery-oriented optimization over failures from weak models. Instead of relying on stronger supervision or carefully engineered data, DenoiseRL learns directly from incorrect reasoning traces by converting them into opportunities for improvement, making training more scalable and less dependent on external resources. This yields a richer and more diverse learning signal, improving exploration efficiency from imperfect model behavior. As a result, DenoiseRL improves reasoning performance and overall training efficiency while reducing the need for expensive data curation or stronger teacher models. Empirically, DenoiseRL consistently outperforms strong on-policy RL baselines across competitive mathematical and general reasoning benchmarks and promotes stronger self-corrective behavior as training difficulty increases, highlighting an effective and scalable alternative pathway for improving reasoning in large language models.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!