2607.14506v1 Jul 16, 2026 cs.LG

검증 가능한 보상을 활용한 강화 학습의 일반화 경계: 비공허성을 갖는 일반화 경계

Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards

Yuxuan Zhu
Yuxuan Zhu
Citations: 247
h-index: 7
Daniel Kang
Daniel Kang
Citations: 251
h-index: 7
Rohan Alur
Rohan Alur
Citations: 124
h-index: 6

검증 가능한 보상을 활용한 강화 학습(RLVR)은 대규모 언어 모델(LLM)의 추론 능력을 향상시키는 데 널리 사용되지만, 결과 모델의 일반화 성능에 대한 이해는 여전히 부족합니다. 본 연구에서는 수십억 개의 파라미터를 가진 모델에서 효율적인 파라미터 기반 RLVR 미세 조정에 대한 최초의 비공허성을 갖는 일반화 경계를 제시합니다. 우리는 PAC-Bayes 압축 경계를 이 설정에 적용하고, 토큰 생성 과정의 고유한 확률적 특성을 해결하기 위해 Gumbel-max 재파라미터화 기법을 사용했습니다. 이러한 경계를 실질적으로 구현하기 위해, 우리는 RLVR과 온라인 정책 증류, TinyLoRA 및 모델 양자화를 통합한 Progressive RLVR 프레임워크를 제안합니다. 실험 결과, Progressive RLVR은 표준 LoRA 미세 조정 성능의 84-97%를 유지하면서 압축률이 14,796배 더 높은 모델을 생성하는 것으로 나타났습니다. 본 연구는 수학 문제 해결, 프로그래밍, 일반 지식 추론 및 텍스트-SQL 변환이라는 네 가지 영역에서 비공허성을 갖는 일반화 경계를 제공합니다. 이러한 경계는 기본 모델의 정확도를 9-51% 이상으로 개선하며, 미세 조정된 모델의 정확도와 6-11% 이내의 오차를 보입니다.

Original Abstract

While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (LLMs), the generalizability of the resulting models remains poorly understood. In this work, we establish the first non-vacuous generalization bounds for parameter-efficient RLVR fine-tuning at the billion-parameter scale. Our approach adapts PAC-Bayes compression bounds to this setting, and addresses the inherent stochasticity of token generation by applying the Gumbel-max reparameterization trick. To operationalize these bounds, we propose the Progressive RLVR framework, which integrates RLVR with on-policy distillation, TinyLoRA, and model quantization. Progressive RLVR empirically retains 84-97% performance of standard LoRA fine-tuning while producing models that are 14,796x more compressible. We show that this framework yields non-vacuous generalization bounds in four domains: mathematical problem-solving, programming, general-knowledge reasoning, and Text-to-SQL. Our bounds exceed the accuracy of the base model by 9-51% and lie within 6-11% of the accuracy of the fine-tuned models.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!