2606.31813v1 Jun 30, 2026 cs.LG

RLVR 환경에서의 저랭크 적응을 위한 기하학적 구조 보존 정규 직교 초기화 방법

Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in RLVR

Jiacheng Zhu
Jiacheng Zhu
Citations: 1,977
h-index: 24
Laixi Shi
Laixi Shi
Citations: 64
h-index: 3
Hanqing Zhu
Hanqing Zhu
Citations: 739
h-index: 15
Ruijia Zhang
Ruijia Zhang
Citations: 0
h-index: 0

저랭크 적응(LoRA) 및 그 변형은 지도 학습 기반 미세 조정(SFT) 패러다임 하에서 대규모 언어 모델의 효율적인 파라미터 업데이트를 가능하게 합니다. 그러나 강화 학습과 검증 가능한 보상(RLVR) 환경에서의 LoRA의 효과와 동작 방식은 아직 잘 알려져 있지 않습니다. 특히, SFT 환경에서 표준 LoRA보다 뛰어난 성능을 보이는 구조적으로 초기화된 LoRA 변형인 PiSSA 및 MiLoRA는 RLVR 환경에서는 오히려 표준 LoRA보다 성능이 낮아지거나 불안정한 학습을 보일 수 있습니다. 이러한 관찰 결과는 RLVR 환경에서 저랭크 행렬을 어떻게 초기화해야 하는지에 대한 명확한 지침이 부족하다는 것을 시사합니다. 본 연구에서는 RLVR 환경에서의 LoRA에 대한 이론적 분석을 수행하여, 정규 직교 초기화가 LoRA의 성능과 전체 미세 조정 간의 격차를 최소화함을 보여줍니다. 이러한 통찰력을 바탕으로, 우리는 RLVR 환경에서 저랭크 적응을 위한 기하학적 구조를 보존하는 정규 직교 초기화 방법을 제안하고, 이를 통해 RLPO 및 RLMO라는 두 가지 새로운 변형을 개발했습니다. 수학적 추론 벤치마크에서의 실험 결과는 제안된 정규 직교 초기화 방법이 RLVR 학습의 안정성을 향상시키고 표준 LoRA보다 우수한 성능을 보이며, PiSSA 및 MiLoRA와는 대조적으로 나타났음을 보여줍니다. 마지막으로, 본 연구에서 제시한 LoRA 초기화에 대한 통합적인 분석은 왜 PiSSA 및 MiLoRA가 RLVR 환경에서 성능이 저하되는지를 설명하며, 이는 독립적인 연구 가치를 지닐 수 있습니다. 코드 및 체크포인트는 다음 링크에서 공개적으로 이용 가능합니다: https://github.com/Richard-ZZZ/geometry-preserving-orthonormal-init-rlvr.

Original Abstract

Low-rank adaptation (LoRA) and its variants enable parameter-efficient fine-tuning of large language models under the supervised fine-tuning (SFT) paradigm. However, their efficacy and behavior under Reinforcement learning with verifiable rewards (RLVR) are less well understood. In particular, two structurally initialized LoRA variants, PiSSA and MiLoRA, which outperform standard LoRA under SFT, can underperform standard LoRA under RLVR and may even exhibit training instability. These observations suggest that how to initialize the low-rank matrices in RLVR remains unclear. In this work, we develop a theoretical analysis of LoRA in RLVR, showing that orthonormal initialization achieves the minimal gap between LoRA outcome and that of full fine-tuning. Guided by this insight, we propose geometry-preserving orthonormal initialization for low-rank adaptation in RLVR, leading to two new variants, RLPO and RLMO. Experiments on mathematical reasoning benchmarks show that the proposed orthonormal initialization stabilizes RLVR training and outperforms standard LoRA, contrasting with PiSSA and MiLoRA. Finally, our unified analysis for LoRA initialization also explains why PiSSA and MiLoRA can underperform in RLVR, which may be of independent interest. Code and checkpoints are publicly available at https://github.com/Richard-ZZZ/geometry-preserving-orthonormal-init-rlvr.

1 Citations
0 Influential
35.4657359028 Altmetric
6.9 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!