2607.26515v1 Jul 29, 2026 cs.LG

대규모 언어 모델의 강화 학습 후속 교육을 위한 HiFloat4 형식: 엔드 투 엔드 방식

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

Guipeng Hu
Guipeng Hu
Citations: 19
h-index: 3
Yunke Peng
Yunke Peng
Citations: 6
h-index: 1
Tianchi Hu
Tianchi Hu
Citations: 19
h-index: 3
Junsong Wang
Junsong Wang
Citations: 19
h-index: 3
H. Le
H. Le
Citations: 7
h-index: 2
Yaoyuan Wang
Yaoyuan Wang
Citations: 61
h-index: 4
Hei Yi Mak
Hei Yi Mak
Citations: 12
h-index: 2
Shadan Golestan
Shadan Golestan
University of Alberta
Citations: 153
h-index: 8
Mehran Taghian Jazi
Mehran Taghian Jazi
Citations: 0
h-index: 0
Fengcheng He
Fengcheng He
Citations: 0
h-index: 0
Tanzila Rahman
Tanzila Rahman
Citations: 645
h-index: 8
Anandharaju Durai Raju
Anandharaju Durai Raju
Simon Fraser University, Canada
Citations: 76
h-index: 3

본 논문에서는, 지금까지 알려진 바로는 처음으로 FP4 정밀도를 사용하는 전체적인 강화 학습 후속 교육 방법을 제시합니다. 이 방법에서 롤아웃(rollout) 정책과 학습 정책 모두, 그리고 이들의 순방향 및 역방향 연산이 모두 4비트 정밀도로 작동됩니다. 체계적인 연구 결과, FP4 강화 학습에서의 성능 저하의 주요 원인은 학습 측면의 양자화 오류가 아니라 롤아웃 활성화 값의 양자화라는 것을 밝혀냈습니다. 직관에 반하여, 학습 정책을 더 높은 정밀도로 복원하면서 롤아웃은 FP4로 유지하면 정확도가 전체 FP4 기준보다 낮아집니다. 이는 롤아웃과 학습 간의 불일치가 주된 실패 원인임을 보여주며, 일반적인 사전 학습 방식의 해결책이 효과적이지 않음을 시사합니다. 우리는 이러한 문제를 '롤아웃 잔차 양자화(Rollout-ResQ)'를 통해 해결합니다. Rollout-ResQ는 하드웨어 친화적인 희소 패턴으로 제한된 단일 잔차 보정 항을 사용하여, 롤아웃 행렬 곱셈에만 적용됩니다. 이는 경량적인 수정 방법이며, 이상치로 인한 언더플로우(underflow)로 발생하는 대부분의 정밀도 손실을 복구하면서도 롤아웃의 연산량을 크게 늘리지 않습니다. Qwen2.5-3B 및 Qwen2.5-Math-7B 모델에서, Rollout-ResQ와 HiFloat4 (HiF4) 형식을 함께 사용하면 정확도 격차가 BF16 기준 4.9%에서 1.1%로 줄어들어, 완전히 양자화된 FP4 강화 학습이 전체 정밀도에 매우 근접하게 됩니다. 동일한 방법을 오픈 표준인 MXFP4에 적용했을 때에도, 정확도 격차가 13.6%에서 5.3%로 감소하며, 이는 FP4 형식 선택이 복구 가능한 정확도의 상한을 결정하는 중요한 요소임을 보여줍니다. 종합적으로 볼 때, 이러한 결과는 HiF4를 엔드 투 엔드 FP4 강화 학습 후속 교육을 가능하게 하는 핵심 형식으로, 그리고 Rollout-ResQ를 BF16과의 격차를 좁힐 수 있게 하는 활성화 측면의 메커니즘으로 확립합니다.

Original Abstract

We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. A systematic study reveals that the dominant source of degradation in FP4 RL is not training-side quantization error but rollout activation quantization: outliers stretch the dynamic range so far that a large number of activation values underflow to zero under FP4. Counterintuitively, restoring the training policy to higher precision while keeping the rollout in FP4 makes accuracy worse than full FP4 baseline, exposing rollout-training mismatch as the principal failure mode and ruling out standard pretraining-style fixes. We address this with Rollout Residual Quantization (Rollout-ResQ): a single residual correction term constrained to a hardware-friendly sparsity pattern, added only to the FP4 rollout matmul -- a lightweight correction that recovers most of the precision lost to outlier-driven underflow without inflating the rollout's compute footprint. On Qwen2.5-3B and Qwen2.5-Math-7B, Rollout-ResQ paired with the HiFloat4 (HiF4) format -- whose three-level hierarchical scaling preserves resolution under FP4's tight 4-bit budget -- closes the accuracy gap to BF16 from 4.9% to 1.1%, bringing fully quantized FP4 RL within striking distance of full precision. Applied to the open-standard MXFP4, the same recipe narrows the gap from 13.6% to 5.3%, revealing that FP4 format choice is a key factor that determines the ceiling on recoverable accuracy. Together, these results establish HiF4 as the enabling format for end-to-end FP4 RL post-training, and Rollout-ResQ as the activation-side mechanism that makes the gap to BF16 closable.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!