2606.20381v1 Jun 18, 2026 cs.AI

LLM FP4 사전 학습에서의 축소 편향 재고: 기하학적 원인, 체계적인 영향 및 UFP4 레시피

Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

Mingliang Gong
Mingliang Gong
Citations: 120
h-index: 2
Zhonghui Jiang
Zhonghui Jiang
Citations: 112
h-index: 4
Changxin Tian
Changxin Tian
Citations: 181
h-index: 6
Kunlong Chen
Kunlong Chen
Citations: 92
h-index: 4
Ziqi Liu
Ziqi Liu
Citations: 82
h-index: 3
Jun Zhou
Jun Zhou
Citations: 135
h-index: 5
Peijie Jiang
Peijie Jiang
Citations: 80
h-index: 5
Zhiqiang Zhang
Zhiqiang Zhang
Citations: 72
h-index: 5
Qian Zhao
Qian Zhao
Citations: 135
h-index: 4
Haitao Zhang
Haitao Zhang
Citations: 0
h-index: 0
Chaofan Yu
Chaofan Yu
Citations: 297
h-index: 7
Jia Liu
Jia Liu
Citations: 46
h-index: 2

FP4 학습은 LLM 사전 학습의 메모리 및 계산 비용을 크게 줄일 수 있는 잠재력을 가지고 있지만, 현재 NVIDIA Blackwell/Rubin 클래스 시스템 및 AMD MI350 시리즈 GPU를 포함한 FP4 하드웨어 경로 및 레시피는 여전히 E2M1 데이터 요소에 집중되어 있습니다. 본 연구에서는 이러한 선택의 근본적인 한계를 밝히고 있습니다. 비균일 형식인 E2M1은 표현 가능한 빈(bin)의 기하학적 비대칭으로 인해 발생하는 체계적인 음수 반올림 오류인 축소 편향(Shrinkage Bias)을 inherent하게 갖습니다. 우리는 이 편향이 레이어를 거치면서 곱셈적으로 누적되고, Random Hadamard Transform (RHT)에 의해 증폭되어 기존 E2M1 기반 FP4 레시피에서 관찰되는 학습 불안정성을 통합적으로 설명한다고 보여줍니다. 반면, 균일 그리드(E1M2/INT4)는 이러한 그리드 기하학 오류를 우회하며 RHT로 인한 향상된 버킷 활용률을 더 높은 양자화 품질로 전환하는 데 도움이 됩니다. 이 발견에 기반하여, 우리는 모든 세 가지 학습 GEMM에 RHT를 적용하고, 동적 반올림을 dY에만 제한하는 균일 4비트 학습 레시피인 UFP4를 제안합니다. Dense 1.5B, MoE 7.9B 및 MoE 124B 모델의 장기 사전 학습에서, UFP4는 강력한 E2M1 기반 기준 성능보다 BF16 대비 손실 저하 측면에서 일관되게 우수한 결과를 보여주었으며, 이는 스케일링 법칙 분석 및 ablation 연구를 통해 뒷받침됩니다. 우리의 결과는 향후 가속기가 E2M1과 함께 E1M2/INT4 스타일의 균일 4비트 그리드를 주요 학습 원시 자료로 지원해야 함을 시사합니다.

Original Abstract

FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hardware paths and recipes, including NVIDIA Blackwell/Rubin-class systems and AMD MI350-series GPUs, remain centered on E2M1 data elements. In this study, we identify a fundamental limitation of that choice: non-uniform formats such as E2M1 inherently suffer from Shrinkage Bias, a systematic negative rounding error caused by the geometric asymmetry of their representable bins. We show that this bias accumulates multiplicatively across layers and is amplified by the Random Hadamard Transform (RHT), providing a unified explanation for the training instability observed in existing E2M1-based FP4 recipes. In contrast, uniform grids (E1M2/INT4) bypass this grid-geometry error and better convert the improved bucket utilization from RHT into higher quantization quality. Based on this finding, we propose UFP4, a uniform 4-bit training recipe that applies RHT to all three training GEMMs while restricting stochastic rounding to dY alone. On Dense 1.5B, MoE 7.9B, and MoE 124B long-run pretraining, UFP4 consistently achieves lower BF16-relative loss degradation than strong E2M1-based baselines, supported by scaling-law analysis and ablation studies. Our results suggest that future accelerators should support E1M2/INT4-style uniform 4-bit grids as first-class training primitives alongside E2M1.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!