2607.24953v1 Jul 27, 2026 cs.LG

전위(Transposition)-불변 블록 양자화를 통한 안정적인 FP4 학습

Stable FP4 Training via Transposition-Invariant Block Quantization

Yufei Cui
Yufei Cui
Citations: 51
h-index: 4
Yunke Peng
Yunke Peng
Citations: 6
h-index: 1
Yaoyuan Wang
Yaoyuan Wang
Citations: 61
h-index: 4
Zhijun Tu
Zhijun Tu
Citations: 260
h-index: 8
Mehran Taghian Jazi
Mehran Taghian Jazi
Citations: 0
h-index: 0
Mehdi Rahimifar
Mehdi Rahimifar
Citations: 0
h-index: 0
Amin Darabi
Amin Darabi
Citations: 26
h-index: 3
Xing Huang
Xing Huang
Citations: 18
h-index: 1
Hongliang Li
Hongliang Li
Citations: 0
h-index: 0

대규모 언어 모델(LLM) 학습의 효율성을 향상시키는 핵심 방법은 학습 정밀도를 낮추는 것입니다. 하지만 FP8을 넘어 4비트 부동소수점(FP4)으로 확장하는 것은 최적화 과정에서의 불안정성 때문에 여전히 어려운 과제입니다. 본 연구에서는 기존의 미세 스케일링 접근 방식에서 발생하는 불안정성의 근본적인 원인을 밝혀냈습니다. 이는 텐서 전위로 인해 발생하는 스케일 불일치 때문입니다. 기존의 1차원 블록 양자화 방법에서는 순전파 및 역전파 과정에서 동일한 값에 대해 다른 스케일링 계수가 적용되어, 전위 연산 후에는 편향되고 불안정한 기울기 업데이트가 발생합니다. 이 문제를 해결하기 위해, 본 연구에서는 2차원 블록 FP4 양자화를 기반으로 하는 저정밀 학습 프레임워크를 제안합니다. 이는 전위 불변 스케일링을 적용하여 순전파 및 역전파 계산 간의 일관성을 유지합니다. 또한, 양자화 오차를 제어하고 편향되지 않은 기울기를 유지하기 위해 잘라내기 없는 스케일링과 확률적 반올림을 결합했습니다. 어텐션 메커니즘의 민감성을 고려하여, 쿼리 및 키 투영에 MXFP8 양자화를 적용하여 실용적인 혼합 정밀도 설계를 구현했습니다. 본 연구에서는 최대 70억 개의 파라미터를 가진 밀집 LLM과 300억 개의 파라미터를 가진 Mixture-of-Experts 모델을 사용하여 제안하는 방법을 평가했습니다. 이 모델들은 최대 1000억 개의 토큰으로 학습되었습니다. 모든 설정에서, 본 연구의 접근 방식은 안정적인 엔드투엔드 FP4 학습을 달성했으며, 퍼플렉시티 및 다운스트림 정확도 측면에서 BF16 성능과 거의 동일한 결과를 보였습니다 (1.3% 미만의 성능 저하). 이러한 결과는 순전파-역전파 스케일링 일관성을 적용하는 것이 대규모 FP4 학습을 가능하게 하는 데 충분하며, 보다 효율적인 LLM 학습을 위한 간단하고 효과적인 방법을 제공한다는 것을 보여줍니다.

Original Abstract

Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating point (FP4) remains challenging due to instability during optimization. We identify a fundamental source of this instability in existing microscaling approaches: scale inconsistency induced by tensor transposition. In conventional 1D block quantization, forward and backward passes assign di erent scaling factors to the same values after transposition, leading to biased and unstable gradient updates. To address this issue, we propose a low-precision training framework based on 2D block FP4 quantization, which enforces transposition-invariant scaling and preserves consistency between forward and backward computations. We further combine this with truncation-free scaling and stochastic rounding to control quantization error and maintain unbiased gradients. To handle the sensitivity of attention mechanisms, we adopt MXFP8 quantization for query and key projections, yielding a practical mixed-precision design. We evaluate our method on dense LLMs up to 7B parameters and a 30B Mixture-of-Experts model, trained on up to 100B tokens. Across all settings, our approach achieves stable end-to-end FP4 training and closely matches BF16 performance, with less than 1.3% degradation in perplexity and downstream accuracy. These results demonstrate that enforcing forwardbackward scaling consistency is su cient to enable practical FP4 training at scale, providing a simple and e ective pathway toward more e cient LLM training.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!