2608.01847v1 Aug 03, 2026 cs.AI

FOCUS: 결합된 이완 및 다중-그레인 스케일링을 통한 FP4 최적화

FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling

Guanghua Yu
Guanghua Yu
Citations: 50
h-index: 4
Hong Liu
Hong Liu
Citations: 36
h-index: 2
Xianglong Yan
Xianglong Yan
Citations: 86
h-index: 5
Chengzhu Bao
Chengzhu Bao
Citations: 4
h-index: 1
Tianao Zhang
Tianao Zhang
Citations: 53
h-index: 4
Yulun Zhang
Yulun Zhang
Citations: 67
h-index: 5
Jianchen Zhu
Jianchen Zhu
Citations: 479
h-index: 7

대규모 언어 모델(LLM)은 뛰어난 성능을 보이지만, 막대한 크기로 인해 배포 비용이 매우 높습니다. MXFP4 및 NVFP4와 같은 형식을 사용하는 FP4 양자화는 최신 가속기에 내장된 하드웨어 지원 기능을 제공하여 매력적인 솔루션을 제시합니다. 그러나 FP4 정밀도에서 정확도를 유지하는 것은 여전히 어렵습니다. 주요 병목 현상은 스케일 최적화에 있습니다. 기존 방법은 양자화 및 역양자화 스케일을 강하게 결합시켜 하드웨어 요구 사항(예: MXFP4의 경우 E8M0)과 같은 이산 저정밀 형식에 맞춰야 합니다. 그러나 양자화 스케일은 저장되지 않으며 이러한 제약을 따를 필요가 없으므로 상당한 잠재적 최적화 공간이 존재합니다. 본 연구에서는 결합된 이완 및 다중-그레인 스케일링을 통한 FP4 최적화를 위한 엔드투엔드 스케일 학습 기능을 갖춘 양자화 프레임워크인 FOCUS를 제안합니다. 결합된 이완 스케일링(CRS)은 학습 가능한 고정밀 계수를 사용하여 양자화 및 역양자화 스케일 간의 강한 결합을 완화하여 하드웨어 호환성을 유지하면서 더욱 효과적인 최적화를 가능하게 합니다. 다중-그레인 스케일링(DGS)은 더 미세한 서브 블록 수준에서 양자화 스케일을 추가로 조정하여 로컬 가중치 분포에 대한 보다 정확한 적응을 허용합니다. 여러 LLM 패밀리 및 벤치마크에서의 실험 결과, FOCUS는 MXFP4 및 NVFP4 형식 모두에서 최첨단 FP4 정확도를 달성하며 추가적인 추론 오버헤드를 발생시키지 않습니다. 코드 및 양자화된 모델은 https://github.com/tencent/AngelSlim 에서 공개될 예정입니다.

Original Abstract

Large language models (LLMs) achieve remarkable performance but are expensive to deploy due to their enormous size. FP4 quantization, with formats such as MXFP4 and NVFP4, offers an appealing solution with native hardware support on modern accelerators. However, maintaining accuracy under FP4 precision remains difficult. A key bottleneck lies in scale optimization: existing methods tightly couple the quantization and dequantization scales, forcing both to conform to the discrete low-precision format required by hardware, such as E8M0 in MXFP4. Yet the quantization scale is never stored and need not obey this constraint, suggesting a significant untapped optimization space. In this work, we propose FOCUS, a post-training quantization framework with end-to-end scale learning for FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling. Coupled-Relaxation Scaling (CRS) relaxes the tight coupling between quantization and dequantization scales with a learnable full-precision coefficient, enabling more effective optimization without breaking hardware compliance. Dual-Granularity Scaling (DGS) further refines the quantization scale at a finer sub-block granularity, allowing more precise adaptation to local weight distributions. Experiments across multiple LLM families and benchmarks show that FOCUS achieves state-of-the-art FP4 accuracy under both MXFP4 and NVFP4 formats, while introducing no additional inference overhead. Code and quantized models will be released at https://github.com/tencent/AngelSlim.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!