2601.07475v1 Jan 12, 2026 cs.LG

ARCQuant: 잔차 채널 증강을 통한 NVFP4 양자화 성능 향상 - LLM을 위한 방법

ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs

Xindian Ma
Xindian Ma
Citations: 285
h-index: 7
Peng Zhang
Peng Zhang
Citations: 36
h-index: 4
Haoqian Meng
Haoqian Meng
Citations: 15
h-index: 2
Yilun Luo
Yilun Luo
Citations: 15
h-index: 2
Yafei Zhao
Yafei Zhao
Citations: 10
h-index: 2
Wenyuan Liu
Wenyuan Liu
Citations: 32
h-index: 3

NVFP4와 같은 정밀한 수치 형식이 등장하면서, 효율적인 대규모 언어 모델(LLM) 추론에 새로운 기회를 제공합니다. 그러나 기존의 양자화(PTQ) 전략을 이러한 형식에 적용하는 것은 어렵습니다. 회전 기반 방법은 정밀한 블록 분리를 저해하며, 스무딩 기법은 상당한 4비트 양자화 오류에 어려움을 겪고, 혼합 정밀도 방식은 종종 통일된 정밀도 연산에 대한 하드웨어 제약과 충돌합니다. 이러한 문제점을 해결하기 위해, 잔차 채널 증강을 통해 NVFP4 성능을 향상시키는 프레임워크인 ARCQuant를 제안합니다. ARCQuant는 블록 분리나 하드웨어 균일성을 저해하는 기존 방법과 달리, 활성화 행렬에 양자화된 잔차 채널을 추가하여 엄격하게 통일된 NVFP4 형식을 유지합니다. 이러한 설계는 오류 보정 과정을 행렬 축소 차원에 직접 통합하여, 최소한의 오버헤드로 표준화된, 고도로 최적화된 GEMM 커널을 사용할 수 있도록 합니다. 이론적 분석 결과, ARCQuant의 2단계 NVFP4 양자화 방식의 최악의 오류 경계는 MXFP8과 같은 표준 8비트 형식과 유사합니다. LLaMA 및 Qwen 모델에 대한 광범위한 실험 결과, ARCQuant는 최첨단 정확도를 달성하며, 퍼플렉서티 및 다운스트림 작업에서 전체 정밀도 기준과 유사한 성능을 보입니다. 또한, RTX 5090 및 RTX PRO 6000 GPU에서의 배포 테스트를 통해 실질적인 이점을 확인했으며, FP16 대비 최대 3배의 속도 향상을 달성했습니다. 저희의 코드는 https://github.com/actypedef/ARCQuant 에서 확인하실 수 있습니다.

Original Abstract

The emergence of fine-grained numerical formats like NVFP4 presents new opportunities for efficient Large Language Model (LLM) inference. However, it is difficult to adapt existing Post-Training Quantization (PTQ) strategies to these formats: rotation-based methods compromise fine-grained block isolation; smoothing techniques struggle with significant 4-bit quantization errors; and mixed-precision approaches often conflict with hardware constraints on unified-precision computation. To address these challenges, we propose ARCQuant, a framework that boosts NVFP4 performance via Augmented Residual Channels. Distinct from methods that compromise block isolation or hardware uniformity, ARCQuant maintains a strictly unified NVFP4 format by augmenting the activation matrix with quantized residual channels. This design integrates the error compensation process directly into the matrix reduction dimension, enabling the use of standard, highly optimized GEMM kernels with minimal overhead. Theoretical analysis confirms that the worst-case error bound of our dual-stage NVFP4 quantization is comparable to that of standard 8-bit formats such as MXFP8. Extensive experiments on LLaMA and Qwen models demonstrate that ARCQuant achieves state-of-the-art accuracy, comparable to full-precision baselines in perplexity and downstream tasks. Furthermore, deployment on RTX 5090 and RTX PRO 6000 GPUs confirms practical benefits, achieving up to 3x speedup over FP16. Our code is available at https://github.com/actypedef/ARCQuant .

8 Citations
0 Influential
37.951858789481 Altmetric
22.0 Score
Original PDF
17

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!