2606.10531v1 Jun 09, 2026 cs.CL

LC-QAT: 선형 제약 벡터 양자화를 이용한 데이터 효율적인 2비트 양자화 인식 학습을 통한 거대 언어 모델 (LLM) 최적화

LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization

Haiyan Zhao
Haiyan Zhao
Citations: 18
h-index: 2
Fengxiang Wang
Fengxiang Wang
Citations: 37
h-index: 2
Xu Han
Xu Han
Citations: 106
h-index: 3
Haoyu Wang
Haoyu Wang
Citations: 21
h-index: 3
Xingyu Yu
Xingyu Yu
Citations: 11
h-index: 1

양자화 인식 학습(QAT)은 극도로 낮은 비트의 거대 언어 모델(LLM)에 필수적입니다. 현재의 QAT 방법들은 주로 스칼라 양자화(SQ)를 기반으로 하며, 이는 효율적인 최적화를 가능하게 하지만 2비트 정밀도에서 심각한 성능 저하를 초래합니다. 반면, 벡터 양자화(VQ)는 훨씬 높은 표현 능력을 제공하지만, 이산 코드북 검색은 엔드-투-엔드 학습을 방해합니다. 우리는 LC-QAT라는 2비트 가중치 기반의 VQ-QAT 프레임워크를 제안합니다. LC-QAT는 양자화된 가중치를 이산 벡터에 대한 학습 가능한 선형 매핑으로 표현하여 고품질의 PTQ 초기화를 제공하며, 학습 과정에서 명시적인 코드북 검색 없이 완전하게 미분 가능한 엔드-투-엔드 최적화를 가능하게 합니다. 이러한 강력한 사후 훈련 초기화 덕분에 LC-QAT는 매우 데이터 효율적입니다. 다양한 LLM에 대한 실험 결과, LC-QAT는 최첨단 QAT 방법보다 일관되게 우수한 성능을 보이며, 학습 데이터의 0.1%에서 10%만을 사용합니다. 이러한 결과는 LC-QAT를 극도로 낮은 비트 모델 배포를 위한 실용적이고 확장 가능한 솔루션으로 확립합니다.

Original Abstract

Quantization-aware training (QAT) is essential for extremely low-bit large language models (LLMs). Current QAT methods are mainly based on scalar quantization (SQ), which enables efficient optimization but suffers from severe performance degradation at 2-bit precision. On the other hand, vector quantization (VQ) provides substantially higher representational capacity, but its discrete codebook lookup prevents end-to-end training. We propose LC-QAT, a 2-bit weight-only VQ-QAT framework that represents quantized weights via a learned affine mapping over discrete vectors, which yields a high-quality PTQ initialization and enables fully differentiable end-to-end optimization without explicit codebook lookup in the training forward pass. This strong post-training initialization makes LC-QAT highly data-efficient. Experiments across diverse LLMs demonstrate that LC-QAT consistently outperforms state-of-the-art QAT methods while using only 0.1%--10% of the training data. Our results establish LC-QAT as a practical and scalable solution for extreme low-bit model deployment.

1 Citations
0 Influential
1.5 Altmetric
8.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!