대규모 언어 모델의 양자화로 인한 성능 저하: 신호-잡음 관점
Quantization Degradation in Large Language Models: A Signal-Noise Perspective
양자화 후 학습은 대규모 언어 모델의 배포 비용을 줄이지만, 양자화된 모델의 성능 저하 정도는 비트 폭(bit-width)만으로 결정되지 않습니다. 본 연구에서는 다양한 비트 폭, 양자화 방법, 모델 크기 및 downstream 작업에 대해 여러 모델 패밀리를 대상으로 가중치 전용 양자화 후 학습을 체계적으로 분석했습니다. 그 결과, 이러한 성능 저하는 다양한 요인에 따라 크게 달라지는 것을 확인했습니다. 4비트 양자화는 일반적으로 성능 저하를 최소화하지만, 2비트는 광범위한 성능 저하를 유발하며, 3비트에서는 성능 저하가 나타나지만 작업 유형, 양자화 방법 및 모델 크기에 따라 현저하게 달라집니다. 이러한 변동성을 설명하기 위해, 본 연구에서는 신호 대 잡음 비율(SNR)을 사용하여 양자화가 전체 정밀도 표현에 얼마나 큰 영향을 미치는지 측정했습니다. 성능 저하의 원인을 분석한 결과, 개별 모듈 내에서 발생하는 양자화 오차와 이 오차가 레이어를 거치면서 어떻게 누적되는지에 대한 두 가지 상호 관련된 과정이 밝혀졌습니다. 첫째, SNR 분해를 통해 새로운 오차가 발생할 확률은 가중치 오차의 크기, 작업 관련 신호의 강도 및 양자화 오차가 작업 관련 활성화 값과 얼마나 일치하는지에 따라 달라짐을 확인했습니다. 이러한 요인들은 각각 다른 방식으로 구성 요소에 영향을 미칩니다. 둘째, 레이어 간 전파 분석 결과, 오차는 레이어를 통과하면서 감쇠되거나 유지되거나 증폭될 수 있으며, 더 큰 모델은 오차 증폭이 약하여 성능 저하가 덜 발생하는 것을 확인했습니다. 종합적으로 볼 때, 본 연구의 결과는 양자화로 인한 성능 저하가 오류가 어디에서 발생하고 네트워크 전체를 통해 어떻게 누적되는지에 따라 결정된다는 것을 보여줍니다.
Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales and downstream tasks on multiple model families. We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation, and at 3-bit, degradation becomes apparent but varies markedly with task type, quantization method and model scale. To explain this variability, we use the signal-to-noise ratio (SNR) to measure how strongly quantization perturbs full-precision representations. We trace degradation back to two linked processes: how quantization errors arise within individual modules, and how they accumulate across layers. First, a source SNR decomposition shows that newly introduced errors depend on three factors: the magnitude of the weight error, the strength of the task-specific signal, and how strongly the quantization error aligns with task-specific activations. Different factors affect these components in distinct ways. Second, a cross-layer propagation analysis shows that these errors can be attenuated, preserved, or amplified as they pass across layers, and that larger models benefit from weaker error amplification. Together, these results establish that quantization degradation is governed by how errors are introduced at the source and how they accumulate across the network.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.