2607.08734v1 Jul 09, 2026 cs.AI

동등성의 환상: LLM에서의 양자화 효과에 대한 통계적 분석

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

C. Akcora
C. Akcora
Citations: 1,559
h-index: 22
Baha Rababah
Baha Rababah
Citations: 127
h-index: 6
C. K. Leung
C. K. Leung
Citations: 331
h-index: 8

사후 학습 양자화는 자원 제약 환경에서 대규모 언어 모델을 배포하는 데 널리 사용되지만, 그 평가는 거의 정확도와 퍼플렉시티 지표에만 의존합니다. 본 연구에서는 이러한 지표들이 양자화로 인해 발생하는 행동 변화를 제대로 반영하지 못함을 보여줍니다. 우리는 '정확성 일치도(correctness agreement)'라는 새로운 평가 지표를 제시하는데, 이는 기준 모델과 양자화된 모델 간의 올바른 예측 결과의 중첩 정도를 측정하며, 절대적인 정확도와는 독립적입니다. 8비트에서 2비트에 이르는 다양한 모델 및 양자화 방식을 사용하여 분석한 결과, 작업 성능이 유지되는 것처럼 보이는 경우에도 중간 수준의 양자화에서 행동적 차이가 나타나는 것을 확인했습니다. 이러한 현상을 설명하기 위해, 우리는 어텐션 가중치에 대한 구조 연산으로서의 양자화를 분석하고, 통계 및 분포 기반 지표를 사용하여 계층별 왜곡을 정량화했습니다. 연구 결과는 낮은 비트 폭에서 비선형적인 경향성을 보여주며, 쿼리(query) 및 키(key) 투영은 값(value) 및 출력(output) 투영보다 일관되게 더 민감한 것을 나타냅니다. 이러한 발견은 기준 모델과 양자화된 모델 간의 동등성에 대한 착시를 드러내며, 기존 성능 지표를 넘어 행동적 평가의 중요성을 강조합니다.

Original Abstract

Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization. We introduce correctness agreement, a decision-level metric that measures overlap in correct predictions between a base model and its quantized variants, independent of absolute accuracy. Across multiple models and quantization schemes from 8-bit to 2-bit, we find that behavioral divergence emerges under moderate quantization even when task performance appears preserved. To explain this effect, we analyze quantization as a structural operator on attention weights and quantify layer-wise distortions using statistical and distributional measures. Our results reveal non-linear breakpoints at low bit-widths and show that query and key projections are consistently more sensitive than value and output projections. These findings expose an illusion of equivalence between base and quantized models and motivate behavioral evaluation beyond conventional performance metrics.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!