2606.25519v1 Jun 24, 2026 cs.AI

양자화가 추론 능력을 저해한다: 저비트 추론 모델의 숨겨진 비용으로서의 토큰 증가

Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models

Minjia Zhang
Minjia Zhang
Citations: 25
h-index: 4
Walid Krichene
Walid Krichene
Citations: 2,358
h-index: 21
Xinyu Lian
Xinyu Lian
Citations: 66
h-index: 5
Beichen Huang
Beichen Huang
Citations: 11
h-index: 2
Masahiro Tanaka
Masahiro Tanaka
Citations: 396
h-index: 7
Olatunji Ruwase
Olatunji Ruwase
Citations: 13,468
h-index: 25
Li Zhang
Li Zhang
Citations: 1
h-index: 1

양자화는 대규모 언어 모델의 추론 비용을 줄이는 데 널리 사용되지만, 그 효과는 최종 답변 정확도나 토큰당 지연 시간만으로는 충분히 설명되지 않습니다. 본 연구에서는 저비트 사후 양자화가 숨겨진 테스트 시간 계산 비용을 유발할 수 있음을 보여줍니다. 양자화된 추론 모델은 여전히 정답을 맞추더라도, 더 긴 사고 과정을 생성하는 경향이 있습니다. 수학적 추론, 코드 생성, 과학 질문 답변 및 에이전트 기반 도구 사용 벤치마크를 통해 INT4/INT3 양자화가 정확도를 유지할 수 있지만, 추론에 필요한 토큰 수를 증가시켜 예상되는 토큰당 속도 향상을 상쇄할 수 있음을 확인했습니다. 이러한 효과를 측정하기 위해, 본 연구에서는 양자화된 모델과 전체 정밀도 모델 간의 추론 길이를 비교하는 CoT 토큰 증가 비율을 도입했습니다. 또한, 토큰 증가는 추론 과정에서의 행동 변화와 함께 발생하며, 이는 더 많은 중간 단계와 의미 중복을 포함합니다. 이러한 변화는 실제 서비스 환경에서 측정 가능한 성능 저하를 초래합니다. 마지막으로, 완화 전략을 평가한 결과, 프롬프트 및 디코딩 시간 샘플링은 일관성 없는 정확도-길이 균형을 제공하는 반면, 양자화 인식 학습은 정확도 저하와 토큰 증가 모두를 줄이는 데 더 유망한 것으로 나타났습니다. 본 연구의 결과는 양자화된 추론 모델을 평가할 때 정확도뿐만 아니라 추론에 사용되는 토큰 수도 함께 보고해야 함을 시사합니다.

Original Abstract

Quantization is widely used to reduce the inference cost of large language models, but its effect on reasoning models is not fully captured by final-answer accuracy or per-token latency. We show that low-bit post-training quantization can introduce a hidden test-time compute cost: quantized reasoning models often generate longer chains of thought even when they still answer correctly. Across mathematical reasoning, code generation, scientific question answering, and agentic tool-use benchmarks, we find that INT4/INT3 quantization can preserve accuracy but increase reasoning-token usage, offsetting the expected per-token speedup. To measure this effect, we introduce the CoT Token Inflation Ratio, which compares reasoning length between quantized and full-precision models averaged across all evaluation benchmarks. We further show that token inflation is accompanied by behavioral changes in the reasoning trace, including more intermediate steps and greater semantic repetition. These changes translate into measurable end-to-end real-world serving penalties. Finally, we evaluate mitigation strategies and find that prompting and decoding-time sampling offer inconsistent accuracy-length trade-offs, while quantization-aware training shows more promise in reducing both accuracy degradation and token inflation. Our results suggest that reasoning-token usage should be reported alongside accuracy when evaluating quantized reasoning models.

0 Citations
0 Influential
12.5 Altmetric
62.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!