2607.27275v1 Jul 29, 2026 cs.LG

플랫 스코어, 증폭된 오류: 양자화 LLM 에이전트에서 오차 예산이 손상을 가리는 방식

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

Heuiseok Lim
Heuiseok Lim
Citations: 33
h-index: 2
Jiwon Jang
Jiwon Jang
Citations: 0
h-index: 0
Kisu Yang
Kisu Yang
Citations: 265
h-index: 8
Hyunwoo Park
Hyunwoo Park
Citations: 343
h-index: 2

훈련 후 양자화를 통해 4비트로 무게를 줄이는 것은 거의 손실이 없다고 널리 보고되어 있습니다. 본 연구에서는 다중 회선, 도구 호출 에이전트에 대해 이 주장을 검증합니다. $τ^2$-벤치마크에서 밀집형 및 MoE 변형의 두 가지 오픈 웨이트 모델 패밀리에 대해, 그리고 두 가지 영역(각 영역당 8개의 셀, 456개 에피소드)에서 16비트, 8비트 및 4비트 무게로 양자화를 수행한 결과, 표준 지표에서는 양자화가 실제로 손실이 없는 것처럼 보입니다. 어떤 셀에서도 통계적 다중 비교 교정을 통과하는 성능 변화는 나타나지 않았으며, 가장 큰 프로세스 손상을 보이는 셀에서 등가성 테스트를 통해 변화 폭은 ±7.5점으로 제한됩니다. 하지만 실제로는 양자화가 모델이 원래 정밀도에서 이미 보여주는 오류(예: 통신 분야의 도구 이름 환각 현상, 그리고 소매 분야의 개체 오류에서도 동일한 경향)를 최대 2.5배까지 증폭시킵니다 (+17.6점/작업), 새로운 유형의 오류는 거의 발생시키지 않습니다. 모든 정밀도에서 오류 세트는 동일하며(랭크 상관관계 ≥ 0.94, 새로운 이벤트 비율 0.18%), 표준 지표에서는 이러한 오류가 감춰집니다. 이는 벤치마크의 10개의 오류 예산이 추가적인 오류를 흡수하기 때문입니다. 오류 예산을 2개로 줄이면 17점의 성능 격차가 다시 드러나며, 이는 양자화로 인해 오류량이 증가한 특정 셀에서만 나타납니다. 이는 마스킹 효과에 대한 예측과 일치합니다. 또한, 특정 오류 수정을 위한 프롬프트를 사용하여 다섯 개의 통신 모델을 모든 정밀도에서 실행한 결과, 손상이 있는 부분에서만 정확하게 손상을 제거할 수 있었습니다. 채널별 오류율 및 축소된 예산 하에서의 성공 여부는 기존 벤치마크가 이미 수집하는 로그 데이터로부터 얻을 수 있으며, 이러한 지표를 작업 보상과 함께 보고할 것을 제안합니다.

Original Abstract

Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it now matters most. On $τ^2$-bench, across two open-weight model families in dense and MoE variants and two domains (eight cells, 456 episodes each, at 16-, 8-, and 4-bit weights), quantization indeed looks free on the standard metric. No cell shows a score change that survives multiple-comparison correction, and in the cell that carries the largest process damage, equivalence testing bounds the change within $\pm$7.5 points. The process tells a different story. Quantization amplifies the failure the model already exhibits at full precision (tool-name hallucination in telecom, with the same directional trend in retail entity errors) by up to 2.5$\times$ in volume (+17.6 points per task), while creating essentially no new failures. The failure set is the same at every precision (rank correlation $\geq$ 0.94, 0.18% novel events). The score stays flat because the benchmark's ten-error budget absorbs the extra failures. Shrinking the budget to two errors re-exposes a score gap of 17 points, and it does so only in the one cell where quantization added error volume, exactly as the masking account predicts. A targeted error-repair prompt, run for five telecom models at every precision, removes the damage exactly and only where it lives. Both diagnostics, the per-channel error rate and success under a shrinking budget, come from logs benchmarks already collect; we suggest reporting them alongside task reward.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!