2608.01631v1 Aug 03, 2026 cs.CL

정확성이 증거와 동일한가? KV 캐시 압축 환경에서의 추론 충실성

Does Accuracy Equal Evidence? Reasoning Faithfulness under KV Cache Compression

Mengting Ai
Mengting Ai
Citations: 195
h-index: 6
Jingrui He
Jingrui He
Citations: 213
h-index: 9
Yue Guo
Yue Guo
Citations: 0
h-index: 0

KV 캐시 압축은 일반적으로 최종 답변의 정확도를 기준으로 평가되는데, 이는 답변을 유지하는 것이 해당 답변을 뒷받침하는 추론 과정도 함께 보존한다는 전제를 내포합니다. 본 연구에서는 대규모 추론 모델에 대해 이 가정이 성립하지 않을 수 있음을 보여줍니다. 압축 과정에서 올바른 답변과 그 답변을 뒷받침하는 명시적인 근거의 유효성이 서로 다른 비율로 유지될 수 있습니다. 우리는 제어된 고정 트레이스 재현 프로토콜을 사용하여 이러한 현상을 연구했습니다. 이 프로토콜은 추론 내용을 고정한 상태에서, 압축이 이미 존재하는 트레이스의 유용한 정보를 얼마나 잘 보존하는지 분리하여 분석합니다. 본 연구에서는 수학적 추론, 과학 질의응답, 임상 계산 및 긴 문맥 검색을 수행하는 세 가지 모델에 대해 열 개의 토큰 삭제 KV 압축 방법과 하나의 양자화 방법을 평가했습니다. 최종 정확도, 답변 체인 일관성 및 교란 충실성을 측정했으며, 다양한 작업에서 토큰 삭제 방식은 경쟁력 있는 최종 답변 정확도를 유지하면서 추론 체인 지원 또는 교란 충실성이 크게 저하되는 경향을 보였습니다. 우리는 이를 '답변-증거 간 격차'라고 부릅니다. 반면, 정보 손실을 최소화하는 양자화 제어는 이러한 현상에 덜 영향을 받는 것으로 나타났습니다. 이는 실패가 KV 메모리 감소 자체보다는 추론 과정의 일부에 대한 접근성 상실과 관련되어 있음을 시사합니다. 코드 및 추가 자료는 https://github.com/famous-blue-raincoat/Safe_KV_Compress 에서 확인할 수 있습니다.

Original Abstract

KV cache compression is commonly evaluated by final-answer accuracy, implicitly assuming that preserving the answer also preserves the reasoning that supports it. We test this assumption for large reasoning models and show that it can fail: under compression, correct answers and the validity of their visible supporting rationales can be preserved at different rates. We study this failure with a controlled fixed-trace replay protocol, which holds reasoning content fixed and isolates whether compression preserves usable information from an already available trace. We evaluate ten token-eviction KV compression methods and one quantization method on three models across mathematical reasoning, scientific QA, clinical calculation, and long-context retrieval. We measure final accuracy, answer-chain consistency, and perturbation faithfulness. Across tasks, token-eviction methods can preserve competitive final-answer accuracy while substantially degrading chain support or perturbation faithfulness. We call this the answer-evidence gap. A coverage-preserving quantization control is substantially less affected, suggesting that the failure is tied less to KV memory reduction itself than to losing access to parts of the reasoning trace. Code is available at https://github.com/famous-blue-raincoat/Safe_KV_Compress.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!